An artificial intelligence system can apologise, justify a recommendation, reject a harmful request and explain its answer in language that sounds strikingly moral. A new peer-reviewed philosophy paper argues that none of those performances, even when reliable, is enough to make the system a genuine moral partner.
Published in Philosophy & Technology on 28 September 2026, Saša Josifović’s article examines a conceptual assumption that sits beneath much contemporary AI alignment work: that morally acceptable behaviour can be learned from human feedback, preferences, discourse and other behavioural traces. The paper argues that this can produce useful norm-conforming behaviour while still falling short of moral agency in the stronger sense required for responsibility, obligation and mutual accountability.
The distinction matters because increasingly fluent AI systems invite users and institutions to treat them as if they understand, own and stand behind the reasons they provide. Josifović’s concern is not primarily that aligned systems will behave badly. It is that reliable moral-looking behaviour may be mistaken for moral commitment, encouraging people to trust machines in ways that blur where responsibility actually belongs.
Alignment can reproduce the appearance of morality
Modern alignment techniques often work by shaping observable behaviour. Reinforcement learning from human feedback rewards outputs that human evaluators prefer. Preference learning infers patterns in human choices. Constitutional approaches use articulated principles and model feedback to steer responses toward desired norms.
These methods can be practically valuable. They can make systems more civil, safer in ordinary interactions and more predictable under familiar conditions. But the paper asks whether success at this task should be interpreted as evidence that an AI has learned morality itself.
Josifović argues that the inference is too strong. Human moral discourse is internally diverse and culturally situated. Human behaviour is also an unreliable substitute for moral truth because people routinely act from habit, self-interest, fear or social pressure rather than principle. Statistical regularities in approval and aversion can reveal what people tend to endorse, but popularity and frequency do not by themselves determine what people ought to do.
An aligned model can therefore become very good at navigating the outward patterns associated with moral conversation without occupying the social position from which moral claims become binding. The paper describes this as a difference between normative performance and normative participation.
A reason is more than a sentence that sounds like one
The article builds its case using several traditions in moral philosophy and developmental psychology. It draws especially on Michael Tomasello’s account of shared intentionality, Hegelian recognition and the second-personal tradition associated with moral accountability.
Across these approaches, morality is not reduced to selecting the correct output. Moral agents participate in practices where reasons can be requested, challenged, revised and owned. People make commitments to one another, recognise claims made upon them and can be called to account when they fail to meet those commitments.
That makes an important distinction between generating a justification and being bound by it. A language model can produce a coherent explanation for why an action is wrong. The philosophical question is whether the generated reason functions as a commitment for the system itself, something it recognises as bearing on what it ought to do and for which it can answer to another party.
Josifović argues that present generative systems do not establish this stronger relationship merely by speaking fluently. Their outputs may fit the grammar of giving reasons while the system remains outside the reciprocal practice in which reasons are demanded, contested and sustained.
Reliability is not the same as normative stability
This distinction becomes especially important when an aligned system performs consistently. A model that reliably avoids harmful or offensive responses can appear not merely predictable but principled. The paper warns against treating those two qualities as interchangeable.
Behavioural alignment can generate regularity and, under favourable conditions, robust reliability. Normative stability is a different claim. It implies steadfastness grounded in commitments that can be owned within a relationship of accountability.
The practical difference may become visible when conditions change. Under distribution shift, adversarial prompting, altered objectives or new incentives, a system can depart from its familiar behaviour. If the original regularity came from optimisation rather than self-binding moral commitment, the departure is better understood as a failure of the behavioural proxy than as a betrayal by a moral partner.
This is why the paper’s central warning is about misplaced trust. Humans are highly responsive to language that signals concern, explanation and accountability. An AI that reproduces those signals can elicit expectations normally reserved for agents who can actually be held to their commitments.
Moral status and moral agency are separate questions
The argument deliberately avoids a broader claim that artificial systems could never deserve moral consideration. Philosophers who defend ethical behaviourism or relational approaches have argued that sufficiently sophisticated outward behaviour or social interaction might provide grounds for including artificial entities within the moral circle.
Josifović sets that debate aside. The article remains agnostic about whether behavioural or relational properties could support moral rights or some form of moral patienthood. Its narrower target is moral agency and partnership, particularly where society is tempted to transfer responsibility to the system.
An entity might, in principle, deserve consideration without being an appropriate bearer of blame, obligation or responsibility. Conversely, an interactive system can influence human behaviour and become socially significant without thereby becoming the kind of subject to whom people can appropriately address moral claims.
This separation prevents the argument from depending on a definitive theory of machine consciousness. The practical question is not whether an AI has an unknowable inner experience. It is whether normatively fluent performance gives institutions sufficient grounds to relocate human responsibility onto the machine. The paper’s answer for present generative systems is no.
The risk grows as the stakes rise
For low-stakes applications, simulated morality may be entirely useful. A system that filters abusive language, maintains polite conversation or follows straightforward procedural constraints does not need to be a moral agent for its behaviour to have practical value.
The problem changes when decisions concern autonomous weapons, critical infrastructure, medical triage or social governance. In such settings, a persuasive explanation can make a system appear to be the source of a judgment rather than a tool operating inside a chain of human design, deployment and oversight.
If institutions begin to speak as though the AI decided, believed, intended or accepted responsibility, accountability can become diffuse. Designers may point to model autonomy, deployers may point to algorithmic recommendations, and users may defer to a system whose language creates the impression of an independent moral viewpoint.
Josifović’s argument is that linguistic sophistication should not be allowed to perform this institutional transfer. Until second-personal answerability is established rather than inferred from behaviour, responsibility should remain attached to identifiable human agents and organisations.
From teaching machines morality to governing machine power
The paper therefore proposes a different way to frame AI alignment. Rather than treating alignment as ethical education for an artificial moral pupil, it should be understood as institutional design for constraining an artefact.
That shift changes the practical questions. Instead of asking whether a model has internalised human values, regulators and organisations should ask whether its actions are bounded, whether high-stakes outputs can be audited, whether decision trails are traceable, whether escalation routes are clear and whether liability remains enforceable.
On this view, technical alignment still matters. The point is not to abandon efforts to make AI safer or more reliable. It is to interpret what those techniques achieve more cautiously. Behavioural alignment can constrain machine behaviour without creating a responsibility-bearing moral subject.
The distinction also changes how failures are described. A harmful output need not be framed as an AI choosing immorally. It can instead be investigated as a failure of design, training, deployment, oversight or institutional control. That vocabulary keeps the search for accountability directed toward actors who can actually respond to claims, revise practices and bear obligations.
What the paper does and does not show
This is a philosophical research article, not an experiment. It has no participant sample, intervention, effect size, confidence interval or prevalence estimate. The paper develops a conceptual argument by examining alignment practices through theories of shared intentionality, recognition, second-personal accountability and moral agency.
That means it cannot establish how often people actually over-trust AI because of moral-sounding language, how users distinguish reliability from moral commitment, or whether particular alignment methods increase inappropriate responsibility attribution. Those are empirical questions that could be tested through behavioural, organisational and legal research.
The argument also depends on a substantive account of moral agency. Philosophers who place greater weight on outward behaviour or relational interaction may dispute whether shared intentionality and second-personal answerability are necessary in the way the paper proposes. Josifović engages that objection but does not eliminate the wider philosophical disagreement.
Nor does the paper claim that future artificial systems could never satisfy stronger conditions for moral partnership. Its conclusions are directed at the inference from normatively fluent performance to moral agency, especially for present generative architectures. Future systems with very different forms of identity, social participation or answerability could reopen the question.
The deeper issue is where responsibility stays
As AI becomes more persuasive, the temptation to describe it in human moral terms will become stronger. Systems can already apologise, offer reasons, express apparent concern and adapt their language to social expectations. Those capabilities matter for usability, but they can also make a statistical system look like a participant in moral life.
The paper’s contribution is to insist that good behaviour and moral partnership answer different questions. One concerns whether an artefact behaves in ways people find acceptable. The other concerns whether it stands inside relationships of recognition, commitment and accountability that make moral reasons binding.
For AI governance, that difference has a concrete consequence. The safer institutional assumption is not that a convincing moral voice carries its own responsibility. It is that the humans and organisations building, deploying and authorising the system remain responsible for the power exercised through it.
Source Information
Study: Simulated Morality, Misplaced Trust: The Risks of Treating AI as a Moral Partner
Author: Saša Josifović
Journal: Philosophy & Technology
Published: 28 September 2026
Volume and article: 39, Article 192
DOI: 10.1007/s13347-026-01180-8









