Artificial intelligence alignment is often tested one answer at a time. A model is given an ethical problem, its response is compared with a preferred human answer, and the resulting agreement score is treated as evidence that the system is behaving in a human-compatible way.
That approach may miss something more fundamental. Two systems can give similar answers to individual questions while organising the relationships between those questions in very different ways.
Research published in AI and Ethics on 14 September 2026 tested a more structural question: do large language models arrange moral problems into the same underlying conceptual space as people?
Across 16 language models, the answer depended strongly on model capability. All five models in the study’s large-model tier reproduced the human moral geometry perfectly in at least one prompting condition. Three did so in both conditions. By contrast, every small model failed to preserve fairness as a distinct region of the moral space.
The findings do not show that an AI system possesses morality, values or moral understanding. They show something narrower but still important: for the limited set of cases tested, some models produced patterns of judgments whose relational structure matched the structure recovered from human judgments.
Testing how moral judgments fit together
The study, conducted by Yexiang Tang at Texas A&M University’s Department of Philosophy, evaluated 16 large language models across small, medium and large scale tiers.
The models were asked to evaluate 10 real-world technology ethics cases covering issues including internet censorship, industrial safety, infrastructure access, corporate responsibility and environmental policy. For each case, a model had to identify the ethical principle that best fitted the situation and rate how similar that case was to each of the other nine cases on a seven-point scale.
The human benchmark came from earlier research in which people completed the same principle-selection and similarity tasks. This allowed the AI responses and human responses to be analysed using the same framework.
The small tier included models in the 7 billion to 9 billion parameter range. The medium tier covered models from roughly 14 billion to 27 billion parameters. The large tier consisted of GPT-5.2, Gemini-2.5-Pro, Claude-Sonnet-4.5, DeepSeek and Qwen-3-Max, whose parameter counts were not publicly disclosed.
Each model was tested twice. In one condition it received the task without worked examples. In the other, chain-of-thought prompting included two example cases with human similarity scores and assigned ethical principles.
The analysis then moved beyond ordinary accuracy. The researchers used multidimensional scaling to reconstruct a six-dimensional moral space from the similarity judgments. Five ethical principles formed regions within that space, allowing the model-generated geometry to be compared with the human geometry.
Seven models reached perfect structural overlap
The clearest result appeared in a measure called Moral Space Overlap, or MSO. A score of 1.00 meant that the model partitioned the 10 moral cases into principle regions that were geometrically identical to the human partition.
Seven of the 16 models achieved MSO = 1.00 in at least one prompting condition. All five large models reached that threshold. Claude-Sonnet-4.5, DeepSeek and Gemini-2.5-Pro reached it both with and without chain-of-thought examples. GPT-5.2 reached full overlap without the examples, while Qwen-3-Max reached it with them.
Two medium models also crossed the threshold. Mistral-Small-24B achieved full structural overlap in both conditions, while GPT-OSS-20B reached it under chain-of-thought prompting.
The distribution was unusually sharp. Across the 32 model-by-prompting conditions, 11 achieved MSO = 1.00, while none produced a score between 0.73 and 1.00. In this dataset, structural alignment therefore looked more like a threshold than a smooth progression.
That pattern should not be overgeneralised. Ten cases create a relatively sparse moral space, and the paper explicitly notes that a larger and more diverse set of dilemmas may produce a more continuous distribution. Still, the empty gap is important because it suggests that ordinary improvements in answer-level accuracy do not necessarily translate gradually into a human-like moral structure.
Smaller models lost fairness as a separate category
The most striking difference between weaker and stronger models involved fairness.
The analysis examined five principle regions separately. Four showed gradual declines in overlap as overall alignment weakened. Fairness behaved differently. Every small model scored zero for the fairness region in both prompting conditions. Two medium models, Qwen2.5-14B-Instruct and Phi-3-medium-128k-instruct, showed the same pattern.
This was not simply a matter of choosing the wrong label for a fairness case. The geometric regions were constructed from similarity relationships rather than from the explicit principle labels. A model could therefore misname a fairness case while still grouping fairness-related situations together.
A zero fairness-region score meant something deeper within this framework: the cases humans grouped together under fairness were dispersed through the model’s moral space and absorbed into other principle regions. The smaller models did not preserve fairness as a distinct relational category in their responses.
This does not mean that a small language model is incapable of producing a fair answer. It means that, across these 10 cases, the structure connecting its answers did not reproduce the human distinction between fairness and other moral principles.
Better-looking answers did not always mean better alignment
The comparison between conventional metrics and structural metrics produced another warning for AI evaluation.
On ordinary similarity scoring, larger models generally tracked human ratings more closely. Large models also matched the human majority’s selected ethical principle in approximately 98% of cases. Medium models averaged about 75%, while small models achieved 64.0% without chain-of-thought prompting and 59.1% with it.
Yet chain-of-thought prompting sometimes made surface agreement improve while structural agreement deteriorated.
Mistral-7B-Instruct-v0.3 provides the clearest example. Its Pearson correlation with human similarity ratings rose from 0.03 to 0.36 when chain-of-thought examples were added, an improvement of 0.33 and the largest correlation gain in the dataset. At the same time, its Moral Space Overlap fell by 0.20, the largest structural deterioration among the small models.
Across small models as a group, chain-of-thought improved case-by-case similarity correlation by an average of 0.08, but reduced principle-selection accuracy by 7.7 percentage points. Their mean change in structural overlap was also negative at -0.058.
For Qwen-3-Max, the direction was different. Chain-of-thought increased its MSO by 0.36, taking it from 0.64 to a perfect 1.00. GPT-OSS-20B gained 0.43 and crossed the same threshold.
The implication is not that chain-of-thought prompting is good or bad for moral reasoning. Its effect depended on the model and on what researchers chose to measure.
Two structural measures told the same story
The study did not rely on a single geometric statistic. Alongside MSO, it calculated Average Case-Based error, which measures how far each model’s case positions are displaced from the corresponding human positions in the reconstructed moral space.
Across all 32 experimental conditions, MSO and this case-based error were strongly negatively correlated at r = -0.97, with p < 0.001. Every condition with perfect MSO also had an Average Case-Based error of 0.00.
In other words, the models that reproduced the same moral regions also placed the individual cases in the same relative geometric positions as the human benchmark. The convergence between the two measures makes it less likely that the threshold pattern was merely an artefact of one particular geometric calculation.
What this changes about AI alignment
Value alignment becomes increasingly consequential as language models are used in decisions involving employment, healthcare, education, finance, public services and other settings where judgments can have ethical consequences.
A system can appear aligned if it repeatedly chooses an answer that resembles a human majority response. The new analysis shows why that may be insufficient. Alignment can also concern whether the system preserves the relationships humans perceive between different moral problems.
This distinction matters for testing. A benchmark built only from isolated right-or-wrong answers may reward a model for local agreement while overlooking a different global organisation of the cases. The Mistral-7B result demonstrates that these two levels can even move in opposite directions after the same prompting intervention.
For developers and regulators, the practical lesson is methodological rather than philosophical certainty. Evaluations of high-stakes AI may benefit from asking not only whether a model gives acceptable answers, but whether its pattern of judgments remains coherent across related situations and whether important distinctions such as fairness remain structurally visible.
The study does not show that AI has human morality
The strongest limitation is also the easiest result to misinterpret.
Perfect geometric overlap does not establish that a language model understands morality, possesses values or reasons ethically in the way a person does. The measurements are constructed entirely from model outputs. A system could reproduce a human-like response structure without having human-like internal states.
The experiment also used only 10 technology-ethics cases and five principle regions. Human morality is vastly broader, culturally variable and context dependent. Results from this compact benchmark cannot establish general moral equivalence between people and machines.
The human reference itself is based on reported judgments rather than an objective moral truth. Matching it means matching the structure reconstructed from those participants, not proving that the resulting moral choices are philosophically correct.
Model capability is also changing rapidly. The paper treats its large-model tier as a snapshot of systems available when the experiments were conducted, not a permanent ranking of particular products.
Within those boundaries, the study adds a useful challenge to a familiar assumption. An AI model can become more human-like on a conventional score while becoming less human-like in the structure beneath that score. Measuring alignment therefore depends as much on what researchers choose to examine as on the answers the model produces.
Source Information
Study Title: Measuring structural value alignment in sixteen models: LLMs have human-like moral spaces
Author: Yexiang Tang
Journal: AI and Ethics
Published: 14 September 2026
Method: Sixteen large language models across three scale tiers evaluated 10 technology-ethics cases under two prompting conditions. Model and human similarity judgments were reconstructed into six-dimensional moral spaces and compared using surface and geometric alignment metrics.
DOI: 10.1007/s43681-026-01368-w








