• Home  
  • Nearly half of cross-sectional social science papers used causal language their designs could not establish
- Science

Nearly half of cross-sectional social science papers used causal language their designs could not establish

An analysis of 194,631 cross-sectional social-science papers found causal language in 46.3% of titles or abstracts, while experiments showed that both readers and AI summaries can interpret these claims more strongly than the underlying designs support.

Research papers, charts and a magnifying glass illustrating the difference between correlation and causation.

Scientific papers often use cautious language for a reason.

If a study finds that two things occur together, researchers can usually say they are associated. Saying that one thing caused the other requires stronger evidence.

New research suggests that this distinction is being blurred surprisingly often across the social sciences.

A study published in Nature Human Behaviour analysed 194,631 cross-sectional social-science articles published between 1980 and 2024. The researchers found that 46.3% contained causal language in their titles or abstracts, despite using study designs that generally cannot establish causality on their own.

The pattern has also become more common. Around 20% of the cross-sectional studies examined used causal language in 2000. By 2024, the proportion had risen to more than 60%.

The researchers then tested whether this wording actually changes how people understand research. It did. Readers were more likely to interpret studies as providing causal evidence when abstracts used causal phrasing, while clearer methodological labels and more explicitly associational wording reduced that tendency.

Large language models introduced an additional problem. When asked to summarise research, several models sometimes removed cautious wording or introduced stronger causal claims than appeared in the original study.

Cross-sectional studies can show patterns, but causality is harder

A cross-sectional study typically measures variables at one point in time.

Researchers might find, for example, that people who use a particular technology more frequently also report higher stress, or that people with one behavioural characteristic tend to report another.

These findings can be useful. They can reveal relationships worth investigating, identify groups that differ from one another and help researchers generate theories.

What they usually cannot establish on their own is whether one variable caused the other.

There may be reverse causality, where the supposed outcome actually influences the supposed cause. A third variable may influence both. The relationship may also reflect selection effects or other differences between the groups being compared.

This is why scientific writing normally distinguishes phrases such as “was associated with” from stronger statements such as “increased”, “reduced”, “led to” or “caused”.

The researchers screened almost 195,000 papers

Calvin Isch and colleagues developed a large-scale classification system to examine how often cross-sectional studies crossed that linguistic boundary.

They searched five ProQuest databases for social-science papers using cross-sectional methods and then used a combination of large language models, expert human coding and a fine-tuned BERT classifier to identify both study design and causal language.

The final dataset contained 194,631 studies that relied exclusively on cross-sectional designs rather than combining cross-sectional data with longitudinal, experimental or quasi-experimental methods.

The researchers focused primarily on titles and abstracts because these are often the parts of a paper that receive the most attention from readers, journalists, students and other researchers.

Nearly half used causal language

Across the full dataset, 46.3% of the cross-sectional studies contained causal language in the title or abstract.

The researchers describe these statements as overreaching because the underlying designs generally provide limited leverage for identifying causal effects.

The figure was not evenly distributed across disciplines.

By 2024, business research showed the highest prevalence in the dataset, with 84% of cross-sectional articles using causal language. Economics followed at 67%, psychology at 55%, political science at 53% and sociology at 46%.

These figures do not mean that the findings in those papers are false. A causal relationship may genuinely exist.

The problem is whether the study design presented in the paper can justify the strength of the claim being made.

Causal wording has increased sharply since 2000

The historical trend was one of the most striking findings.

From roughly 1980 to 2000, around one in five cross-sectional papers in the dataset used causal language.

After 2000, the proportion began rising rapidly. By the 2020s, it had roughly tripled.

The increase appeared across all five disciplines examined.

The study does not establish exactly why this happened. Possible explanations include changing publication incentives, pressure to communicate stronger implications, shifts in academic writing conventions and growing emphasis on practical impact.

Importantly, the researchers also found that causal overclaiming was not confined to lower-status journals. The pattern appeared in prestigious outlets as well.

The wording changed what readers thought the study proved

The researchers did not stop at counting words.

They recruited 1,105 college-educated adults in the United States and randomly assigned them to read one of several versions of research abstracts taken from elite journals.

Some participants saw the original abstract containing causal language. Others saw a version rewritten using strictly associational wording.

Another group received the original abstract together with a methodological note explaining that the research was cross-sectional and therefore could not, on its own, establish causality.

A fourth condition included the methodological note together with AI-generated feedback highlighting possible methodological and interpretive problems.

Readers exposed to the causal wording were more likely to agree that the research established a causal relationship. Rewriting the abstract in associational terms and adding a methodological warning both reduced that interpretation.

Even cautious wording did not completely solve the problem

One of the more important findings was that people showed a strong tendency to think causally even when the wording became more careful.

Participants frequently reintroduced causal language when they were asked to summarise the study themselves.

This suggests that causal overstatement may not be caused only by careless academic writing.

Readers may naturally prefer causal explanations because they are simpler and more useful for understanding the world. Saying that two variables are merely associated leaves uncertainty. Saying that one produces the other creates a much clearer story.

The problem is that a clearer story can also be a less accurate one.

AI summaries sometimes made the claims stronger

The researchers also examined what happens when large language models become intermediaries between scientific papers and readers.

Five language models were asked to summarise a stratified sample of research articles using different prompting styles.

The models sometimes strengthened the causal interpretation of the original material.

Conditional language and hedges could disappear in the summary. In some cases, models introduced an unqualified causal claim even when the original abstract had used strictly associational wording.

This matters because many readers now use AI tools to explain academic papers, simplify technical findings or translate research into practical recommendations.

If the summarisation process quietly changes “is associated with” into “causes”, the scientific conclusion has changed even if the summary still sounds plausible.

The prompt given to the AI made a difference

The study also offers a practical lesson for people who use AI to interpret research.

The models were tested with several types of prompts, including basic summaries, simplified explanations, practical interpretations and a prompt explicitly asking the model to be careful and sceptical about causal claims.

The careful prompt reduced causal overstatement.

By contrast, prompts asking for simplified or practical summaries tended to produce stronger causal language more often.

This creates an important trade-off. The very prompts people use to make research easier to understand may also encourage the model to remove uncertainty and qualification.

“Practical” summaries can be especially vulnerable

Real-world recommendations require a model to move from description toward action.

If a study finds that two things are associated, a practical summary may be tempted to convert that relationship into advice about changing one to improve the other.

But that recommendation assumes causality.

For example, if a cross-sectional study finds that people who sleep more also report greater wellbeing, it does not necessarily prove that increasing sleep will produce the observed difference. Health, income, employment, stress and many other factors could influence both variables.

A useful summary therefore needs to preserve the distinction between what the study observed and what an intervention has actually been shown to change.

This matters beyond academic journals

Most people never read an entire research paper.

Scientific findings are filtered through abstracts, university press releases, news reports, social-media posts, textbooks, presentations and increasingly AI-generated explanations.

Every layer creates another opportunity for a cautious statistical relationship to become a confident causal statement.

That can affect how people understand health, education, business, technology and public policy.

A correlational finding presented as causal can make an intervention appear more certain than the evidence supports, potentially influencing decisions about money, behaviour or policy.

The study is also a warning for science communicators

Making research accessible often requires simplifying it.

But simplification and exaggeration are not the same thing.

A sentence such as “social-media use is associated with anxiety” is less dramatic than “social media causes anxiety”, but the two statements communicate fundamentally different evidence.

For journalists, universities and research websites, preserving that distinction is part of accurate reporting rather than unnecessary academic caution.

The same principle applies to headlines. A stronger verb may produce a cleaner headline while simultaneously making the scientific claim less defensible.

There are limitations to the analysis

The researchers used automated systems to classify a very large body of academic writing, so classification error is unavoidable.

However, the models were validated against expert human coding and showed strong agreement.

The paper also focuses on cross-sectional studies because these designs generally offer weak causal identification. There can be unusual cases in which additional assumptions, methods or context allow stronger inferences than the design label alone suggests.

Likewise, the presence of causal language does not prove that authors intentionally exaggerated their work.

Some wording may reflect disciplinary convention, imprecise phrasing or a belief that previous evidence supports the implied mechanism.

The study is therefore best understood as measuring a mismatch between the strength of language and the identification power of the design, not as evidence of deliberate misconduct.

What this means for students and researchers

The findings offer a straightforward reading habit: whenever a paper says that one thing increased, reduced, influenced or caused another, ask what research design produced that conclusion.

Randomised experiments, strong natural experiments and carefully designed longitudinal or quasi-experimental studies can sometimes support causal inference.

A single cross-sectional association usually cannot.

For researchers, the study also reinforces the value of writing conclusions that match the evidence actually collected.

That may mean choosing a less dramatic verb, but it also makes the finding more scientifically precise.

AI can help interpret research, but it needs the right instruction

The AI results do not suggest that language models are inherently incapable of handling causal uncertainty.

In fact, the models improved when explicitly instructed to preserve distinctions between association and causation.

That makes the problem partly one of workflow.

Instead of asking an AI system simply to “summarise this study”, a more reliable request would ask it to identify the study design, distinguish associations from causal evidence and avoid strengthening claims beyond the original methods.

The wording of the prompt can therefore become part of research literacy.

The bigger lesson is about certainty

Scientific language often sounds cautious because science is built around uncertainty.

That caution can be frustrating when readers want a clear answer, but removing it does not make the evidence stronger.

The new study shows how easily an association can acquire the language of causation as research moves from paper to reader and from paper to AI summary.

In a scientific information environment increasingly shaped by automated summaries, preserving small words such as “associated”, “may” and “could” may matter more than ever.

Source Information

Study Title: Quantifying the prevalence and impact of overreaching causal claims in social science
Authors: Calvin Isch, Timothy Dörr, Neil Fasching, Grace Jennings and Duncan J. Watts
Journal: Nature Human Behaviour
Published: 24 August 2026
Dataset: 194,631 cross-sectional social-science articles published between 1980 and 2024, a reader experiment involving 1,105 college-educated US adults, and experiments with five large language models.
Method: The researchers combined automated classification with expert human validation to identify causal language, experimentally tested how wording and methodological warnings affected readers’ interpretations, and examined whether AI-generated summaries strengthened or weakened causal claims.
Main finding: 46.3% of cross-sectional studies contained causal language in their titles or abstracts, with annual prevalence rising from around 20% in 2000 to more than 60% by 2024. Causal phrasing increased readers’ causal interpretations, while AI summaries sometimes amplified causal overstatement.
DOI: 10.1038/s41562-026-02553-x

Research Today is a South African digital publication that makes credible research easier to understand.

 

ResearchToday.bus@gmail.com

TERMS OF USE & PRIVACY POLICY

follow us