• Home  
  • Correlation is increasingly being written like causation, study of 194,631 papers finds
- Education

Correlation is increasingly being written like causation, study of 194,631 papers finds

Researchers analysed 194,631 cross-sectional social-science papers and found causal language in nearly half, with the rate climbing above 60% by 2024. A separate experiment showed that the wording can change what readers believe the evidence proves.

“People who exercise more are happier.”

That sentence sounds simple, but scientifically it leaves an important question unanswered.

Does exercise make people happier, or are happier people simply more likely to exercise?

New research suggests social-science papers are increasingly using language that can blur this distinction.

Researchers analysed 194,631 cross-sectional studies published between 1980 and 2024 and found that 46.3% used causal language in their titles or abstracts, despite relying on research designs that generally cannot establish cause and effect on their own.

The trend has also accelerated dramatically.

Around 20% of the papers used causal language at the beginning of the 2000s. By 2024, the proportion had risen to more than 60%.

The study, published in Nature Human Behaviour, suggests that one of science’s most familiar warnings remains surprisingly relevant:

Correlation does not automatically mean causation.

Seeing two things together does not tell you which one caused the other

Cross-sectional studies are extremely common across fields such as psychology, business, economics, political science and sociology.

They usually examine people, organisations or other observations at a particular point in time.

A researcher might survey 2,000 adults and discover that people who spend more time on social media report greater loneliness.

That relationship can be important.

But the study alone cannot tell us whether social media increased loneliness.

Lonely people may use social media more frequently. A third factor, such as poor health or social circumstances, might influence both. Several explanations can fit exactly the same statistical association.

This is why establishing causality usually requires stronger evidence about timing and alternative explanations, often using experiments, natural experiments, longitudinal data or other methods designed specifically for causal inference.

A single snapshot can show that two things move together.

It is much less capable of proving which one moved the other.

The researchers examined almost 200,000 papers

Calvin Isch and colleagues at the University of Pennsylvania searched large academic databases covering five areas of social science.

They first identified studies that relied exclusively on cross-sectional data and excluded papers containing longitudinal or experimental components.

The researchers then examined the titles and abstracts of those papers for language implying that one variable caused, changed, increased, reduced or otherwise affected another.

Because manually reading almost 200,000 articles would be impractical, the team developed automated classifiers and validated them against expert human judgements.

The final dataset contained 194,631 cross-sectional studies.

Across those papers, 46.3% contained language the researchers classified as making or implying a causal claim about the study’s own results.

The problem has become much more common

The historical trend was one of the study’s most striking findings.

From roughly 1980 until 2000, the percentage of cross-sectional papers using causal language remained near 20%.

After 2000, it began climbing rapidly.

By the 2020s, the rate had approximately tripled, reaching more than 60% by 2024.

The increase appeared across every discipline the researchers examined.

That suggests the effect is not confined to one unusual corner of social science.

Instead, stronger language appears to have become increasingly common across a substantial part of the literature.

Business research had the highest rate

The disciplines were not identical.

By 2024, the researchers found causal language in approximately 84% of the cross-sectional business papers included in their analysis.

Economics followed at 67%.

Psychology stood at approximately 55%, political science at 53% and sociology at 46%.

These figures should not be interpreted as meaning that 84% of all business research is wrong.

The analysis concerned a specific subset: papers using cross-sectional designs.

Business research also includes experiments, longitudinal studies, econometric designs and many other methods capable of addressing causal questions more convincingly.

But among the cross-sectional articles the researchers identified, causal wording was particularly common.

Prestigious journals were not immune

One might expect higher-impact journals to be more cautious about this kind of language.

The researchers found the opposite association.

Among journals with impact factors above the median in their dataset, 54.4% of cross-sectional abstracts contained causal language.

Among journals below the median, the figure was 43.3%.

The researchers do not establish why this relationship exists.

But it raises an interesting possibility.

Academic publishing rewards findings that sound important.

“X causes Y” naturally sounds more decisive than “X was associated with Y”.

The second statement may be scientifically more appropriate for a particular study, but the first can produce a much stronger headline.

The pressure to tell a compelling story may therefore conflict with the need to communicate uncertainty precisely.

Readers actually believed the stronger wording

The researchers then tested whether this language makes a practical difference.

They recruited 1,105 college-educated adults in the United States and gave them abstracts from cross-sectional studies published in prominent journals.

Some participants saw the original abstract containing causal language.

Others saw a rewritten version that described the findings only as associations.

Another group received the original abstract together with a short note explaining that the study was cross-sectional and therefore could not establish causality on its own.

Readers exposed to the causal wording were more likely to conclude that the research had demonstrated a genuine cause-and-effect relationship.

Changing the wording or clearly identifying the methodological limitation reduced that tendency.

The language therefore did more than make the paper sound stronger.

It changed what readers thought the evidence had actually shown.

One word can substantially change the meaning

Consider two hypothetical headlines:

“Working from home increases productivity.”

And:

“Working from home is associated with higher productivity.”

They may sound almost interchangeable in everyday conversation.

Scientifically, they make different claims.

The first implies that changing where someone works will change their productivity.

The second says only that the two were observed together.

If the underlying research surveyed workers once, the second formulation is usually the safer interpretation.

Remote workers may have different occupations, employers, personalities, working conditions or levels of seniority. Any of those differences could contribute to the observed relationship.

Small changes in wording therefore carry surprisingly large assumptions.

This matters once research leaves the journal

Most people will never read the methods section of an academic paper.

They encounter research through abstracts, university press releases, newspaper articles, social-media posts and headlines.

That means the title and abstract have disproportionate influence.

If the original paper already describes an association as though it were causal, every additional layer of summarisation creates another opportunity for the claim to become even stronger.

“Researchers found an association” can gradually become “scientists prove” by the time the finding reaches the public.

This is especially important when research influences health choices, education, business decisions or public policy.

A policy designed around a correlation may fail if changing the supposed cause does not actually change the outcome.

South African research readers should recognise the problem

The issue is highly relevant in South Africa, where surveys are frequently used to understand consumer behaviour, education, health, employment and public opinion.

A study might find that financially satisfied consumers report greater loyalty to their bank.

That does not automatically prove that increasing financial satisfaction will cause loyalty.

Loyal customers may evaluate the bank more positively in the first place, while income, service quality or previous experiences could influence both.

Likewise, if students who spend more time studying achieve higher marks, the relationship does not tell us exactly how much additional studying would improve the marks of a particular learner.

None of this makes observational research useless.

Associations are often the first indication that something important deserves further investigation.

The problem begins when the conclusion becomes stronger than the design that produced it.

This does not mean you should distrust every observational study

There are several important qualifications.

The researchers themselves acknowledge that associational evidence can contribute to a larger causal argument when combined with theory, previous experiments or other evidence.

Not every causal-sounding statement is necessarily false simply because one study within that argument is cross-sectional.

The analysis also relied partly on automated classification.

Although the models were validated against expert reviewers and performed strongly, language can be ambiguous and no classifier will interpret every paper perfectly.

The dataset was drawn from five major academic databases and should therefore not be treated as a census of every social-science article ever published.

Most importantly, the study examined language.

It did not attempt to determine whether the underlying substantive conclusions of all 194,631 papers were ultimately correct or incorrect.

The easiest defence is surprisingly simple

For ordinary readers, the research suggests a useful habit.

Whenever an article says that something “causes”, “increases”, “reduces” or “leads to” something else, ask one additional question:

How did the researchers actually know that?

If people were randomly assigned to different conditions, there may be a strong basis for causal interpretation.

If the researchers followed changes over time or exploited a natural experiment, the case may also be persuasive.

If they simply surveyed people once and discovered that two variables were related, more caution is required.

The difference sounds technical.

It is actually one of the most important distinctions in understanding research.

Science is not weakened when researchers say that two things are associated.

Sometimes that is exactly what the evidence shows.

The danger begins when a more exciting sentence quietly claims that it showed something more.

Source Information

Study Title: Quantifying the prevalence and impact of overreaching causal claims in social science
Authors: Calvin Isch, Timothy Dörr, Neil Fasching, Grace Jennings and Duncan J. Watts
Journal: Nature Human Behaviour
Published: 24 August 2026
Cross-sectional articles analysed: 194,631
Human experiment: 1,105 participants
Period analysed: 1980–2024
DOI: 10.1038/s41562-026-02553-x

Research Today is a South African digital publication that makes credible research easier to understand.

 

ResearchToday.bus@gmail.com

TERMS OF USE & PRIVACY POLICY

follow us