Automated video interviews promise employers a faster way to screen large applicant pools, but the language people naturally use in those interviews may change with age in ways that matter for automated assessment.
A peer-reviewed study published in the Journal of Business and Psychology on 25 September 2026 found systematic age-related differences in the language workers used when answering the same automated video interview questions. In a second phase, the researchers used AI-generated personas to test whether rewriting the interview prompts could reduce those differences before an assessment reached real applicants.
The results point to an important design problem for employers adopting AI in recruitment. A system can be technically consistent while still receiving systematically different linguistic inputs from people at different stages of adulthood. If those differences are unrelated to the capability a job actually requires, they can become a source of construct-irrelevant variation.
What changed as workers got older
Jingyi Li, Daphne X. Hou and Louis Tay recruited 373 working adults through CloudResearch Connect. Participants were aged from 18 to 66 years and older, worked full-time or part-time, and spoke English as their first language. After quality screening, 368 participants remained in the main analysis.
Each participant answered six open-ended interview questions on video. The questions covered subjects such as career aspirations, preferred work environments and work-related activities. Participants then completed measures of Big Five personality traits and vocational interests, alongside demographic questions covering age, gender, race, education and socioeconomic status.
The researchers transcribed the interviews and extracted 41 linguistic, semantic and timing features. These included vocabulary categories, topics identified through latent Dirichlet allocation, topic relevance, lexical diversity, idea density, syntactic complexity, response pacing and the time participants took before beginning an answer.
Age had a significant overall linear relationship with this combined set of features even after the analysis controlled for demographic characteristics, personality and vocational interests. The multivariate test produced a Pillai’s Trace of 0.26, F(40, 300) = 2.58, p < 0.001, with a partial eta-squared of 0.26. A quadratic age effect was not statistically significant, so the follow-up analysis focused on linear differences across adulthood.
Fourteen features showed significant linear age effects under the study’s false-discovery-rate correction. The strongest vocabulary effect involved references to the past. Past-focused language increased with age, with a standardised coefficient of β = 0.32 and an additional 8% of variance explained after the control variables were included.
Other patterns moved in the opposite direction. Cognitive-process language declined with age at β = -0.21, while curiosity-related language declined at β = -0.19. Family-related language increased at β = 0.16. The analysis also identified smaller age relationships involving money, home, achievement and work-related wording.
These are not simply differences between two arbitrarily defined age groups. Age was analysed continuously. The researchers nevertheless calculated age-group comparisons to make the magnitude easier to interpret. Older workers used fewer cognitive-process words than younger workers, with Cohen’s d = -0.59, and fewer long words, with d = -0.48.
The difference extended beyond word choice
The interviews also differed in how closely responses aligned with the question. Topic relevance decreased with age at β = -0.18. In descriptive comparisons, older workers’ answers were less topically relevant than those of middle-aged workers by d = -0.47 and younger workers by d = -0.56.
Syntactic structure showed another measurable pattern. Clausal subordination, which captures one aspect of sentence complexity, declined with age at β = -0.25. The older-versus-younger group difference reached d = -0.68. Global length units and phrasal constructions also declined.
Not every commonly assumed age difference appeared. Lexical diversity was not significantly related to age after correction, nor was idea density. The researchers also found no significant age relationship for response pacing or response initiation latency. In other words, older participants were not simply slower to begin or complete their answers in this setting.
That distinction matters because an automated assessment can potentially treat linguistic characteristics as signals about an applicant. A difference in sentence structure, topic framing or vocabulary is not automatically evidence of a difference in job capability. The validity question is whether a feature is genuinely relevant to the construct an employer intends to measure.
Could better questions narrow the difference?
The second study tested whether the wording of interview questions itself could be an upstream intervention. Rather than immediately recruiting another human sample, the researchers created an AI persona for each of the 368 participants using gpt-4o-mini.
Each persona was conditioned on its corresponding participant’s age, gender, race, education, socioeconomic status, personality scores, vocational interests and original interview transcript. The personas then answered both the original interview questions and revised versions designed to reduce age-linked linguistic differences.
The simulation reproduced relative differences imperfectly but meaningfully. Across linguistic features, the average correlation between human and simulated scores was 0.62. However, the average standardised mean difference between human and simulated scores was 0.52, showing that the AI responses did not reproduce the absolute human levels closely across every feature.
The researchers therefore treated the personas as a prototyping tool rather than as substitutes for human validation.
The revised questions were deliberately more structured. One question that originally asked participants to describe several work-related activities was narrowed to one activity. Other revisions added tighter time anchors, encouraged direct wording, emphasised job-relevant experiences and standardised the number of examples requested.
Across the features collectively, the interaction between prompt condition and age was statistically significant, Pillai’s Trace = 0.17, F(39, 291) = 1.52, p = 0.030, partial eta-squared = 0.17. That indicates the relationship between age and language changed when the prompts changed.
Some gaps became much smaller, but not all
The largest attenuation occurred for past-focused language. Its age coefficient fell from β = 0.20 under the original prompts to β = 0.05 under the revised prompts, a 77% reduction.
Home-related language showed a 73% attenuation, big-word use 67%, phrasal constructions 62% and achievement-related language 51%. Smaller reductions appeared for global length units at 43%, money-related language at 34%, work-related language at 32%, curiosity language at 29%, topic relevance at 23% and clausal subordination at 19%.
Several effects that had been statistically significant with the original prompts were no longer significant after modification, including past focus, global length units, phrasal constructions and achievement-related language.
But prompt engineering did not provide a universal fix. The age effect for family-related language remained at β = 0.16 in both conditions. More strikingly, the age relationship for cognitive-process language became stronger, shifting from β = -0.04 under the original prompts to β = -0.15 under the revised prompts. The learning topic also changed direction rather than simply shrinking.
Those mixed results are one of the study’s most useful findings. Reducing one demographic difference in an assessment does not guarantee that another will disappear. A prompt revision can alter several dimensions of a response simultaneously, sometimes in unintended directions.
What this means for AI-assisted hiring
For organisations using automated video interviews, the research shifts part of the fairness discussion upstream. Auditing a scoring model after it has been built is important, but the question presented to applicants can already shape the linguistic patterns that the model receives.
That makes prompt design part of assessment design rather than a cosmetic writing decision. Questions that implicitly reward particular ways of discussing curiosity, achievement, relationships or past experience may interact with communication patterns that vary across adulthood.
The evidence does not show that automated video interviews discriminate against older applicants, nor does it demonstrate that the linguistic differences identified here cause different hiring outcomes. It shows something narrower but practically important: workers of different ages responded differently to identical interview prompts, and changing those prompts altered several of the observed age relationships.
For South African employers, where recruitment increasingly combines digital screening with large and demographically diverse applicant pools, that distinction is relevant. Fairness testing should examine not only the algorithm and its final recommendations but also the questions, instructions, time constraints and linguistic features that feed the system.
The study is a proof of concept, not a hiring audit
Several limitations constrain the conclusions. The human participants completed a research task rather than a real high-stakes job application. People competing for employment may monitor their language more carefully, manage impressions more actively or respond differently under time and evaluative pressure.
The sample was also drawn from an online research platform and consisted of first-language English speakers. Cultural, occupational and industry-specific communication norms could produce different age patterns elsewhere.
Most importantly, the revised prompts were tested on AI personas rather than a new human sample. The personas preserved relative patterns reasonably well, but their absolute linguistic scores differed from the human data. LLMs can also reproduce stereotypes embedded in training data and can respond to instructions differently from people.
The 77% attenuation therefore should not be interpreted as evidence that the same prompt revision would reduce an age difference by 77% among real applicants. It is a simulation result that identifies promising designs for subsequent human testing.
That is also where the study’s practical value lies. Generative AI may offer assessment developers a relatively inexpensive way to stress-test alternative questions before exposing real candidates to them. The final standard, however, remains empirical validation with people and evidence that the assessment measures job-relevant characteristics fairly across groups.
Source Information
Study Title: Age-inclusive AI assessments: testing and refining open-ended prompts with human and LLM-simulated data
Authors: Jingyi Li, Daphne X. Hou and Louis Tay
Journal: Journal of Business and Psychology
Published: 25 September 2026
DOI: 10.1007/s10869-026-10152-w
Study design: Two-study investigation combining automated video interview data from working adults with LLM persona simulations for prompt redesign
Human sample: 373 working adults recruited, with 368 retained after quality screening, aged 18 to 66 years and older
Main finding: Age was associated with systematic differences in interview language after demographic and individual-difference controls. In AI-persona simulations, redesigned prompts reduced several age-linked linguistic effects, including a 77% attenuation in past-focused language, but some differences persisted or increased.








