• Home  
  • Reinforcing food preferences first made later learning less flexible
- People

Reinforcing food preferences first made later learning less flexible

In a preregistered experiment, people found it harder to update reward rules for foods they liked, especially after their initial preferences were reinforced.

Person considering several snack foods at a table

A snack you love might seem an obvious winner in a reward game. But what happens when the rules change and the same snack starts predicting a loss? New peer-reviewed research suggests that people have more difficulty learning associations that contradict existing food preferences, particularly when their first experiences in a task reinforce those preferences.

In a preregistered study published in Communications Psychology on 9 October 2026, researchers examined how people updated associations between personally liked or disliked food images and positive or negative point outcomes. The main analyses included 279 participants across two cohorts. Average prediction accuracy was 68% when rules matched food preferences and 61% when they contradicted them. Yet the more important finding concerned learning flexibility: an initial round of feedback that confirmed existing preferences appeared to make subsequent adaptation more difficult.

The experiment did not test whether participants changed their diets, ate more nutritious foods or maintained new habits. It examined how prior preferences affected learning in a controlled online task. That narrower conclusion is relevant because successful behavioural change often requires people to revise familiar associations rather than merely learn information from scratch.

Why food is a difficult test of flexible learning

Many experiments about reward learning use neutral symbols, such as coloured shapes. People usually enter those tasks without strong feelings about the stimuli. Food is different. A person may love strawberries but dislike olives, while someone else has the opposite preferences. These associations reflect a history of sensory experiences, expectations and memories. When new information conflicts with that history, learning may become more difficult.

The researchers, Alexandra Rich, Sohum Kapadia, Ryan Henry, Ohad J. Dan and Ifat Levy, wanted to separate two questions. First, can someone learn which outcome usually follows a particular food image? Second, can they update that association when the rule reverses? High accuracy when a rule confirms existing preferences does not necessarily indicate flexibility when that rule changes.

How the experiment was designed

Participants were recruited online through Prolific in two cohorts. The first cohort informed exploratory analyses that were preregistered before the second cohort was collected. Across the two cohorts, 298 people completed the task; 279 contributed to the principal learning analyses. Among those analysed, 145 started with a preference-consistent block and 134 with a preference-inconsistent block. These numbers should not be confused with the larger number initially recruited.

Participants first rated 65 snack photographs according to how much they liked each food. Researchers then selected four personalised stimuli for each participant: one strongly liked, one liked, one disliked and one strongly disliked food. This avoided assuming that nutritional content or calorie density automatically determined whether a food was appealing.

In the subsequent learning task, each food image was followed by a positive or negative point outcome. During preference-consistent blocks, liked foods were usually followed by positive points and disliked foods by negative points. During preference-inconsistent blocks, the pattern was reversed. The prevailing association held on 90% of trials, with misleading feedback on the remaining 10%.

Participants completed six blocks, three of each type, with 32 to 36 trials per block. This amounted to approximately 192 to 216 predictions per participant. They were told that the associations could switch during the task but did not know when a reversal would happen. The order of the first block was counterbalanced so that some participants began with a familiar rule and others with a contradictory one.

The main outcome was policy-adherence accuracy, the percentage of predictions consistent with the underlying rule. This differs from matching the actual outcome of every trial, since the experimental design deliberately included occasional misleading outcomes. The distinction is important when interpreting the reported percentages.

Prediction accuracy was seven points higher for familiar rules

Across the analysed sample, mean accuracy was 0.68 in preference-consistent blocks and 0.61 in preference-inconsistent blocks. These correspond to 68% and 61%, a difference of seven percentage points. The standard deviations were 0.15 and 0.16 respectively, showing that performance varied between individuals.

The difference was statistically significant, with p < 0.001. The authors also reported F(1, 554) = 33.78 and a partial eta-squared of 0.057 for the block-type effect. The latter is a statistical effect-size measure within the model, not a percentage change in actual eating behaviour.

Importantly, performance generally remained above the 50% chance level even when rules contradicted personal preferences. Participants were capable of learning unfamiliar associations. The issue was that learning them was harder and less reliable than learning associations that matched what they already liked or disliked.

The starting rule changed the pattern of learning

Participants who began with a preference-consistent block achieved an average accuracy of 66% across the task, compared with 62% among those who began with a preference-inconsistent block. A superficial reading would conclude that starting with a familiar rule made people better learners. The researchers found a more nuanced pattern.

The higher overall accuracy among preference-consistent starters was concentrated in blocks that reinforced existing preferences. Participants who began with contradictory rules showed more even accuracy across both kinds of block. In other words, starting with a familiar association may have improved performance when the rule remained familiar while making subsequent reversal more difficult.

The starting-condition effect was statistically significant, F(1, 554) = 8.17, p = 0.004, with partial eta-squared 0.015. Its magnitude was smaller than the overall block-type effect, but the pattern supports the authors’ interpretation that initial reinforcement of prior preferences can reduce learning flexibility.

Liked foods were especially resistant to change

The researchers separately analysed liked and disliked foods. For liked foods among people who started with preference-consistent rules, accuracy was 78% when the rule matched their preference but only 55% when it contradicted it. Among those who started with inconsistent rules, accuracy was 68% for consistent blocks and 57% for inconsistent blocks. Both differences were statistically significant.

That pattern suggests that positive associations with favourite foods were relatively rigid. Participants could still learn that a liked food predicted negative points, but the association was harder to use accurately. This should not be confused with evidence that participants disliked the foods more or changed their actual eating patterns.

Disliked foods behaved differently. Among preference-consistent starters, accuracy was 69% when disliked foods predicted negative points and 63% when they predicted positive points. Among preference-inconsistent starters, accuracy was 68% when disliked foods predicted positive points and 57% when they predicted negative points. These findings suggest that associations involving disliked foods were more sensitive to the rule encountered first.

The statistical interaction between block type and food rating was significant, F(1, 1108) = 19.51, p < 0.001. The important distinction is not simply that liked foods were easier or harder overall, but that liked and disliked foods responded differently to a change in the underlying rule.

Positive feedback was easier to learn

Across food categories and starting conditions, participants also tended to learn associations with positive point outcomes more accurately than associations with negative ones. The researchers interpreted this as a positivity bias in reward learning. However, the outcomes were only plus five or minus five points, not real snacks, money lost or health consequences. Participants received a bonus based on prediction accuracy rather than accumulating positive points.

The findings therefore cannot establish that encouraging nutrition messages outperform warnings in real life. They suggest a plausible research question: could highlighting the rewarding features of unfamiliar foods help people update their preferences more effectively than emphasising potential harms? Testing that possibility would require actual dietary interventions.

Did food valuations change after the task?

The researchers collected additional ratings of enjoyment, expected satisfaction and willingness to pay after the learning game. These analyses involved 274 participants. Some measures shifted, particularly for initially disliked foods among participants who started with preference-inconsistent rules. That pattern is consistent with the idea that positive associations may sometimes influence subjective valuations.

The changes were not uniform. Among preference-inconsistent starters, expected satisfaction for liked foods decreased slightly after the task, p = 0.030, with within-person effect size d = -0.19. Among preference-consistent starters, expected satisfaction for disliked foods increased slightly, p = 0.009, with d = 0.22. These were modest changes in immediate ratings, not demonstrations of durable food preference reversal.

No one was observed selecting groceries, preparing meals or maintaining a new diet. A shift in a questionnaire response after a short points game is not equivalent to long-term behavioural change.

Why the results matter beyond the experiment

The study offers a useful distinction for education, public health and behaviour-change research: learning a familiar rule and adapting to a contradictory rule are different achievements. A programme may appear successful when people repeat an already comfortable response, yet still fail to prepare them for changed circumstances.

Research Today has previously covered experiments showing that people continued using teaching shortcuts after they stopped working. The contexts differ, but both lines of research illustrate how a previously successful response can persist beyond the circumstances that made it useful.

For future nutrition research, the findings suggest that individual preferences and the sequence of feedback deserve attention. A disliked but nutritious food may be more responsive to positive associations than an already-liked food is to negative ones. That possibility should be tested with meaningful outcomes, not assumed to be established by this task.

Limitations and questions for future research

The experiment used food photographs and artificial point outcomes, which may be less motivating than actual foods, social experiences or money. Its online participants may not represent the wider population. The study did not measure changes in consumption, weight, nutritional quality or long-term adherence.

The researchers also lacked a neutral-stimulus baseline. Consequently, they could not establish whether beginning with familiar rules improved learning relative to a neutral starting point or whether beginning with contradictory rules impaired it. The mechanisms behind the order effect, such as anchoring, confirmation bias or confidence, remain hypotheses rather than directly established causes.

Food preferences were measured using rating categories rather than comprehensive assessments of taste, nutrition beliefs, culture and willingness to pay. The experiment did not record participants’ confidence in each prediction. These design choices limit how precisely the results can be generalised to everyday decisions.

An exploratory analysis found a weak association between binge-eating severity and accuracy in one specific condition, r = -0.197, p = 0.019. The authors cautioned that the effect was small and involved multiple comparisons. It should not be used as evidence that the task diagnoses or treats an eating disorder.

The use of two cohorts and preregistration improves transparency, but replication in larger and more diverse samples remains necessary. Studies involving actual food choices, more consequential rewards and repeated observations over time would be needed before the findings could guide interventions.

The central lesson is flexibility

The clearest result is that existing preferences can shape new learning even when participants know that rules might change. In this study, people predicted outcomes more accurately when associations matched their food preferences, and starting with confirmation of those preferences was linked to less flexible performance after reversals. Disliked foods appeared more sensitive to the initial learning environment than liked foods.

Whether those dynamics can help people make healthier choices remains an open question. For now, the experiment shows why performance under familiar conditions should not be confused with an ability to adapt when circumstances change.

Source Information

Original study: Prior preferences interfere with the associative learning of food values.
Authors: Alexandra Rich, Sohum Kapadia, Ryan Henry, Ohad J. Dan and Ifat Levy.
Journal: Communications Psychology, volume 4, article 132 (2026).
Publication date: 9 October 2026.
Study design: Preregistered online probabilistic reversal-learning experiment using personalised food images, with 279 participants in the main learning analyses across two cohorts.
DOI: 10.1038/s44271-026-00533-5.

Contact Us

Research Today is a South African digital publication that makes credible research easier to understand.

TERMS OF USE & PRIVACY POLICY

follow us