What people choose to show about themselves online may contain signals about how they feel about their bodies, but those signals do not necessarily look the same across cultures. A new peer-reviewed study suggests that combining information from social media selfies and self-descriptions can predict self-reported body satisfaction more effectively than analysing either type of information alone.
The study, published in Scientific Reports on 4 October 2026, examined 240 participants from Western and East Asian cultural contexts. Rather than treating social media photographs as a universal visual language, the researchers tested whether the relationship between online self-presentation and body satisfaction differed depending on cultural context.
Why online self-presentation matters
Body satisfaction refers to how positively or negatively people evaluate their own bodies. It is an important component of psychological well-being and has been studied extensively in relation to social comparison, appearance ideals and social media use. Much of that work, however, relies on questionnaires about media exposure or body image rather than examining the photographs and words people actually use to present themselves online.
That distinction matters because a social media profile is inherently multimodal. A photograph can convey facial expression, posture, framing and setting, while accompanying text can reveal tone, self-description and contextual information. Looking at only one of those channels may therefore miss part of the pattern.
Culture adds another layer. Norms surrounding modesty, individual expression, appearance and self-promotion differ across societies, so the same type of photograph or written description may not carry the same relationship with body satisfaction in different cultural groups.
Researchers combined images, language and self-reported body satisfaction
The researchers recruited 240 participants from Western and East Asian cultural contexts. Participants provided self-reported body satisfaction measures together with selected social media images and textual materials.
The team then analysed the two forms of online self-presentation separately before combining them. Visual features were extracted from the images using convolutional neural networks, a class of machine-learning model widely used to identify patterns in visual data. Linguistic features were derived from participants’ text using multilingual transformer-based language models.
The researchers integrated these streams using a late-fusion strategy. In practical terms, this allowed information derived from images and information derived from language to contribute to a combined prediction rather than forcing the two types of data into a single representation from the beginning.
They evaluated the models in both regression and classification tasks. Regression assessed how well the models could predict variation in body satisfaction scores, while classification tested their ability to distinguish categories of body satisfaction. Explainability analyses were then used to investigate which features contributed to the predictions. The quantitative modelling was supplemented with qualitative examination of a purposively selected subset of cases, including misclassified examples and cases receiving high model attention.
Combining photographs and text improved prediction
The central quantitative result was consistent across the two modelling approaches: models combining visual and linguistic information outperformed models relying on a single modality. In other words, neither selfies nor written self-descriptions captured the full predictive information available in participants’ social media self-presentation.
This is important because the study was not simply asking whether people with different levels of body satisfaction post different photographs. It tested whether several forms of naturally occurring self-presentation could be analysed together and whether that combined information improved prediction of an independently reported psychological measure.
The explainability analysis also showed that the predictive signals were not culturally uniform. Among Western participants, visual characteristics such as facial expression and body posture contributed more strongly to predictions. Among East Asian participants, contextual and linguistic information made a larger contribution.
Those differences suggest that an identical computational approach may not interpret self-presentation equally well across cultural settings. A model trained to place heavy emphasis on overt visual expression, for example, could overlook informative linguistic or contextual signals in populations where self-presentation follows different social conventions.
The qualitative cases helped explain the cultural differences
The researchers did not rely solely on model performance statistics. They also examined selected cases qualitatively to understand why some profiles attracted greater model attention or were classified incorrectly.
These cases pointed to culturally differentiated patterns involving self-expression, modesty and self-objectification. That finding helps explain why a visually similar profile feature may not have an identical psychological meaning across groups. An image that appears highly expressive in one cultural setting may function differently in another, particularly when accompanying language changes the context in which it is interpreted.
The mixed-methods design therefore adds an important qualification to the machine-learning results. Better predictive accuracy does not mean that a model has discovered a universal visual marker of body satisfaction. Instead, the study suggests that body satisfaction is associated with a combination of online behaviours whose relevance varies with cultural context.
What the findings could mean for digital mental-health research
Computational analysis of social media is increasingly used to study psychological well-being because online platforms contain large amounts of behavioural data produced outside conventional laboratory settings. The present findings show why multimodal and culturally sensitive approaches may be preferable to systems that infer psychological characteristics from photographs or language alone.
For researchers, the results indicate that cultural context should be treated as part of the modelling problem rather than as a demographic variable added after a model has been developed. Predictors that perform well in one cultural group may rely on different information in another group even when the underlying outcome, in this case body satisfaction, is measured in the same way.
The findings may also be relevant to the design of digital well-being tools. A system intended to identify patterns associated with body-image concerns could generate misleading conclusions if it assumes that facial expression, posture, language and contextual cues have universal meanings. Culturally differentiated modelling may therefore be important for both accuracy and fairness.
Prediction is not diagnosis
The authors explicitly caution against interpreting the models as clinical or diagnostic tools. The outcome was self-reported body satisfaction, and the study identified predictive associations between that measure and participants’ social media materials. It did not demonstrate that social media behaviour causes body dissatisfaction or that an algorithm can diagnose a mental-health condition from a profile.
The sample of 240 participants is also relatively modest for machine-learning research, particularly once participants are divided across cultural contexts. Larger and more diverse samples would be needed to determine how reliably the observed patterns generalise to different countries, age groups, platforms and forms of social media use.
Participants supplied selected images and textual materials, meaning the analysis reflects what they chose to provide rather than an unrestricted record of their online behaviour. Platform conventions can also shape self-presentation, and patterns observed in one social media environment may not transfer directly to another.
Finally, explainable machine learning can identify features that contribute to a model’s predictions, but feature importance does not establish a psychological mechanism. Facial expression, posture or language may correlate with other characteristics not directly measured in the study.
A more culturally specific picture of digital self-presentation
The study offers a useful demonstration of how social media research can move beyond counting posts or analysing a single type of content. By combining selfies, written self-descriptions, self-reported body satisfaction and qualitative interpretation, the researchers found that online self-presentation contained meaningful predictive information while also showing that the relevant signals differed across cultural groups.
The broader implication is not that a selfie reveals how satisfied someone is with their body. Rather, patterns distributed across images, language and context may be associated with body satisfaction, and understanding those patterns requires attention to the cultural environment in which people present themselves.
Source Information
Study: Cross-cultural modeling of body satisfaction using social media selfies and self-descriptions
Authors: Huang Qiuyang, Wu Hua and Chen Zhengjun
Journal: Scientific Reports
Published: 4 October 2026
DOI: 10.1038/s41598-026-74109-y








