REVIEW 5 major objections 5 minor 39 references
Semiotic Reconstruction of Destination Expectation Constructs An LLM-Driven Computational Paradigm for Social Media Tourism Analytics
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Using GPT-4 to score travel posts, this paper finds that leisure and social expectations drive likes more than natural or emotional ones.
desk verdict The central claim is unsupported: LLM-derived expectation scores are never validated against an external criterion, so the leisure/social regression may just reflect GPT-4's semantic associations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a two-stage LLM pipeline. In the unsupervised stage, GPT-4 is prompted to extract latent expectations from each post, and the outputs are consolidated into five categories (Emotional, Natural, Exotic Cultural, Leisure, Social). In the supervised stage, 1,287 crowdworker ratings on a 7-point scale are used to fine-tune GPT-4 via LoRA, so the model learns to assign intensity scores to each category. The resulting scores for all posts are then related to like counts using multivariate linear regression, a random forest model, partial dependence plots, and SHAP values.
What would settle it
A concrete falsifier is a pre-registered replication that scores the original Chinese texts (no translation step) with bilingual human raters and fails to reproduce the significant positive coefficients for Leisure and Social expectations, or finds that the effect vanishes once post topic (e.g., travel tips versus trip reports) is statistically controlled.
Extended reading notes
Core claim
The paper's central claim is that leisure/social expectations, as scored by GPT-4 on a 1–7 intensity scale, are the expectation dimensions that actually drive engagement on Chinese social media tourism posts. In the linear regression, Leisure Expectation ($b = 7.08$, $p < 0.05$) and Social Expectation ($b = 5.61$, $p < 0.05$) are the only significant positive predictors of like counts, and the SHAP analysis of the random forest model shows these two categories making the most consistently positive contributions, whereas Natural and Emotional Expectations show high feature-importance rankings but unstable and often negative effects. The paper interprets this as evidence that foundational destination attributes have become cognitively entrenched but behaviorally inert, while hedonic and social needs dominate platform-mediated engagement.
Load-bearing premise
The load-bearing premise is that the GPT-4-derived expectation scores, computed on English translations of Chinese posts, faithfully measure the actual expectations of the Chinese-speaking posters and readers; if this measurement is biased, the relationship with likes is an artifact of the instrument.
Editorial extensions
If this is right
- Destination marketing organizations could prioritize posts and content that express leisure and social expectations when aiming to maximize engagement.
- The five-category expectation scoring pipeline offers a scalable alternative to manual content analysis and survey-based instruments for measuring tourist expectations from UGC.
- The engagement hierarchy identified here implies that pandemic-era tourism content on Chinese social media functions more as a space for hedonic and social fulfillment than for scenic or emotional appreciation.
- Because the pipeline is domain-general, the same dual-method LLM approach could be applied to other consumer domains where user-generated content encodes latent expectations.
Reading between the lines
- Editorial inference: the English-translation step in preprocessing could be the source of the result; a replication that scores the original Chinese text would test whether leisure/social dominance survives translation.
- Editorial inference: since the five expectation categories themselves were derived from GPT-4 outputs and then human labels were anchored to those categories, the category set may omit other relevant dimensions (budget, uniqueness, status) that appeared in earlier extraction.
- Editorial inference: the causal direction is not established; posts may attract likes because they use social/leisure language, or users may adopt such language after seeing what performs well.
- Editorial inference: because the data window covers the COVID-19 period, the leisure/social dominance may be specific to pandemic constraints; post-pandemic data could show a return to natural and emotional expectations as the strongest engagement drivers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a two-phase LLM-based pipeline for measuring destination expectations from Chinese social media UGC collected from Weibo and Xiaohongshu. In the unsupervised phase, zero-shot GPT-4 prompts extract expectation categories from posts, and the top five categories are selected by word frequency. In the supervised phase, crowdworkers rate posts on those same categories, and a LoRA fine-tuned GPT-4 model is trained to reproduce the ratings. The resulting expectation-intensity scores are then used as predictors of like counts in linear regression and random forest analyses. The paper's headline claim is that leisure/social expectations drive engagement more than natural/emotional expectations.
Significance. If the measurement validity of LLM-derived expectation scores could be established, the framework would offer a scalable complement to surveys and manual content analysis for tourism research. The paper is transparent about some limitations, and it does provide a Bland-Altman consistency check and chain-of-thought rationales as audit trails. However, in its current form the central claim is not supported: the expectation instrument is validated only against itself, the engagement analysis lacks the statistical detail needed to compare drivers, and the random forest results contradict the regression narrative without a quantitative resolution.
major comments (5)
- [§3.1.3, §3.2.1, §3.3] The construct-validity loop is closed inside GPT-4's own semantic space. GPT-4 generates the expectation categories in §3.2.1, crowdworkers in §3.1.3 rate posts using exactly those categories, and the Bland-Altman check in §3.3 compares the fine-tuned model's scores to those same crowdworker labels. No comparison to an established psychometric instrument or to an independently derived category system is reported. Consequently, the significant coefficients for Leisure and Social expectations in Table 3 may reflect systematic linguistic association biases introduced by GPT-4's translation and category generation rather than real tourist expectations; the paper's central claim therefore depends on an assumption that is never tested.
- [Table 3, §4.3] The linear regression reports only coefficients and p-values. No R², adjusted R², or model F-statistic is given, and no control variables such as post length, platform, posting date, author characteristics, or time effects are included. The raw coefficients (7.08 and 5.61 for Leisure and Social) are small relative to the intercept of 258.14, and the like-count scaling is not precisely described, so the reader cannot assess effect size. Without these, the statement that leisure/social expectations 'drive engagement more' than other expectations is not quantitatively supported.
- [§4.3, Figures 4–6] The random-forest feature importance in Figure 4 assigns the highest contributions to Exotic Cultural and Natural expectations (0.207 each) and the lowest to Social (0.197) and Leisure (0.185), directly contradicting the regression narrative. The authors reinterpret this using SHAP and PDP plots in a purely qualitative way, and no statistical test, effect-size metric, or quantitative comparison of SHAP values is provided. The central claim is therefore left resting on a contradiction between two analyses, with no formal resolution.
- [§4.2, §3.1.3] The sample-size reporting is internally inconsistent and blocks reproducibility. The corpus has 12,843 posts, but §4.2 says 7,135 entries (70% of total samples), which is not 70% of 12,843. The training subset of 100 and validation set of 1,425 do not sum to 7,135 or to any clear fraction of the corpus. Additionally, §3.1.3 reports 1,287 validated questionnaire responses, while §4.2 reports 1,500 distributed evaluators with 15 ratings per sample. These discrepancies need to be reconciled.
- [§6, §3.2.1] The paper explicitly acknowledges prompt sensitivity as a limitation but reports no sensitivity analysis. Because the entire measurement pipeline—category selection and score generation—depends on the GPT-4 zero-shot prompts in §3.2.1, different prompt formulations could plausibly change which categories emerge and how intensities are assigned. Without a stability check across prompts, the robustness of the five-category structure and of the subsequent regression results is unknown.
minor comments (5)
- [§3.1.2] The preprocessing description is internally contradictory: noise removal is said to strip emojis, while the normalization stage just before it converts emojis to semantic descriptors. Clarify which pipeline is actually used.
- [§3.2.1] The claim that GPT-4 has 1.76 trillion parameters is not publicly documented and is irrelevant to the method; remove it or cite a reliable source.
- [References] Several references do not support the statements they are attached to; for example, references 2–4 concern hate-speech detection and magnetic recording rather than the cited tourism/UGC claims. The reference list needs a thorough correction.
- [Table 1, Figure 2] Table 1's columns 'Word Count' and 'Content (%)' are not defined, and Figure 2 lacks axis labels and numeric values; without these, the category-selection step cannot be audited.
- [§4.2] Averaging multiple raters before the Bland-Altman comparison hides inter-rater disagreement; report the distribution of ratings or the inter-rater reliability for the 15 ratings per sample.
Circularity Check
The engagement regression is not independent: expectation labels were collected with likes visible, so the central claim partially reduces to measurement contamination.
-
self definitional
[Section 3.1.3 (Questionnaire Design and Labeling) feeding Section 3.2.2 fine-tuning and Section 4.3 (Relationship Between Tourism Expectations and Likes)]
"Each questionnaire item followed a structured format: 'To what extent does the following post express strong expectation about [ExpectationCategory]?' accompanied by the original social media text and its contextual metadata (post date, engagement metrics). ... Initial exploration through multivariate linear regression modeled like counts as a function of expectation intensity scores."
The 'true' expectation scores used to fine-tune the scoring model were collected from human raters who saw each post's engagement metrics, which include the likes count that Section 4.3 later uses as the dependent variable. The fine-tuned GPT-4 reproduces these questionnaire scores, and those model scores are then regressed on likes. Thus the predictor variable is constructed from labels that already contained information about the outcome; the significant Leisure and Social coefficients may reflect labelers' reactions to popularity rather than an independent expectation-to-engagement effect. The regression is not a clean behavioral validation because the measurement instrument for the predictor was contaminated by the criterion it is claimed to predict.
full rationale
The paper's central claim that leisure/social expectations drive engagement more than natural/emotional ones rests on a regression where the predictor scores were not independently measured. The human questionnaire used to fine-tune the scoring model displayed engagement metrics alongside each post, so the learned expectation scores encode popularity signals. Consequently, the subsequent regression of likes on those scores is partially circular: the independent variable contains information about the dependent variable by construction. No external psychometric validation or out-of-sample behavioral prediction is provided to break this loop. While the unsupervised extraction and the Bland-Altman agreement check are internally consistent, they do not validate the scores against an external criterion. The contradiction between random forest importance and SHAP/PDP interpretations further weakens the empirical claim, but the primary circularity is the label contamination.
Assumptions & free parameters
free parameters (6)
- Top-k expectation selection =
k=5
- Like-count scaling factor =
100
- Minimum post length =
50 characters
- Duplicate detection threshold =
Jaccard similarity >0.85
- Repeated model evaluations =
10
- Crowdsourced ratings per sample =
15
assumptions (5)
- domain assumption GPT-4 can recover latent tourist expectations directly from preprocessed Chinese social media text.
- ad hoc to paper The five expectation categories map cleanly onto tourist psychology and are mutually exclusive.
- domain assumption Translation of Chinese UGC into English preserves expectation-relevant semantics.
- domain assumption Crowdworkers' self-reported expectation intensity is ground truth for latent expectations.
- domain assumption The number of likes is a valid proxy for engagement and behavioral impact.
invented entities (4)
-
Expectation saturation threshold theory
-
Compensatory gratification hypothesis
-
Digital hermeneutic gap
-
Computational tourism phenomenology
Cite this review
Pith. "Pith review of Semiotic Reconstruction of Destination Expectation Constructs An LLM-Driven Computational Paradigm for Social Media Tourism Analytics." pith.science (2026). https://pith.science/paper/UJZDO2BU
@misc{pith2026250516118,
author = {Pith},
title = {Pith review of: Semiotic Reconstruction of Destination Expectation Constructs An LLM-Driven Computational Paradigm for Social Media Tourism Analytics},
year = {2026},
howpublished = {\url{https://pith.science/paper/UJZDO2BU}},
note = {Machine review of arXiv:2505.16118}
}
read the original abstract
Social media's rise establishes user-generated content (UGC) as pivotal for travel decisions, yet analytical methods lack scalability. This study introduces a dual-method LLM framework: unsupervised expectation extraction from UGC paired with survey-informed supervised fine-tuning. Findings reveal leisure/social expectations drive engagement more than foundational natural/emotional factors. By establishing LLMs as precision tools for expectation quantification, we advance tourism analytics methodology and propose targeted strategies for experience personalization and social travel promotion. The framework's adaptability extends to consumer behavior research, demonstrating computational social science's transformative potential in marketing optimization.
Reference graph
Works this paper leans on
-
[1]
Khan, M.M.; Siddique, M.; Yasir, M.; Qureshi, M.I.; Khan, N.; Safdar, M.Z.J.S. The significance of digital marketing in shaping ecotourism behaviour through destinationimage.2022,14,7395.https://doi.org/10.3390/su14127395
-
[4]
Gitari,N.D.;Zuping,Z.;Damien,H.;Long,J.J.I.J.o.M.;Engineering,U.Alexicon- based approach for hate speech detection. 2015, 10, 215-230. http://dx.doi.org/10.14257/ijmue.2015.10.4.21
-
[5]
Destructive de-energizingrelationships:Howthrivingbufferstheireffectonperformance.2015, 100,1423
Gerbasi, A.; Porath, C.L.; Parker, A.; Spreitzer, G.; Cross, R.J.J.o.A.P. Destructive de-energizingrelationships:Howthrivingbufferstheireffectonperformance.2015, 100,1423
work page 2015
-
[6]
Feldman,T.; Gibson,G.J.; USENIX,l.t.m.o.;SAGE.Shingledmagnetic recording: Arealdensityincreaserequiresnewdatamanagement.2013,38,22-30
work page 2013
-
[7]
Martinez-Torres, M.d.R.; Toral, S.L.J.T.M. A machine learning approach for the identification of the deceptive reviews in the hospitality sector using unique attributes and sentiment orientation. 2019, 75, 393-403. https://doi.org/10.1016/j.tourman.2019.06.003
-
[8]
Analysis of complaints in primary care using statistical process control
JE, V.R.J.R.d.C.A.O.d.l.S.E.d.C.A. Analysis of complaints in primary care using statistical process control. 2009, 24, 155-161. https://doi.org/10.1016/s1134- 282x(09)71799-3
-
[9]
Mena, R.A.; Ornelas, E.L.; Hernández, S.Z.J.P.E. Análisis cualitativo para la detección de factores que afectan el rendimiento escolar: estudio de caso de la licenciaturaentecnologíasysistemasdeinformación.2018,38
work page 2018
-
[10]
Social media in travel decision makingprocess.2017,7,193-201
Dwityas, N.A.; Briandana, R.J.I.J.o.H.; Science, S. Social media in travel decision makingprocess.2017,7,193-201
work page 2017
Show all 39 references
-
[11]
Social media envy: How experience sharing on social networking sitesdrives millennials’ aspirational tourism consumption
Liu, H.; Wu, L.; Li, X.J.J.o.t.r. Social media envy: How experience sharing on social networking sitesdrives millennials’ aspirational tourism consumption. 2019, 58,355-369.https://doi.org/10.1177/0047287518761615
2019 doi
-
[12]
Leisure Mobility of Chinese Millennials
Shi, J.; Fan, A.; Cai, L.A. Leisure Mobility of Chinese Millennials. Journal of ChinaTourismResearch2020,16,527-546,doi:10.1080/19388160.2019.1687060
2019
-
[13]
A flash of culinary tourism: Understanding the influences of online food photography on people's travel planning process on flickr
Liu, I.; Norman, W.C.; Pennington-Gray, L.J.T.C.; Communication. A flash of culinary tourism: Understanding the influences of online food photography on people's travel planning process on flickr. 2013, 13, 5-18. https://doi.org/10.3727/109830413X13769180530567
2013 doi
-
[14]
Impact of Short Video Marketing on Tourist Destination Perception in the Post-pandemic Era
Chen, H.; Wu, X.; Zhang, Y.J.S. Impact of Short Video Marketing on Tourist Destination Perception in the Post-pandemic Era. 2023, 15, 10220. https://doi.org/10.3390/su151310220
2023 doi
-
[15]
Tourism and the smartphone app: Capabilities, emerging practice and scope in the traveldomain.2014,17,84-101.https://doi.org/10.1080/13683500.2012.718323
Dickinson, J.E.; Ghali, K.; Cherrett, T.; Speed, C.; Davies, N.; Norgate, S.J.C.i.i.t. Tourism and the smartphone app: Capabilities, emerging practice and scope in the traveldomain.2014,17,84-101.https://doi.org/10.1080/13683500.2012.718323
2014
-
[16]
Virtual destination image a new measurement approach
Govers, R.; Go, F.M.; Kumar, K.J.A.o.t.r. Virtual destination image a new measurement approach. 2007, 34, 977-997. https://doi.org/10.1016/j.annals.2007.06.001
2007 doi
-
[17]
Semantic analysis onsocialnetworks:Asurvey.2020,33,e4424.https://doi.org/10.1002/dac.4424
Bayrakdar, S.; Yucedag, I.; Simsek, M.; Dogru, I.A.J.I.J.o.C.S. Semantic analysis onsocialnetworks:Asurvey.2020,33,e4424.https://doi.org/10.1002/dac.4424
2020 doi
-
[18]
Frias-Martinez,V.; Soto,V.; Hohwald,H.;Frias-Martinez,E.Characterizingurban landscapes using geolocated tweets. In Proceedings of the 2012 International conferenceonprivacy, security, riskandtrustand2012internationalconferneceon socialcomputing,2012;pp.239-248.10.1109/SocialCo...
2012 doi
-
[19]
Effects of tourism information quality in social media on destination image formation: The case of SinaWeibo.2017,54,687-702.https://doi.org/10.1016/j.im.2017.02.009
Kim, S.-E.; Lee, K.Y.; Shin, S.I.; Yang, S.-B.J.I.; management. Effects of tourism information quality in social media on destination image formation: The case of SinaWeibo.2017,54,687-702.https://doi.org/10.1016/j.im.2017.02.009
2017 doi
-
[20]
Reportal,D.Globalsocialmediastats—datareportal—globaldigitalinsights.2021
2021
-
[21]
2013, 30, 3-22
Leung,D.;Law,R.;VanHoof,H.;Buhalis,D.J.J.o.t.;marketing,t.Socialmediain tourism and hospitality: A literature review. 2013, 30, 3-22. https://doi.org/10.1080/10548408.2013.750919
2013
-
[22]
Big Data y administraciones públicas en redessociales.2018,3,1-29
Criado, J.; Pastor, V.; Villodre, J.J.C.N. Big Data y administraciones públicas en redessociales.2018,3,1-29
2018
-
[23]
Acomparative analysis of major onlinereview platforms: Implications for social media analytics in hospitality and tourism
Xiang, Z.; Du,Q.; Ma, Y.; Fan, W. Acomparative analysis of major onlinereview platforms: Implications for social media analytics in hospitality and tourism. Tourism Management 2017, 58, 51-65. https://doi.org/10.1016/j.tourman.2016.10.001
2017 doi
-
[24]
Role of social media in online travel information search.2010,31,179-188.https://doi.org/10.1016/j.tourman.2009.02.016
Xiang, Z.; Gretzel, U.J.T.m. Role of social media in online travel information search.2010,31,179-188.https://doi.org/10.1016/j.tourman.2009.02.016
2010 doi
-
[25]
Exploring the Characteristics and Extent of Travel Influencers’ImpactonGenerationZTouristDecisions.Sustainability2025,17,66
Băltescu, C.A.; Untaru, E.-N. Exploring the Characteristics and Extent of Travel Influencers’ImpactonGenerationZTouristDecisions.Sustainability2025,17,66. https://doi.org/10.3390/su17010066
-
[26]
Modeling individuals’ willingness to share trips with strangers in an autonomous vehicle future
Lavieri, P.S.; Bhat, C.R.J.T.r.p.A.p.; practice. Modeling individuals’ willingness to share trips with strangers in an autonomous vehicle future. 2019, 124, 242-261. https://doi.org/10.1016/j.tra.2019.03.009
2019 doi
-
[27]
Image as a factor in tourism development
Hunt, J.D.J.J.o.t.r. Image as a factor in tourism development. 1975, 13, 1-7. https://doi.org/10.1177/004728757501300301
1975 doi
-
[28]
A general model of traveler destination choice.1989,27,8-14.https://doi.org/10.1177/004728758902700402
Woodside, A.G.; Lysonski, S.J.J.o.t.R. A general model of traveler destination choice.1989,27,8-14.https://doi.org/10.1177/004728758902700402
1989 doi
-
[29]
The Relationships among destination image, visitors' satisfaction and behavioral intentions: The effects of destination image segmentation.2009,9,1-22
Chang, S.J.T.A.o.M.J. The Relationships among destination image, visitors' satisfaction and behavioral intentions: The effects of destination image segmentation.2009,9,1-22
2009
-
[30]
Cyber hate speech on twitter: An application of machine classification and statistical modeling for policy and decisionmaking.2015,7,223-242.https://doi.org/10.1002/poi3.85
Burnap, P.; Williams, M.L.J.P.; internet. Cyber hate speech on twitter: An application of machine classification and statistical modeling for policy and decisionmaking.2015,7,223-242.https://doi.org/10.1002/poi3.85
2015 doi
-
[31]
Large-scale machine learning at twitter
Lin, J.; Kolcz, A. Large-scale machine learning at twitter. In Proceedings of the Proceedingsof the 2012ACM SIGMODInternational Conference onManagement ofData,2012;pp.793-804.https://doi.org/10.1145/2213836.2213958
2012
-
[32]
Predicting hotel review helpfulness: The impact of review visibility, and interaction between hotel stars and review ratings
Hu, Y.-H.; Chen, K. Predicting hotel review helpfulness: The impact of review visibility, and interaction between hotel stars and review ratings. International Journal of Information Management 2016, 36, 929-944. https://doi.org/10.1016/j.ijinfomgt.2016.06.003
2016 doi
- [33]
-
[34]
The power of generative AI: A review of requirements, models, input–output formats, evaluation metrics, and challenges.FutureInternet2023,15,260.https://doi.org/10.3390/fi15080260
Bandi, A.; Adapa, P.V.S.R.; Kuchi, Y.E.V.P.K. The power of generative AI: A review of requirements, models, input–output formats, evaluation metrics, and challenges.FutureInternet2023,15,260.https://doi.org/10.3390/fi15080260
-
[35]
AI-Augmented Surveys: Leveraging Large Language ModelsforOpinionPredictioninNationallyRepresentativeSurveys.2023
Kim, J.; Lee, B.J.a.p.a. AI-Augmented Surveys: Leveraging Large Language ModelsforOpinionPredictioninNationallyRepresentativeSurveys.2023
2023
-
[36]
An Aspect-Based Review Analysis Using ChatGPT for the Exploration of Hotel Service Failures
Jeong, N.; Lee, J. An Aspect-Based Review Analysis Using ChatGPT for the Exploration of Hotel Service Failures. Sustainability 2024, 16, 1640. https://doi.org/10.3390/su16041640
2024 doi
-
[37]
Algorithms 2025, 18, 46
Muhammad,I.; Rospocher,M.OnAssessingthePerformance ofLLMsforTarget- Level Sentiment Analysis in Financial News Headlines. Algorithms 2025, 18, 46. https://doi.org/10.3390/a18010046
2025 doi
-
[38]
Roumeliotis, K.I.; Tselikas, N.D.; Nasiopoulos, D.K. Fake News Detection and Classification: A Comparative Study of Convolutional Neural Networks, Large Language Models,andNatural Language ProcessingModels.Future Internet 2025, 17,28.10.3390/fi17010028
2025 doi
-
[39]
Leveraging ChatGPT and other generative artificial intelligence (AI)-based applications in the hospitality and tourismindustry: Practices, challengesand researchagenda.Int
Dwivedi, Y.K.; Pandey, N.; Currie, W.; Micu, A. Leveraging ChatGPT and other generative artificial intelligence (AI)-based applications in the hospitality and tourismindustry: Practices, challengesand researchagenda.Int. J.Contemp.Hosp. Manag.2023,36,1–12.https://doi.org/10.11...
2023 doi
-
[40]
Xiaohongshu
Lin, B.; Shen, B.J.B.S. Study of consumers’ purchase intentions on community E- commerce platform with the SOR model: a case study of China’s “Xiaohongshu” app.2023,13,103.https://doi.org/10.3390/bs13020103
2023 doi
-
[41]
2021, 231, 107438
Ye,W.;Liu,Z.;Pan,L.J.K.-B.S.Whoarethecelebrities?Identifyingvitaluserson Sina Weibo microblogging network. 2021, 231, 107438. https://doi.org/10.1016/j.knosys.2021.107438
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.