REVIEW 4 major objections 5 minor 24 references
Sentiment Analysis on the young people's perception about the mobile Internet costs in Senegal
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Senegal's youth call Orange a thief over mobile data prices
desk verdict A genuinely new Senegalese telecom sentiment corpus, but the headline Orange-vs-others finding is confounded by sampling period and the 'young people' framing is unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a translated sentiment-analysis pipeline over a scraped social media corpus: 10,184 comments collected from Facebook and Twitter posts where operators announced Internet packages; GlotLID for language identification; Google Translate to convert Wolof text to French; and XLM-T, a multilingual Twitter-pretrained model based on XLM-RoBERTa, to label each comment positive, negative, or neutral. Word clouds are used as a first-pass qualitative check, and the sentiment distributions by operator are then read against the word clouds to connect recurrent vocabulary (e.g., 'voleur' for Orange) to measured polarity.
What would settle it
Take a random sample of, say, 300 comments from the corpus, have two Wolof- and French-speaking annotators label sentiment by hand, and compare their labels to the XLM-T output; if agreement on Wolof comments is not substantially above chance, the sentiment distributions in Figures 11 through 14 would not survive.
Extended reading notes
Core claim
The central discovery is that social media commentary by Senegalese users expresses widespread dissatisfaction with the trade-off between mobile Internet cost and network quality, and that hostility is concentrated on Orange, the incumbent operator. Orange comments contain words like 'voleur' (thief), 'arnaque' (scam), and 'boycotte', while Free is framed as the fighter against high prices and Expresso and Promobile are criticized mainly for poor coverage. Despite this hostility, the authors find users do not switch in large numbers because competitors' networks are worse, which they offer as an explanation for Orange's continued dominance. They also report that a substantial share of the corpus is in Wolof, often mixed with French, and that applying a multilingual sentiment model after machine translation yields strongly negative sentiment distributions for Orange and mixed but more positive ones for the challengers.
Load-bearing premise
The conclusions rest on the assumption that the scraped comments, drawn mainly from Facebook posts about operator price announcements, fairly represent what young Senegalese users think and that the machine-translated sentiment labels preserve the original tone.
Editorial extensions
If this is right
- If the sentiment distributions are accurate, Orange's market dominance coexists with a reputation for high prices and perceived unfair consumption of data packages.
- Free and Expresso are seen as cheaper alternatives, but their networks are perceived as weaker, so price alone does not drive switching.
- The pipeline's heavy reliance on Facebook comments (over 10,000 of the 10,184 posts) means the findings speak mainly to Facebook users' behavior, not to the whole young population.
- The authors' proposed next step—manual annotation of the corpus into an open Wolof sentiment dataset—would be needed before the negative-sentiment ratios can be treated as stable measurements.
Reading between the lines
- My inference: the temporal imbalance in the data—Orange comments mostly from 2019 and Free comments from 2021 to 2024—could confound the cross-operator comparison, since sentiment about an operator may reflect the specific price changes and network conditions of the year in which comments were posted.
- My inference: because the classifier was never validated on the actual corpus, the reported sentiment distributions may systematically underestimate negative sentiment in Wolof, which the authors note tends to be classified as neutral before translation.
- My inference: a direct test would be to hand-label a random sample of 200 to 300 comments in Wolof and French, then compare the XLM-T labels; if agreement is low for Wolof, the main quantitative claim would need re-estimation.
- My inference: the paper's framing suggests a testable extension—linking sentiment toward each operator to actual subscriber churn data from Senegal's telecom regulator would show whether hostile sentiment predicts switching or is, as the authors suggest, contained by network-quality differences.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a sentiment-analysis study of Senegalese social-media comments about mobile Internet prices and quality. The authors scraped roughly 10,000 Facebook and Twitter comments related to posts by four operators (Orange, Free, Expresso, Promobile), used GlotLID for language identification, translated Wolof text into French with the Google Translate API, classified sentiment with XLM-T, and produced word clouds and sentiment distributions by operator. They conclude that users are generally dissatisfied with the quality/price trade-off and that the strongest hostility is directed at Orange, whose comments contain terms such as 'voleur' and 'arnaque'. The paper includes a short limitations section and states that a future open Wolof sentiment dataset is planned.
Significance. If the central findings were robust, the paper would provide a useful, policy-relevant measurement of consumer sentiment about mobile Internet in Senegal, a low-resource language setting that is underrepresented in NLP research. The authors are transparent about several methodological choices and limitations, and the qualitative word-cloud evidence gives the claims some plausibility. However, the quantitative sentiment distributions and the cross-operator comparison currently rest on unvalidated labels and on a temporally confounded sample, and the paper's 'young people' framing is not supported by any demographic measurement. These issues limit the paper's contribution to an exploratory case study rather than a validated empirical claim; nonetheless, the gaps are addressable with additional validation and re-analysis, so major revision is appropriate.
major comments (4)
- [§5, Figs. 11–14] The sentiment classifier is never validated on the actual corpus. Section 5 reports that GPT-4o overclassified Wolof texts as neutral and that XLM-T was then used after Google Translate, but no accuracy, agreement, or manual spot-check is reported for XLM-T on this dataset. Since the central quantitative claims come from these labels, the distributions in Figures 11–14 could be artifacts of translation or model bias. I request a stratified human-annotated sample (e.g., a few hundred comments across operators and languages) with reported inter-annotator agreement and per-language accuracy, or an explicit demonstration that the main conclusions are robust to annotation noise.
- [§4, Fig. 5; §5, Figs. 11–14] The cross-operator comparison is temporally confounded. Figure 5 shows that most Orange comments date from 2019, while Free comments are from 2021–2024, Expresso from 2023–2024, and Promobile from 2024; the sentiment distributions in Figures 11–14 are then compared directly, and the conclusion that hostility is sharpest toward Orange is drawn. Because the comments were scraped from operator posts about price changes, the operator variable is confounded with year and with the specific price event that generated the comments. The paper itself warns in §4 that 'the period is therefore very important to consider,' but no period-adjusted analysis is presented. I request a robustness check restricted to a common time window, or an event-level analysis that accounts for the type of post and the year.
- [Title, Abstract, §1, §3] The paper's claim to measure 'young people's' perception is unsupported by the data. The collection procedure in §3 targets operator posts and their comments without any demographic filter, user-age variable, or proxy for age, so the commenters cannot be shown to be young. The median-age statistic in §1 is population-level and does not establish the age distribution of commenters. The conclusions should be reframed to 'social media users commenting on operator pages' unless age-relevant evidence is added.
- [§3, §4] The sampling is event- and platform-specific in a way that directly affects the 'general dissatisfaction' claim. The 250 tweets are described as comments on a single Orange post about lowering package prices collected in one week in September 2024, while the Facebook data dominate the corpus (Fig. 3); the resulting sentiment may reflect reactions to a particular promotional or price event rather than stable user perceptions. I request that results be reported separately by platform and by event, and that the pooled 'general dissatisfaction' claim be justified in light of this selection bias.
minor comments (5)
- [§4, Figs. 4 and 5] Figures 4 and 5 have identical captions ('Data distribution across Senegal's 04 leading telecom operators'); Figure 5 actually shows the temporal distribution, so its caption should state the time range.
- [§1 and §4] The introduction says there are five operators active in Senegal, but the data collection covers four; please clarify which operator is omitted and why.
- [§5] The text says 'sentient extraction' where 'sentiment extraction' is meant; also, please report the exact XLM-T checkpoint and any decision thresholds used for positive/negative/neutral classification.
- [§3 and §6] The number '10.000' should be formatted as '10,000', and the limitations section could more explicitly connect the Facebook over-representation to the temporal imbalance shown in Fig. 5.
- [General] No mention is made of data or code availability; for reproducibility, the authors should state whether the corpus or scraping scripts will be released.
Circularity Check
No significant circularity: sentiment labels come from externally pretrained models and direct corpus statistics, with self-citations only in preprocessing background.
full rationale
The paper's central results are the sentiment distributions in Figures 11-14. These are produced by applying externally pretrained models (Google Translate and XLM-T) to scraped comments; the paper fits no parameters to the target conclusion, and no equation or fitted value is reused as a prediction. The word-cloud observations are direct frequency statistics over the corpus, and the concluding 'general dissatisfaction' is an interpretation of those observed and classified labels rather than a quantity defined in terms of itself. The two self-citations ([15] Beqi spelling corrector and [21] Wolof machine translation) appear only as background in the preprocessing discussion and do not justify the central claim. The paper's stated limitations (Facebook over-representation, translate-test error propagation, and absence of corpus-specific validation of the classifier) are validity concerns, not circularity. A skeptic's temporal-confounding objection would be an empirical validity threat to the cross-operator comparison, but it does not make the derivation equivalent to its inputs by construction. No specific reduction of the required kind can be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Comments on operator posts about internet packages are representative of young Senegalese users' general perception of mobile internet costs.
- domain assumption XLM-T sentiment labels on Google-Translated French text preserve the sentiment of the original French/Wolof comments.
- domain assumption Word-cloud term frequencies capture meaningful sentiment beyond random variation.
Cite this review
Pith. "Pith review of Sentiment Analysis on the young people's perception about the mobile Internet costs in Senegal." pith.science (2026). https://pith.science/paper/F2TFLAOG
@misc{pith2026250413284,
author = {Pith},
title = {Pith review of: Sentiment Analysis on the young people's perception about the mobile Internet costs in Senegal},
year = {2026},
howpublished = {\url{https://pith.science/paper/F2TFLAOG}},
note = {Machine review of arXiv:2504.13284}
}
read the original abstract
Internet penetration rates in Africa are rising steadily, and mobile Internet is getting an even bigger boost with the availability of smartphones. Young people are increasingly using the Internet, especially social networks, and Senegal is no exception to this revolution. Social networks have become the main means of expression for young people. Despite this evolution in Internet access, there are few operators on the market, which limits the alternatives available in terms of value for money. In this paper, we will look at how young people feel about the price of mobile Internet in Senegal, in relation to the perceived quality of the service, through their comments on social networks. We scanned a set of Twitter and Facebook comments related to the subject and applied a sentiment analysis model to gather their general feelings.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Pak A, Paroubek P (2010) Twitter as a corpus for sentiment analysis and opinion mining. In: Calzolari N, Choukri K, Maegaard B, Mariani J, Odijk J,PiperidisS,RosnerM,TapiasD(eds)ProceedingsoftheSeventhInterna- tional Conference on Language Resources and Evaluation (LREC’10), Euro- pean Language Resources Association (ELRA), Valletta, Malta, URLhttp: //www...
work page 2010
-
[2]
Muhammad S, Abdulmumin I, Ayele A, Ousidhoum N, Adelani D, Yimam S, Ahmad I, Beloucif M, Mohammad S, Ruder S, Hourrane O, Jorge A, Brazdil P, Ali F, David D, Osei S, Shehu-Bello B, Lawan F, Gwadabe T, Ru- tunda S, Belay TD, Messelle W, Balcha H, Chala S, Gebremichael H, Opoku B, Arthur S (2023) AfriSenti: A Twitter sentiment analysis benchmark for African...
2023
-
[3]
Yuri MN, Rosli MM (2022) Telcosentiment: Sentiment analysis on mo- bile telecommunication services. Journal of Applied Research and Multidis- ciplinary Studies 6(3), URLhttps://mail.journalppw.com/index.php/ jpsp/article/view/5113
work page 2022
-
[4]
Amalia Z, Irfan M, Maylawati DS, Wahana A, Zulfikar WB, Ramdhani MA (2022) Sentiment analysis of the use of telecommunication providers on twitter social media using convolutional neural network. In: 2022 IEEE 8th International Conference on Computing, Engineering and Design (ICCED), pp 1–6, https://doi.org/10.1109/ICCED56140.2022.10010357
arXiv 2022
-
[5]
Saleem I, Jamil A, Mehmood MA (2023) Employing sentiment analy- sis to enhance customer relationships for mobile phone operators work- ing in pakistan. Journal of Applied Research and Multidisciplinary Studies 4(1), https://doi.org/10.32350/jarms.41.09, URL https:// journals.umt.edu.pk/index.php/jarms/article/view/4336
-
[6]
Skoularikis K, Savvas IK, Garani G, Kakarontzas G (2021) A scalable framework for customer sentiment analysis in the telecommunication in- dustry. In: 2021 29th Telecommunications Forum (TELFOR), pp 1–4, https://doi.org/10.1109/TELFOR52709.2021.9653423
arXiv 2021
-
[7]
Bisbee J, Munger K (2024) The vibes are off: Did elon musk push academics off twitter? PS: Political Science & Politics p 1–8,https://doi.org/ 10.1017/S1049096524000416
-
[8]
Amara A, Hadj Taieb MA, Ben Aouicha M (2021) Multilingual topic modeling for tracking covid-19 trends based on facebook data analysis. Applied intelligence (Dordrecht, Netherlands) 51(5):3052—3073,https:// doi.org/10.1007/s10489-020-02033-3, URL https://europepmc.org/ articles/PMC7881346 Mobile Internet Sentiment Analysis 17
Show all 24 references
-
[9]
International Journal of Internet Mar- keting and Advertising 11:183, https://doi.org/10.1504/IJIMA.2017
Ayo C, Ezenwoke A, Ibukun A (2017) Competitive analysis of social me- dia data in the banking industry. International Journal of Internet Mar- keting and Advertising 11:183, https://doi.org/10.1504/IJIMA.2017. 10006719
2017 doi
-
[10]
Journal of Physics: Conference Series 1641(1):012,012, https: //doi.org/10.1088/1742-6596/1641/1/012012, URL https://dx.doi
Syahriani, Yana AA, Santoso T (2020) Sentiment analysis of facebook comments on indonesian presidential candidates using the naïve bayes method. Journal of Physics: Conference Series 1641(1):012,012, https: //doi.org/10.1088/1742-6596/1641/1/012012, URL https://dx.doi. org/10....
2020 doi
-
[11]
Sandoval-Almazan R, Valle-Cruz D (2018) Facebook impact and senti- ment analysis on political campaigns. In: Proceedings of the 19th An- nual International Conference on Digital Government Research: Gover- nance in the Data Age, Association for Computing Machinery, New York, N...
2018
-
[12]
Baj-Rogowska A (2017) Sentiment analysis of facebook posts: The uber case.In:2017EighthInternationalConferenceonIntelligentComputingand Information Systems (ICICIS), pp 391–395, https://doi.org/10.1109/ INTELCIS.2017.8260068
2017
-
[13]
Afful-Dadzie E, Nabareseh S, Oplatková ZK, Klímek P (2014) Enterprise competitive analysis and consumer sentiments on social media - insights from telecommunication companies. In: Proceedings of 3rd International Conference on Data Management Technologies and Applications - Vo...
2014
-
[14]
In: Hassanien AE, Darwish A, Abd El-Kader SM, Alboaneen DA (eds) Enabling Machine Learning Applications in Data Science, Springer Singapore, Singapore, pp 369–378
Saif Eldin Mukhtar Heamida I, Samani Abd Elmutalib Ahmed AL (2021) The classification model sentiment analysis of the sudanese dialect used into the internet service in sudan. In: Hassanien AE, Darwish A, Abd El-Kader SM, Alboaneen DA (eds) Enabling Machine Learning Applicatio...
2021
-
[15]
URL https://arxiv.org/abs/2305
Mbaye D, Diallo M (2023) Beqi: Revitalize the senegalese wolof language with a robust spelling corrector. URL https://arxiv.org/abs/2305. 08518, 2305.08518
2023 arXiv
-
[16]
URLhttps://arxiv.org/abs/ 1904.00784, 1904.00784
Sitaram S, Chandu KR, Rallabandi SK, Black AW (2020) A survey of code- switched speech and language processing. URLhttps://arxiv.org/abs/ 1904.00784, 1904.00784
2020 arXiv
-
[17]
Kargaran AH, Imani A, Yvon F, Schuetze H (2023) GlotLID: Language identification for low-resource languages. In: Bouamor H, Pino J, Bali K (eds) Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computational Linguistics, Singapore, pp 6155...
2023 doi
-
[18]
URLhttps: //arxiv.org/abs/cs/0205028, cs/0205028
Loper E, Bird S (2002) Nltk: The natural language toolkit. URLhttps: //arxiv.org/abs/cs/0205028, cs/0205028
2002 arXiv
-
[19]
Mbaye et al
OpenAI, :, Hurst A, Lerer A, Goucher AP, Perelman A, Ramesh A, Clark A, Ostrow A, Welihinda A, Hayes A, Radford A, Madry A, Baker-Whitcomb 18 D. Mbaye et al. A, Beutel A, Borzunov A, Carney A, Chow A, Kirillov A, Nichol A, Paino A, Renzin A, Passos AT, Kirillov A, Christakis A...
2024 arXiv
-
[20]
URLhttps://arxiv.org/abs/2406.03368, 2406.03368
Adelani DI, Ojo J, Azime IA, Zhuang JY, Alabi JO, He X, Ochieng M, Hooker S, Bukula A, Lee ESA, Chukwuneke C, Buzaaba H, Sibanda B, Kalipe G, Mukiibi J, Kabongo S, Yuehgoh F, Setaka M, Ndolela L, Odu N, Mabuya R, Muhammad SH, Osei S, Samb S, Guge TK, Stenetorp P (2024) Irokobe...
2024 arXiv
-
[21]
In: Yang XS, Sherratt RS, Dey N, Joshi A (eds) Proceedings of Eighth International Congress on Information and CommunicationTechnology,SpringerNatureSingapore,Singapore,pp243– 255
Mbaye D, Diallo M, Diop TI (2024) Low-resourced machine translation for senegalese wolof language. In: Yang XS, Sherratt RS, Dey N, Joshi A (eds) Proceedings of Eighth International Congress on Information and CommunicationTechnology,SpringerNatureSingapore,Singapore,pp243– 255
2024
-
[22]
URLhttps: //arxiv.org/abs/2104.12250, 2104.12250
Barbieri F, Anke LE, Camacho-Collados J (2022) Xlm-t: Multilingual lan- guage models in twitter for sentiment analysis and beyond. URLhttps: //arxiv.org/abs/2104.12250, 2104.12250
2022 arXiv
-
[23]
URL https://arxiv.org/abs/ 1911.02116, 1911.02116
Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, Grave E, Ott M, Zettlemoyer L, Stoyanov V (2020) Unsupervised cross- lingual representation learning at scale. URL https://arxiv.org/abs/ 1911.02116, 1911.02116
2020 arXiv
-
[24]
URLhttps://arxiv.org/abs/2305
Chen Y, Shah V, Ritter A (2024) Translation and fusion improves zero-shot cross-lingual information extraction. URLhttps://arxiv.org/abs/2305. 13582, 2305.13582
2024 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.