{"id":"8d23f123-615e-4316-963d-6eaaa1f6a53c","arxiv_id":"2504.13284","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A sentiment analysis of 10,184 social media comments shows general Senegalese user dissatisfaction with mobile Internet price-to-quality ratios, with the strongest hostility aimed at operator Orange.","lead":"The paper analyzes more than 10,000 Facebook and Twitter comments about mobile Internet prices in Senegal and reports that users, especially subscribers of the leading operator Orange, express widespread dissatisfaction with the price-to-quality balance. It applies standard sentiment analysis tools to a low-resource language setting (Wolof) and offers a snapshot of public opinion that regulators and telecom operators could use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Operator-level sentiment comparisons are temporally confounded: Orange comments are mostly from 2019 while competitors' are from 2021–2024, so the claimed Orange hostility may reflect sampled price events, not stable perception.","rationale":"The reader's weakest assumption was that the scraped comments are representative of young Senegalese users' general perception, with unvalidated XLM-T labels after Google Translate. My concern is a more specific version of the representativeness problem: the sampling is not only demographically unfiltered but temporally and event-wise imbalanced across operators. That imbalance directly threatens the central comparison in Figs. 11–14, because the Orange sample is concentrated in 2019 while competitors are concentrated in later years. The paper acknowledges this in Section 4 but does not control for it in the analysis. This is not an internal inconsistency, but it is a load-bearing validity threat. The classifier-validation issue raised by the reader is also real, but the temporal/event confound is more decisive for the paper's headline comparison across operators. The concern is fixable by restricting to a common time window and comparable post types, or by collecting a balanced sample, so the existing CONDITIONAL verdict remains appropriate. I would not change the reader's verdict, hence UNCHANGED.","tokens_in":9866,"tokens_out":3143,"duration_ms":31120,"concrete_test":"Recompute the sentiment distributions in Figs. 11–14 using only Facebook comments from a common time window (e.g., 2023–2024) for all four operators, and only comments attached to posts of the same type (e.g., package-price announcements). If Orange no longer shows the most negative distribution, or if the overall negative share drops materially, the central claim of general dissatisfaction and sharpest hostility toward Orange is an artifact of sampling different price events in different years.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4, Fig. 5, reports that the majority of Orange comments date from 2019, while Free comments are from 2021–2024, Expresso from 2023–2024, and Promobile from 2024. Section 5 then presents Figs. 11–14 as if these distributions are directly comparable, and the paper concludes that users are generally dissatisfied and that hostility is sharpest toward Orange. The paper itself warns that 'the period is therefore very important to consider, as external circumstances can influence user sentiment towards a particular operator.' This is the load-bearing weakness: operator identity is confounded with year and with the specific price-change post being commented on. The Orange sample is dominated by reactions to one 2019 price event, whereas competitors' samples come from different years and different promotional or price events. Negative comments about a price increase in 2019 would likely appear regardless of operator, so the contrast between Orange and its competitors in Figs. 11–14 may be an artifact of what was scraped, not a stable property of Senegalese users' perceptions. The claim about 'young people' is also unsupported by any demographic filter, but the temporal/event confounding is the more decisive threat to the central cross-operator comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a sentiment-analysis study of Senegalese social-media comments about mobile Internet prices and quality. The authors scraped roughly 10,000 Facebook and Twitter comments related to posts by four operators (Orange, Free, Expresso, Promobile), used GlotLID for language identification, translated Wolof text into French with the Google Translate API, classified sentiment with XLM-T, and produced word clouds and sentiment distributions by operator. They conclude that users are generally dissatisfied with the quality/price trade-off and that the strongest hostility is directed at Orange, whose comments contain terms such as 'voleur' and 'arnaque'. The paper includes a short limitations section and states that a future open Wolof sentiment dataset is planned.","tokens_in":10156,"tokens_out":4174,"duration_ms":39438,"significance":"If the central findings were robust, the paper would provide a useful, policy-relevant measurement of consumer sentiment about mobile Internet in Senegal, a low-resource language setting that is underrepresented in NLP research. The authors are transparent about several methodological choices and limitations, and the qualitative word-cloud evidence gives the claims some plausibility. However, the quantitative sentiment distributions and the cross-operator comparison currently rest on unvalidated labels and on a temporally confounded sample, and the paper's 'young people' framing is not supported by any demographic measurement. These issues limit the paper's contribution to an exploratory case study rather than a validated empirical claim; nonetheless, the gaps are addressable with additional validation and re-analysis, so major revision is appropriate.","major_comments":[{"comment":"The sentiment classifier is never validated on the actual corpus. Section 5 reports that GPT-4o overclassified Wolof texts as neutral and that XLM-T was then used after Google Translate, but no accuracy, agreement, or manual spot-check is reported for XLM-T on this dataset. Since the central quantitative claims come from these labels, the distributions in Figures 11–14 could be artifacts of translation or model bias. I request a stratified human-annotated sample (e.g., a few hundred comments across operators and languages) with reported inter-annotator agreement and per-language accuracy, or an explicit demonstration that the main conclusions are robust to annotation noise.","section":"§5, Figs. 11–14"},{"comment":"The cross-operator comparison is temporally confounded. Figure 5 shows that most Orange comments date from 2019, while Free comments are from 2021–2024, Expresso from 2023–2024, and Promobile from 2024; the sentiment distributions in Figures 11–14 are then compared directly, and the conclusion that hostility is sharpest toward Orange is drawn. Because the comments were scraped from operator posts about price changes, the operator variable is confounded with year and with the specific price event that generated the comments. The paper itself warns in §4 that 'the period is therefore very important to consider,' but no period-adjusted analysis is presented. I request a robustness check restricted to a common time window, or an event-level analysis that accounts for the type of post and the year.","section":"§4, Fig. 5; §5, Figs. 11–14"},{"comment":"The paper's claim to measure 'young people's' perception is unsupported by the data. The collection procedure in §3 targets operator posts and their comments without any demographic filter, user-age variable, or proxy for age, so the commenters cannot be shown to be young. The median-age statistic in §1 is population-level and does not establish the age distribution of commenters. The conclusions should be reframed to 'social media users commenting on operator pages' unless age-relevant evidence is added.","section":"Title, Abstract, §1, §3"},{"comment":"The sampling is event- and platform-specific in a way that directly affects the 'general dissatisfaction' claim. The 250 tweets are described as comments on a single Orange post about lowering package prices collected in one week in September 2024, while the Facebook data dominate the corpus (Fig. 3); the resulting sentiment may reflect reactions to a particular promotional or price event rather than stable user perceptions. I request that results be reported separately by platform and by event, and that the pooled 'general dissatisfaction' claim be justified in light of this selection bias.","section":"§3, §4"}],"minor_comments":[{"comment":"Figures 4 and 5 have identical captions ('Data distribution across Senegal's 04 leading telecom operators'); Figure 5 actually shows the temporal distribution, so its caption should state the time range.","section":"§4, Figs. 4 and 5"},{"comment":"The introduction says there are five operators active in Senegal, but the data collection covers four; please clarify which operator is omitted and why.","section":"§1 and §4"},{"comment":"The text says 'sentient extraction' where 'sentiment extraction' is meant; also, please report the exact XLM-T checkpoint and any decision thresholds used for positive/negative/neutral classification.","section":"§5"},{"comment":"The number '10.000' should be formatted as '10,000', and the limitations section could more explicitly connect the Facebook over-representation to the temporal imbalance shown in Fig. 5.","section":"§3 and §6"},{"comment":"No mention is made of data or code availability; for reproducibility, the authors should state whether the corpus or scraping scripts will be released.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a modest observational NLP study with clearly acknowledged limitations; its main weakness is inferential validity rather than novelty. I do not see any issue with the citation pattern or any sign of circular reasoning, but the central cross-operator claim needs substantially more validation before the paper can be accepted. A workshop venue or a revised journal version with human annotation and period-adjusted analysis would be more appropriate than acceptance in the current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the paper actually does what it says: it collected about 10,000 Facebook and Twitter comments about four Senegalese telecom operators, mostly in Wolof, and ran a translate-then-classify pipeline with XLM-T. That corpus is new, and the word clouds give a vivid sense of how users talk about Orange (“voleur,” “arnaque”). Second, the central comparison—Orange as the most negatively viewed operator—is not supported by the data as analyzed. Orange comments are mostly from 2019; Free, Expresso, and Promobile comments come from 2021–2024. The paper acknowledges this temporal imbalance in Section 4 but then ignores it in Section 5. The negativity directed at Orange could easily be a reaction to a specific 2019 price event, not a stable perception of the operator.\n\nCredit where it is due. The paper is honest about limitations: it flags translation error propagation, Facebook over-representation, and the difficulty of Wolof preprocessing. The use of GlotLID for language identification is sensible, and the authors do not overclaim—the conclusion is a cautious statement about general dissatisfaction.\n\nThe soft spots are real but not fatal to the entire paper. The load-bearing problem is the temporal and event confounding, which the stress-test note correctly identifies. There is also no validation of XLM-T on this specific corpus; we do not know whether the translated labels are accurate for Wolof-heavy comments. The “young people” framing is unsupported—no age or demographic filtering is described. Finally, the dataset and code are not released, which limits independent checking. These are all fixable in a revision: add a human-annotated evaluation set, filter or justify the demographic framing, and release the corpus.\n\nThis is a modest but genuine applied case study. It is most useful for readers working on African-language sentiment analysis, low-resource NLP, or telecom regulation in West Africa. It deserves a serious referee because the empirical result is new and the authors are clearly thinking carefully about their context. I would send it to peer review with a conditional-accept expectation, asking for the confounding analysis to be addressed before publication.","headline":"A genuinely new Senegalese telecom sentiment corpus, but the headline Orange-vs-others finding is confounded by sampling period and the 'young people' framing is unsupported.","tokens_in":10653,"tokens_out":1894,"would_cite":false,"duration_ms":18710,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Senegal's youth call Orange a thief over mobile data prices","keywords":["Sentiment analysis","Social media","Low-resource languages","Wolof","Mobile Internet","Senegal","Telecommunications","Language models"],"falsifier":"Take a random sample of, say, 300 comments from the corpus, have two Wolof- and French-speaking annotators label sentiment by hand, and compare their labels to the XLM-T output; if agreement on Wolof comments is not substantially above chance, the sentiment distributions in Figures 11 through 14 would not survive.","tokens_in":9638,"feed_emoji":"😤","tokens_out":5316,"duration_ms":44099,"temperature":0.7,"pith_summary":"The paper aims to show how young Senegalese users feel about mobile Internet prices relative to network quality, using their own social media comments. It builds a corpus of 10,184 Facebook and Twitter comments about the four main operators and finds a general mood of dissatisfaction, with the harshest language aimed at Orange, the market leader. If the analysis is right, consumer sentiment in Senegal systematically favors cheaper, less reliable competitors over the dominant incumbent, who is nonetheless hard to abandon because rivals offer worse coverage. The authors also argue that such social media commentary can serve as public pressure on telecom operators and support calls for better regulation.","feed_headline":"Senegal's youth call Orange a thief over mobile data prices","feed_subtitle":"A 10,000-comment read finds anger at data prices and praise for cheaper rivals despite weak networks.","key_machinery":"The carrying mechanism is a translated sentiment-analysis pipeline over a scraped social media corpus: 10,184 comments collected from Facebook and Twitter posts where operators announced Internet packages; GlotLID for language identification; Google Translate to convert Wolof text to French; and XLM-T, a multilingual Twitter-pretrained model based on XLM-RoBERTa, to label each comment positive, negative, or neutral. Word clouds are used as a first-pass qualitative check, and the sentiment distributions by operator are then read against the word clouds to connect recurrent vocabulary (e.g., 'voleur' for Orange) to measured polarity.","core_discovery":"The central discovery is that social media commentary by Senegalese users expresses widespread dissatisfaction with the trade-off between mobile Internet cost and network quality, and that hostility is concentrated on Orange, the incumbent operator. Orange comments contain words like 'voleur' (thief), 'arnaque' (scam), and 'boycotte', while Free is framed as the fighter against high prices and Expresso and Promobile are criticized mainly for poor coverage. Despite this hostility, the authors find users do not switch in large numbers because competitors' networks are worse, which they offer as an explanation for Orange's continued dominance. They also report that a substantial share of the corpus is in Wolof, often mixed with French, and that applying a multilingual sentiment model after machine translation yields strongly negative sentiment distributions for Orange and mixed but more positive ones for the challengers.","pith_inferences":["My inference: the temporal imbalance in the data—Orange comments mostly from 2019 and Free comments from 2021 to 2024—could confound the cross-operator comparison, since sentiment about an operator may reflect the specific price changes and network conditions of the year in which comments were posted.","My inference: because the classifier was never validated on the actual corpus, the reported sentiment distributions may systematically underestimate negative sentiment in Wolof, which the authors note tends to be classified as neutral before translation.","My inference: a direct test would be to hand-label a random sample of 200 to 300 comments in Wolof and French, then compare the XLM-T labels; if agreement is low for Wolof, the main quantitative claim would need re-estimation.","My inference: the paper's framing suggests a testable extension—linking sentiment toward each operator to actual subscriber churn data from Senegal's telecom regulator would show whether hostile sentiment predicts switching or is, as the authors suggest, contained by network-quality differences."],"forward_implications":["If the sentiment distributions are accurate, Orange's market dominance coexists with a reputation for high prices and perceived unfair consumption of data packages.","Free and Expresso are seen as cheaper alternatives, but their networks are perceived as weaker, so price alone does not drive switching.","The pipeline's heavy reliance on Facebook comments (over 10,000 of the 10,184 posts) means the findings speak mainly to Facebook users' behavior, not to the whole young population.","The authors' proposed next step—manual annotation of the corpus into an open Wolof sentiment dataset—would be needed before the negative-sentiment ratios can be treated as stable measurements."],"supporting_citations":[{"why":"Supplies the XLM-T multilingual sentiment classifier that produces the positive, negative, and neutral distributions in Figures 11 through 14.","marker":"[22]"},{"why":"XLM-T is built on XLM-RoBERTa, so this reference supplies the base multilingual model behind the classifier.","marker":"[23]"},{"why":"GlotLID is used to identify French versus Wolof comments, determining which texts get translated before sentiment labeling.","marker":"[17]"},{"why":"Supports the translate-test principle the paper relies on for converting Wolof comments into French for the classifier.","marker":"[24]"},{"why":"GPT-4o was the first sentiment model tested; its tendency to over-label Wolof texts as neutral motivated the switch to XLM-T.","marker":"[19]"},{"why":"Documents limitations of large language models on African low-resource languages, supporting the claim that Wolof is poorly handled and justifying the translation step.","marker":"[20]"},{"why":"Provides the AfriSenti benchmark context for sentiment analysis in African languages, grounding the low-resource challenges the paper addresses.","marker":"[2]"},{"why":"Establishes the Twitter-as-corpus approach for sentiment analysis that the paper adapts to collect comments from operator posts.","marker":"[1]"}],"fun_headline_variants":["Senegal youth rage at Orange data prices on social media","Study: Senegal youth see Orange as thief, but stay for network","Social media sentiment: Senegal youth angry at Orange prices","Senegal youth slam Orange data costs despite cheap rivals","Orange faces 'thief' label from Senegal youth in tweets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions rest on the assumption that the scraped comments, drawn mainly from Facebook posts about operator price announcements, fairly represent what young Senegalese users think and that the machine-translated sentiment labels preserve the original tone.","fun_headline_variants_meta":{"raw":{"variants":["Senegal youth rage at Orange data prices on social media","Study: Senegal youth see Orange as thief, but stay for network","Social media sentiment: Senegal youth angry at Orange prices","Senegal youth slam Orange data costs despite cheap rivals","Orange faces 'thief' label from Senegal youth in tweets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001227,"raw_usage":{"total_tokens":4988,"prompt_tokens":833,"completion_tokens":4155,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":4073}},"tokens_in":449,"tokens_out":4155,"duration_ms":28110,"temperature":1.0,"reasoning_tokens":4073,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:11:08.418432+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 300 comments from the corpus, have two Wolof- and French-speaking annotators label sentiment by hand, and compare their labels to the XLM-T output; if agreement on Wolof comments is not substantially above chance, the sentiment distributions in Figures 11 through 14 would not survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the Twitter-as-corpus approach for sentiment analysis that the paper adapts to collect comments from operator posts."}],"review_version":1}