{"id":"77bd6f2e-5181-4a19-9599-c21dd47e6b84","arxiv_id":"2412.15259","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GLARE releases 76 million Google Play reviews, 69 million in Arabic, from 9,980 Saudi Android apps, with descriptive statistics and engineered features.","lead":"This paper introduces GLARE, a dataset of 76 million Google Play Store reviews, 69 million of them in Arabic, collected from the Saudi app store. It is one of the largest Arabic review corpora and is meant to support Arabic NLP tasks and software engineering research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 3 contradicts the 69M Arabic subset count: feature statistics are computed on 76M non-empty reviews, so the headline 'largest Arabic reviews' claim is internally unsupported.","rationale":"I approached the paper as a resource-description paper: the central claim is the existence and scale of GLARE, specifically 69M Arabic reviews. I checked whether the paper's own evidence supports that number. The undocumented language filter flagged by the reader is real, but there is a sharper issue: Table 3's totals (76.5M total, 76.4M non-empty) and the 40.5% one-word percentage are internally consistent only with a denominator of ≈76.4M, not 69M. This means §5's feature statistics were likely computed before or without the Arabic-only filter, so the paper does not actually demonstrate that the analyzed GLARE corpus contains 69M Arabic reviews. I would not reject the paper: the dataset may be exactly as claimed, and the issue is addressable by releasing the filter code and recomputing Table 3. However, the existing CONDITIONAL verdict is appropriate; I therefore set verdict_should_be to UNCHANGED. My agreement is partial because the reader correctly identified the language-detection gap but did not catch the concrete numerical contradiction.","tokens_in":6375,"tokens_out":5711,"duration_ms":51456,"concrete_test":"Download the released GLARE files from GitHub/HuggingFace; inspect the repository for the language-identification code used in §3.1. Then, with an independent detector (e.g., langid or a fastText Arabic/Arabic-script filter), count records in the 'engineered reviews' file and compute the number of Arabic reviews and the one-word percentage. If the Arabic subset is ≈69M and one-word percentage ≈40.5%, the table contradiction is resolved; if the full file is ≈76.4M and the one-word percentage remains 40.5%, the 69M claim is unsupported and Table 3 was computed before the Arabic filter.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 states that after preprocessing, including 'keeping only Arabic reviews,' the dataset contains 'over 69M Arabic app reviews.' The paper's own Table 3 ('Raw and Engineered Features Statistics'), however, reports Total Number of Reviews = 76,512,077 and Total Non-empty Reviews = 76,387,928. The reported one-word statistics confirm the denominator: 30,953,303 one-word reviews / 76,387,928 ≈ 40.5%, exactly the percentage listed; using the 69M Arabic subset would give ≈44.8%. Thus the feature engineering and vocabulary analysis in §5 were performed on the unfiltered 76M corpus, not on the 69M Arabic subset claimed in the abstract and conclusion. Either the 69M figure is a separate, undocumented filter that was never used in the analyses, or the dataset is not actually Arabic-filtered. In addition, §3.1 gives no method, tool, or threshold for detecting Arabic, so the count cannot be reproduced even if the files are downloaded. Since the paper's central contribution is precisely the scale of the Arabic review subset, this internal inconsistency is load-bearing: the headline number is not corroborated by the paper's own statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GLARE, a corpus of Google Play app reviews collected from the Saudi Google Play Store in March-April 2021. The authors scraped reviews for 9,980 unique apps across 59 categories, producing about 76M review records and 17 GB of raw data. They claim that after cleaning and 'keeping only Arabic reviews' the corpus contains over 69M Arabic reviews, and they present exploratory statistics (ratings, thumbs-up counts, developer replies), feature engineering (vocabulary, review length, duplicate apps, category combinations), and potential use cases for Arabic NLP. The dataset is released via GitHub and Hugging Face.","tokens_in":6573,"tokens_out":4626,"duration_ms":38180,"significance":"If the 69M Arabic subset is real and reproducible, GLARE would be a substantial new resource for Arabic NLP, particularly for app-review sentiment analysis and software-engineering tasks, and it would be larger than the Arabic review datasets cited in the related work. The paper's strengths are the clear collection outline, the useful metadata fields, and the public release of raw and engineered files. The descriptive statistics in Sections 4 and 5 are internally consistent with the raw 76M corpus. However, the central 'largest Arabic reviews dataset' claim is not yet established: the paper does not document how Arabic reviews were identified, and the feature-engineering table is based on the unfiltered corpus rather than the claimed 69M subset. These are load-bearing but correctable issues.","major_comments":[{"comment":"The 69M Arabic-review count is contradicted by the paper's own statistics. Table 3 reports Total Number of Reviews = 76,512,077 and Total Non-empty Reviews = 76,387,928, and the one-word percentage (40.5%) is computed from 30,953,303 / 76,387,928 = 40.52%. Using the claimed 69M Arabic subset would give about 44.8%. Section 5.1 explicitly states that the vocabulary is 'extracted from all the reviews in GLARE dataset', and Section 5.2 describes tokenizing 'the reviews' without limiting to the Arabic subset. Therefore the EDA and feature engineering were performed on the full 76M corpus, and the abstract/conclusion claim that 69M Arabic reviews were isolated is not corroborated by any table or analysis. The authors should either release and analyze the Arabic-filtered subset separately and report its statistics, or revise the central claim to describe a raw 76M corpus with an Arabic subset whose size is verified by a documented filter.","section":"§3.1 and §5 / Table 3"},{"comment":"The Arabic language filter is not described. 'Keeping only Arabic reviews' could mean script-based filtering, a language identification model, a lexicon, or manual rules, and each choice materially changes the 69M count and the vocabulary statistics. The paper should name the tool/library/model (with version and language code), the threshold or confidence score, and how mixed-script, transliterated, and dialectal Arabic reviews were handled. Without this, the headline number is not reproducible.","section":"§3.1"},{"comment":"The 'largest Arabic reviews dataset to date' claim is not benchmarked against the full range of existing Arabic review corpora. The related-work section lists several small review datasets, but it does not survey the Masader catalogue or other large review resources to justify the superlative. Because the comparison corpus is not enumerated and the size of the Arabic subset is itself in question, this claim needs a concrete comparison table or an explicit scope, such as 'largest Arabic app-review dataset'.","section":"§1 and §2"}],"minor_comments":[{"comment":"The license is described as 'Creative Commons Attribution-Non-Commercial 4.0 International License (CC BY 4.0)', but CC BY 4.0 is not a non-commercial license; please correct the license name and URL to match the actual license.","section":"§8"},{"comment":"The collection period 'March 21 - April 21' should include the year, e.g., March 21, 2021 – April 21, 2021, along with the exact scraper and library versions used.","section":"§3.1 / Table 2"},{"comment":"The definition of 'noise' is ambiguous: it is first reported as 17% of the vocabulary, then 15% of vocabulary words 'had digits', and later defined as 'any character that is non-Arabic alphanumeric'. Please define these categories as mutually exclusive or explain the overlap.","section":"§5.1"},{"comment":"The top frequent terms are given as glyph-name sequences rather than in Arabic script; rendering the actual Arabic words would improve readability.","section":"§5.1"},{"comment":"The paper says 'over 11K apps' in Section 3.1 and '11,736 apps' in Section 5.3; please harmonize the two numbers.","section":"§3.1 / §5.3"},{"comment":"No scraping or cleaning code is provided in the repository; for a dataset paper, adding the collection and cleanup scripts, or at least a detailed pseudocode, would support reproducibility.","section":"§3.1 / §7"}],"recommendation":"major_revision","confidential_remarks":"The dataset has clear potential value for Arabic NLP, but the internal inconsistency between the 69M Arabic-review headline and the 76M-based statistics in Table 3 must be resolved before the paper can be accepted. If the 69M figure cannot be reproduced with a documented filter, the abstract and conclusion should be revised to present the corpus as a raw 76M review collection with an unverified Arabic subset, or the claim should be withdrawn."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read on GLARE. The useful part is real: a large, publicly released corpus of Google Play reviews in Arabic, about 10k apps and claimed 76M reviews, from a domain that is under-served for Arabic. The scraping procedure is described well enough to follow, and the exploratory statistics give a first look at ratings, developer replies, and vocabulary. If the data is as described, it has genuine value for Arabic sentiment analysis, ABSA, and software engineering research.\n\nThe soft spot is load-bearing. The abstract and conclusion say 69M of the 76M reviews are Arabic, and Section 3.1 says preprocessing included \"keeping only Arabic reviews.\" But Table 3, the paper's own feature-statistics table, reports Total Number of Reviews = 76,512,077 and Total Non-empty Reviews = 76,387,928, and every subsequent statistic is computed on that 76M denominator. The one-word review count is 30,953,303, which is 40.5% of 76,387,928 exactly as reported; against a 69M Arabic subset it would be about 44.8%. So Sections 4 and 5 were run on the full unfiltered corpus, not the Arabic-filtered subset claimed in the headline. Either the Arabic filter is a separate, undocumented step that was never used in the analysis, or the 69M count is not actually the basis for the dataset's statistics. And no method, tool, or threshold for detecting Arabic is given anywhere, so the 69M figure can't be reproduced even with the files at hand.\n\nThe 'largest Arabic reviews dataset' claim also isn't benchmarked: the only direct comparison is Al-Shamani et al. at 51K, with no systematic pass over Masader's catalogue. Minor, but worth noting: no scraping or cleaning code is shipped, and the vocabulary analysis is done on all reviews, including the non-Arabic portion, so the \"noise\" estimates mix scripts.\n\nMy take: the resource could be valuable, and the paper deserves a serious referee, but the central count is currently unverified in the paper's own text. I'd want the authors to reconcile 76M versus 69M and document the language filter before trusting the headline. I wouldn't cite it until that is fixed.","headline":"Useful dataset at heart, but Table 3 contradicts the 69M Arabic review headline, leaving the central claim unverified.","tokens_in":7098,"tokens_out":3151,"would_cite":false,"duration_ms":27036,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GLARE releases 69 million Arabic app reviews, the largest Arabic corpus of its kind.","keywords":["Arabic NLP","app reviews","Google Play Store","sentiment analysis","aspect based sentiment analysis","feature engineering","exploratory data analysis","low-resource languages"],"falsifier":"Download a random sample of the raw 76M reviews, label it with a transparent Arabic-script and dialect criterion, estimate the Arabic proportion from the sample, and compare it with the 69M/76M ratio; a material mismatch would show the headline count is an artifact of the unspecified filter.","tokens_in":6188,"feed_emoji":"📱","tokens_out":13699,"duration_ms":111480,"temperature":0.7,"pith_summary":"GLARE is a new Arabic-language resource: 76,512,077 reviews harvested from the Saudi Google Play Store, of which the authors count 69 million as Arabic, across 9,980 unique Android applications. The paper's aim is to give Arabic NLP a corpus of app-store reviews at a scale that earlier Arabic review datasets do not approach, with the closest comparable dataset containing about 51,000 reviews. The corpus includes ratings, thumbs-up counts, developer replies, and app metadata, and the paper reports exploratory statistics and engineered features meant to support sentiment analysis, aspect-based sentiment analysis, and software-engineering analysis of user feedback. If the size claim holds, GLARE would supply the Arabic review domain with enough text for deep-learning-scale models rather than the small annotated sets that have dominated the area.","feed_headline":"GLARE puts 69 million Arabic app reviews into one corpus","feed_subtitle":"Saudi Google Play crawl spans 9,980 apps, giving Arabic NLP a review corpus far larger than earlier datasets.","key_machinery":"The load-bearing mechanism is the collection and filtering pipeline behind the corpus: a category-driven crawl of the Saudi Google Play Store that targets the top 200 free apps in each of 59 main and sub categories, maps apps to their category sets, drops duplicates to reach 9,980 unique applications, and scrapes reviews together with ratings, thumbs-up counts, developer replies, and app metadata. The step that turns the raw 76.5M reviews into the headline 69M Arabic figure is a preprocessing filter described only as 'keeping only Arabic reviews'; the paper does not specify how Arabic is detected, and that filter is what the size claim rests on. The feature-engineering pass builds a term-frequency dictionary with the token-counting utility CountVectorizer, producing 8.7M unique tokens and statistics on word length, review length, and noise that quantify how much cleaning downstream users will need.","core_discovery":"The claim on the paper's own terms is that GLARE is the largest Arabic reviews dataset to date. The crawl collected 76,512,077 reviews between March 21 and April 21 from the Saudi Google Play Store, starting from the top 200 free apps in each of 59 categories; after removing duplicated apps, the corpus covers 9,980 unique apps and 17 GB of raw review text. A preprocessing step that drops duplicates, nulls, symbols, numbers, and 'noise' and keeps only Arabic reviews yields the headline 69M Arabic reviews. The paper does not train or evaluate a model; its contribution is the resource itself plus descriptive analysis of ratings, thumbs-up votes, developer replies, vocabulary, review length, and app categories.","pith_inferences":["Until the Arabic-detection filter is documented, the 69M figure should be read as an order-of-magnitude estimate rather than an exact measurement.","Because the storefront is the Saudi Google Play, the Arabic text likely skews toward Saudi-region vocabulary and Modern Standard Arabic, making GLARE a geography-specific sample rather than a balanced pan-Arab corpus.","The authors list a domain-specific Arabic language model and an aspect-based sentiment analysis benchmark as future work, so the dataset's utility for those tasks remains untested in this paper.","If the one-word reviews are mostly generic praise, task-specific filtering could shrink the practically usable corpus well below 69M, so the effective size for NLP may differ from the headline size."],"forward_implications":["Arabic sentiment analysis and aspect-based sentiment analysis can be trained on tens of millions of in-domain reviews rather than thousands.","The rating attached to each review gives a free distant-supervision signal for opinion mining.","App-store reviews provide a text genre distinct from the Twitter posts that dominate most existing Arabic corpora, with a higher character ceiling and reply metadata.","Software engineering researchers and practitioners can use the corpus to study feature requests, bugs, rating dynamics, and developer-response behavior.","The reported vocabulary and noise statistics imply that downstream users will need substantial cleaning, especially because over 40% of reviews are a single word."],"supporting_citations":[{"why":"Prior Arabic Google Play review dataset with 51K reviews; supplies the scale baseline the paper must exceed to claim 'largest'.","marker":"Al-Shamani et al., 2022"},{"why":"Masader catalogue of over 200 Arabic NLP datasets; grounds the claim that most existing Arabic data comes from Twitter and that GLARE fills an underrepresented store-review slot.","marker":"Alyafeai et al., 2021"},{"why":"Supplies the motivation that Arabic subjective sentiment analysis needs large, good-quality datasets and that ABSA remains underexplored.","marker":"Nassif et al., 2021"},{"why":"LABR, a 63K Arabic book review corpus; one of the larger prior review-domain datasets GLARE is compared against by implication.","marker":"Aly and Atiya, 2013"},{"why":"Arabic hotel reviews dataset with roughly 38K reviews; another earlier review corpus establishing the scale GLARE exceeds.","marker":"Elnagar et al., 2018"},{"why":"Provides the CountVectorizer implementation used to build the 8.7M-word term dictionary and review-length statistics.","marker":"Pedregosa et al., 2011"},{"why":"Systematic literature review that motivates app reviews' value for software maintenance and evolution, the non-NLP use case for GLARE.","marker":"Dąbrowski et al., 2022"},{"why":"Shows app store effects on software engineering practices, supporting the claim that the data can aid developers' maintenance decisions.","marker":"Al-Subaihin et al., 2019"}],"fun_headline_variants":["GLARE's 69M Arabic reviews from Saudi Google Play","69M Arabic reviews from 9,980 apps in Saudi Google Play","Largest Arabic app review dataset to date: GLARE's 69M","GLARE dataset: 69M Arabic reviews, 9,980 Saudi apps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 69 million Arabic count rests on an undocumented language-filtering step, so the reported Arabic share depends on a filter the reader cannot inspect or reproduce.","fun_headline_variants_meta":{"raw":{"variants":["GLARE's 69M Arabic reviews from Saudi Google Play","69M Arabic reviews from 9,980 apps in Saudi Google Play","Largest Arabic app review dataset to date: GLARE's 69M","GLARE dataset: 69M Arabic reviews, 9,980 Saudi apps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001623,"raw_usage":{"total_tokens":6355,"prompt_tokens":742,"completion_tokens":5613,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":358,"completion_tokens_details":{"reasoning_tokens":5533}},"tokens_in":358,"tokens_out":5613,"duration_ms":36956,"temperature":1.0,"reasoning_tokens":5533,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:32:16.339020+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Download a random sample of the raw 76M reviews, label it with a transparent Arabic-script and dialect criterion, estimate the Arabic proportion from the sample, and compare it with the 69M/76M ratio; a material mismatch would show the headline count is an artifact of the unspecified filter.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior Arabic Google Play review dataset with 51K reviews; supplies the scale baseline the paper must exceed to claim 'largest'."},{"cited_title":"B., Elnagar, A., Shahin, I., and Henno, S","cited_arxiv_id":null,"evidence_quote":"Supplies the motivation that Arabic subjective sentiment analysis needs large, good-quality datasets and that ABSA remains underexplored."},{"cited_title":"and Atiya, A","cited_arxiv_id":null,"evidence_quote":"LABR, a 63K Arabic book review corpus; one of the larger prior review-domain datasets GLARE is compared against by implication."},{"cited_title":"S., and Einea, A","cited_arxiv_id":null,"evidence_quote":"Arabic hotel reviews dataset with roughly 38K reviews; another earlier review corpus establishing the scale GLARE exceeds."},{"cited_title":"A., Sarro, F., Black, S., Capra, L., and Harman, M","cited_arxiv_id":null,"evidence_quote":"Shows app store effects on software engineering practices, supporting the claim that the data can aid developers' maintenance decisions."}],"review_version":1}