{"id":"e78459ab-41ff-4f85-b85a-0b020772430a","arxiv_id":"2607.24571","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Seven popular GUI visualisation tools poorly support Arabic/RTL (especially Eastern numerals and maps), forcing 11 practitioners into multi-tool labour, simpler charts, and language compromises that the authors frame as infrastructural power.","lead":"Arabic-speaking visualisation practitioners must constantly work around tools built for English and left-to-right charts, manually fixing text, numerals, mirrors, and maps. The study shows how those defaults shrink designers’ agency and quietly encode linguistic and geopolitical norms.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The \"operationalise linguistic and geopolitical power\" claim is underdetermined: generic i18n debt and shared upstream geodata predict the same observations, and the one distinctly geopolitical evidence (map labelling) actually varies across tools, cutting against the \"single worldview\" framing in §","rationale":"The reader's weakest_assumption identified precisely the load-bearing soft spot: the power/colonial-legacy framing is interpretive synthesis over N=11 interviews plus a defaults audit, with no discrimination against i18n debt, market incentives, or education-system conventions. My pass confirms and sharpens it with two internal specifics: (a) the map-labelling evidence in §3.5.2 is bidirectionally inconsistent across tools, which undercuts the \"single worldview\" sentence in §5.2 even while supporting the weaker silent-override claim; (b) the paper's own practitioner quotes (P4/P8/P9) locate LTR habits in school-taught Cartesian conventions, leaving the direction of the §5.2 \"reinforcing chain\" ambiguous. These are refinements, not new objections, so the reader's CONDITIONAL verdict stands: the empirical core (tool audit + labour/compromise findings) is solid and the design implications are well supported; the power framing should be held as provisional. The proposed control-script audit and map-provenance trace are cheap, feasible extensions of the paper's own protocol and would directly settle whether the geopolitical framing carries explanatory weight beyond generic under-localisation. No correctness defects found in the tool-audit methodology itself (independent dual coding, default-settings protocol, OSF artifacts); sample skew and novelty incrementality were already priced into the verdict.","tokens_in":23383,"tokens_out":1595,"duration_ms":61971,"concrete_test":"Two checks. (1) Control-script audit: rerun the §3 protocol (same codebook, same three chart types) with a Hebrew RTL dataset and a low-resource LTR-script dataset (e.g., Thai). If Hebrew mirroring/numeral support fails in the same pattern as Arabic, the failures track generic i18n debt/market incentives rather than Arabic-specific geopolitical power, and the §5.2 framing should be narrowed accordingly. (2) Map-provenance trace: for each of the seven tools, identify the boundary/label data source (e.g., Natural Earth, proprietary geocoder) and check whether the Palestine/Western Sahara relabellings originate in the shared upstream dataset rather than per-tool design choices. If divergent labels trace to dataset versions/locale parameters, \"single embedded worldview\" fails and the correct claim is unexamined data dependency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that tool defaults \"operationalise linguistic and geopolitical power\" via a \"structurally reinforcing chain\" (§5.2) of cultural dominance → conventions → defaults → normalised practice. For this to hold, the observed failures must be better explained by power/dominance than by cheaper mechanisms: incomplete internationalisation, market-size-driven feature prioritisation, and inherited third-party data dependencies. The paper never runs that discrimination. Nearly all tool-audit findings — Eastern numerals parsed as text, absent mirroring, disconnected letter shaping, paywalled controls, English-only UIs — are exactly what generic i18n debt predicts for any low-resource script; no power analysis is needed. The one evidence type that is distinctly geopolitical rather than merely linguistic is map labelling (Palestine → \"West Bank and Gaza\"; Western Sahara as Morocco vs. independent). But §3.5.2 reports this behaviour as inconsistent across tools in both directions: Tableau and Google Sheets recognise \"Palestine\", Excel and Datawrapper do not; Excel folds Western Sahara into Morocco while four others show it separately. §5.2 then asserts \"a single worldview is embedded in most visualisation authoring tools\" — the paper's own Table 3 data shows multiple, divergent worldviews, which weakens the coherence of the embedded-power claim even as it strengthens the (smaller) claim that defaults are unexamined and silently override users. There is also a causal-direction problem: practitioner quotes (P4, P8, P9 in §4.2.2) attribute LTR habits to school-taught Cartesian conventions, i.e., the \"chain\" may run education → convention with tools as a late, reflecting node rather than the reinforcing engine. None of this makes the power framing wrong — it is a legitimate critical reading — but it is load-bearing for the paper's framing and is asserted, not discriminated from alternatives the paper itself documents.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper investigates how Arabic-speaking visualisation practitioners design for right-to-left (RTL) scripts, combining two studies: (1) an analytical evaluation of seven GUI-based visualisation authoring tools (Excel, Tableau, Power BI, Flourish, Datawrapper, Google Sheets, RAWGraphs) using an Arabic books dataset rendered in both Eastern and Western numerals, coded with an explicit deductive codebook by two native-Arabic-speaking coders; and (2) semi-structured interviews with 11 Arabic-speaking practitioners (journalism, design, data analysis) analysed via thematic analysis to saturation (26 codes, 4 themes). The main findings are: tool support for Arabic is fragmented and inconsistent, particularly for Eastern numerals (frequently parsed as text or silently converted), RTL mirroring (partial, heterogeneously labelled, never default), and map defaults (inconsistent treatment of contested regions such as Palestine and Western Sahara). Practitioners perform substantial manual labour (multi-tool workflows, manual mirroring, axis removal) and make strategic compromises on language, chart complexity, and interactivity. The discussion (§5.1–5.2) frames these findings as constrained agency and argues that tool defaults \"operationalise linguistic and geopolitical power,\" closing with a research agenda (RTL replication of perception studies, RTL-adapted literacy instruments) and concrete design recommendations.","tokens_in":23705,"tokens_out":4345,"duration_ms":156385,"significance":"If the results hold, this is a valuable contribution on two fronts. Empirically, it documents — with falsifiable, screenshot-verifiable evidence — that mainstream GUI visualisation tools fail on foundational tasks for Arabic-script work (Eastern numerals parsed as text, silent numeral conversion, no default mirroring, inconsistent map labelling), and triangulates this with practitioner accounts of the compensating labour (multi-tool workflows, manual mirroring, chart-type and language compromises). This triangulation between an artefact audit and lived practice is methodologically stronger than either alone. The paper ships an explicit codebook, dual independent native-speaker coding with reconciliation, a documented path from 26 codes to 4 themes with a saturation criterion, and OSF supplementary material including the Eastern-numeral dataset and all charts. The design implications in §5.3 (consistent controls, native numeral/city-name handling, transparency over silent overrides) are concrete and actionable for tool builders. The work also usefully extends the cross-cultural visualisation literature beyond perception studies to the authoring infrastructure itself. The interpretit","major_comments":[{"comment":"§5.2's claim that 'a single worldview is embedded in most visualisation authoring tools' is contradicted by the paper's own Table 3/§3.5.2 results: Tableau and Google Sheets recognise 'Palestine' while Excel and Datawrapper convert it to 'West Bank and Gaza', and Excel folds Western Sahara into Morocco while four other tools display it separately. This is multiple, divergent worldviews, not a single one. Relatedly, §5.2 presents the 'structurally reinforcing chain' of cultural dominance -> conventions -> defaults as the explanation, but most of the §3.5 findings (Eastern numerals parsed as text, absent mirroring, disconnected letter shaping, English-only UIs) are equally and more parsimoniously predicted by generic low-resource-script internationalisation debt, market-size-driven feature prioritisation, and inherited third-party geodata (e.g., Natural Earth boundaries). The paper never d","section":"§5.2 / Table 3"},{"comment":"§4.1 states that group meetings 'led to identifying and refining the themes to those supporting our research thesis.' As written, this suggests themes were selected for fit to a pre-held thesis rather than derived from the data, which is a rigor concern for a thematic analysis and is load-bearing for the interpretive Theme 4.2.4 (embedded power). Please clarify the analytic stance: was 'embedded power' an a priori sensitising concept (deductive element) or emergent? How were disconfirming cases handled (e.g., P1 and P8, who actively oppose mirroring and whose positions sit uneasily with the constraint narrative)? The discarded data-journalism theme is a good example of the kind of reporting needed; the wording should be corrected even if the underlying procedure was sound.","section":"§4.1"},{"comment":"The translation pipeline in §4.1 (standardise colloquial dialects to Modern Standard Arabic, machine-translate via Google Translate [70], revise by the first author) is reasonable in outline, but the verification rests on a single native speaker, and direct quotations carry substantial evidentiary weight in §4.2.4 (e.g., the P9/P10 'English is more elegant' quotes that anchor the colonial-legacy interpretation). Please describe an independent check — e.g., a second native speaker verifying a sample of translated excerpts, or participant member-checking of quoted passages — or acknowledge this as a limitation. This matters most where quoted material supports the paper's most contestable interpretive claims.","section":"§4.1"}],"minor_comments":[{"comment":"§3.5.2: 'All the axis-based charts we rendered with default settings applied the LTR convention (e.g. RTL X-axis and Left-aligned Y-axis)' — 'RTL X-axis' here contradicts 'LTR convention'; presumably 'LTR X-axis' is meant. Please fix.","section":"§3.5.2"},{"comment":"Tables 2 and 3: the encoding glyphs (checkmarks/crosses) did not render in this version, and Table 2's purple/blue colour coding is the sole carrier of category information, which will fail in grayscale print and for colourblind readers. Add a redundant symbol or label encoding. Also, with n=7 tools, phrases like 'above 71%' and 'above 57%' should be reported as counts (5/7, 4/7).","section":"Tables 2–3"},{"comment":"The 'over two billion Arabic script users' figure ([3], a W3C page on RTL scripts) appears to cover users of all RTL scripts (Urdu, Persian, etc.), not Arabic-script users specifically; please either adjust the wording to match the source or cite a population figure that supports the claim as stated.","section":"§1 / Abstract"},{"comment":"Eastern (Arabic-Indic) numerals are not uniform across the Arab world — the Maghreb predominantly uses Western digits, which is visible in your own data (P10/P11 describe producing French/Western-numeral outputs in Tunisia). A sentence in the limitations acknowledging that 'Arabic visualisation' practice is itself internally heterogeneous would strengthen the generalisation discussion in §5.3.","section":"§5.3"},{"comment":"Tool evaluation: please report tool versions and the evaluation window explicitly (§3.3 mentions coders worked in the same period but no dates or versions are given). Tool behaviour — e.g., Google Sheets' silent Eastern-to-Western numeral conversion — is exactly the kind of finding that may change between versions, and versioning makes the audit reproducible and updatable.","section":"§3.3–3.5"},{"comment":"§5.2, final paragraph: the ChatGPT/Matplotlib disconnected-letter example is anecdotal (a single informal observation). Either soften the claim ('we cannot yet rely on generative AI...') or support it with the same default-settings procedure used in §3, applied to one or two LLM pipelines.","section":"§5.2"},{"comment":"Figure 1 is referenced for several distinct phenomena (mixed presentation, disconnected letters, map relabelling) but the caption is generic; please annotate subpanels so readers can locate each phenomenon.","section":"Figure 1"},{"comment":"References [36] and [37] are the same Huang et al. study listed twice (2022 and 2025 versions); consolidate.","section":"References"},{"comment":"The tool-audit coding reports that the two coders reconciled disagreements by recreating charts; reporting the raw disagreement count (or an agreement statistic) before reconciliation would give readers a sense of codebook reliability without much additional space.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The empirical core of this paper is solid and well suited to the journal. The interpretive framing in §5.2 (colonial legacy, embedded geopolitical power) is where reviewers are likely to split; my report asks the authors to bring that framing into line with their own Table 3 evidence rather than to remove it, which I think is the proportionate ask. Self-citation to Alebri et al. [7] is legitimate lineage (the open question motivating the interview study), not padding. The manuscript header indicates this is an author's version of an accepted TVCG article; if this submission is a duplicate of an already-accepted version, the editor may wish to confirm status."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the paired evidence: a coded default-settings audit of seven named GUI tools on Arabic text and Eastern numerals, plus eleven practitioner interviews that document the real labour (manual mirroring, multi-tool stitching, dropping interactivity, switching to English or basic charts). That core is new relative to the 2024 pattern paper and to generic i18n checklists, and it is useful for vendors and for anyone working on non-LTR visualisation practice.\n\nWhat they did well: explicit codebook, dual native-speaker coding, Eastern vs Western numeral runs, and saturation-minded thematic analysis. The OSF materials and the concrete failure modes (numerals-as-text, inconsistent mirror labels, silent map renames) make the empirical claims checkable. The agency/creativity discussion lands because the quotes and the tool tables point the same way.\n\nSoft spots, in proportion. The load-bearing claim that tools “operationalise linguistic and geopolitical power” via a reinforcing chain is interpretive synthesis, not a discrimination against cheaper accounts (incomplete i18n, market prioritisation, school-taught Cartesian axes, upstream geodata). The paper’s own map results cut against “a single worldview”: tools diverge on Palestine and Western Sahara. Practitioner literacy themes already supply the education-convention path. Treat the power section as a critical frame, not as established mechanism. Sample is small and male-skewed; generalisation beyond Arabic GUI practice should stay modest. Novelty is real but incremental on their prior work.\n\nMath/data/citations look fine for qualitative systems HCI—no circular equations, self-cites are lineage not force-fitting, related work is adequate. This is for VIS/HCI people who care about tooling, localisation, or non-WEIRD practice, and for product teams. It deserves a serious referee. I would engage: cite the audit and the labour findings; hold the power thesis lightly until someone tests alternatives.","headline":"Solid dual study on Arabic/RTL tool gaps and practitioner labour; the power framing is a fair reading but overclaims relative to what the audit actually shows.","tokens_in":24630,"tokens_out":483,"would_cite":true,"duration_ms":14547,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Popular visualisation tools force Arabic-script designers into heavy workarounds that shrink agency and encode LTR and geopolitical defaults as normal.","keywords":["Arabic visualisations","right-to-left","design practice","tool comparison","Eastern numerals","RTL mirroring","data visualisation defaults","localisation"],"falsifier":"A controlled study with Arabic-script audiences comparing comprehension and trust for mirrored versus default LTR charts (and for native versus tool-default map labels) would show whether LTR defaults actually harm understanding or engagement, or whether practitioners’ fears and workarounds are largely habit and tooling friction without audience cost.","tokens_in":24335,"feed_emoji":"📊","tokens_out":901,"duration_ms":18568,"temperature":0.7,"pith_summary":"This paper shows that Arabic-speaking visualisation practitioners design under constant friction between right-to-left reading habits, assumed “universal” left-to-right chart norms, and tools that only half-support Arabic. An evaluation of seven mainstream GUI authoring tools with an Arabic dataset finds inconsistent RTL mirroring, weak handling of Eastern numerals, and map defaults that rename or redraw contested places without notice. Interviews with eleven practitioners in journalism, design, and analysis show the human cost: manual mirroring, multi-tool stitching, loss of interactivity, and strategic retreats to simpler charts or English. The authors argue these are not minor localisation gaps but ways tools operationalise linguistic and geopolitical power, constraining creativity and agency for over two billion Arabic-script users. The work matters because it turns a largely invisible labour tax into a concrete research and design agenda for RTL visualisation support.","feed_headline":"Arabic chart tools force hours of manual fixes","feed_subtitle":"Seven tools and 11 practitioners show LTR defaults tax agency, numerals, maps, and creativity","key_machinery":"A paired method: analytical evaluation of seven popular GUI tools (Excel, Tableau, Power BI, Flourish, Datawrapper, Google Sheets, RAWGraphs) on an Arabic dataset for default RTL/numeral/map behaviour, plus thematic interviews with eleven Arabic-speaking practitioners whose reported tools overlap that set. Together they surface how LTR defaults and silent overrides structure everyday design decisions.","core_discovery":"Visualisation authoring tools and their defaults operationalise linguistic and geopolitical power in RTL contexts: fragmented support for Arabic text, Eastern numerals, RTL mirroring, and map labelling pushes practitioners into substantial multi-tool labour and compromises on language, interactivity, and chart type that limit agency and creativity.","pith_inferences":["The same default asymmetry likely taxes other RTL scripts (Persian, Urdu, Hebrew) even when the paper’s Arabic case is the only one studied; script-specific native support may beat one-size “RTL mode.”","Because LLM and auto-viz pipelines often inherit Matplotlib-style LTR/English defaults, Arabic chart generation failures will keep propagating unless reshaping and numeral support become first-class.","Paywalled advanced controls turn linguistic inclusion into a class issue: free tiers leave Arabic practitioners with more manual labour than English peers on the same product."],"forward_implications":["Tool makers should treat Eastern numerals and Arabic place names as native data types, expose consistent RTL mirroring controls, and surface silent conversions instead of overriding without notice.","Empirically tested RTL design guidance is needed so practitioners stop relying on personal assumptions about audience literacy and direction habits.","Visualisation literacy instruments and core perception studies should be replicated and validated with RTL script users rather than assumed transferable from LTR populations.","Map defaults that re-label or omit contested regions become political acts when non-experts accept them under time pressure; tools need disclosure and alternate framings.","Constrained outputs that circulate as “normal” Arabic charts risk design fixation, narrowing what future Arabic visualisations look like."],"fun_headline_variants":["Arabic viz tools force manual RTL fixes and multi-tool labour","LTR defaults tax agency for Arabic chart practitioners","Seven tools show fragmented RTL, numeral, and map support","Practitioners compromise language and chart type for Arabic data","Viz defaults operationalise linguistic power in RTL contexts"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the tool gaps, workarounds, and map labelling choices are best read as a reinforcing chain of cultural dominance rather than mainly incomplete internationalisation, market incentives, or school-taught Cartesian habits.","fun_headline_variants_meta":{"raw":{"variants":["Arabic viz tools force manual RTL fixes and multi-tool labour","LTR defaults tax agency for Arabic chart practitioners","Seven tools show fragmented RTL, numeral, and map support","Practitioners compromise language and chart type for Arabic data","Viz defaults operationalise linguistic power in RTL contexts"]},"model":"grok-4.5","effort":"low","cost_usd":0.00406,"raw_usage":{"total_tokens":1259,"prompt_tokens":819,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":40604000,"prompt_tokens_details":{"text_tokens":819,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":380,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":819,"tokens_out":60,"duration_ms":5929,"temperature":1.0,"reasoning_tokens":380,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T11:28:48.290369+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled study with Arabic-script audiences comparing comprehension and trust for mirrored versus default LTR charts (and for native versus tool-default map labels) would show whether LTR defaults actually harm understanding or engagement, or whether practitioners’ fears and workarounds are largely habit and tooling friction without audience cost.","supporting_citations":[],"review_version":1}