{"id":"fdb62a5b-2059-48d9-958b-59196748d211","arxiv_id":"2501.17785","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Current vision-language and text-only models perform poorly on deciphering non-Unicode rare scripts, and Unicode encoding helps only for common languages, not low-resource ones.","lead":"This paper tests how well vision-language models like GPT-4o, Gemini, and Claude 3.5 can decode rare writing systems that are not encoded in Unicode, using custom glyph tokens and picture- or description-based prompts. It finds that all models struggle with visual token descriptions and geometric reasoning, while Unicode encoding only helps for widely known languages, not low-resource ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated glyph-token segmentation (§3) may not match real orthographic units, so the puzzle results may not measure decipherment of the scripts as written.","rationale":"The reader's weakest_assumption identifies the glyph-token segmentation rule as the point where the central claim is least secure. I agree: the paper's method is a novel but unvalidated preprocessing step, and every downstream measurement depends on it. The paper itself flags the subjectivity of the exception rule and notes that mirrored Avoiuli glyphs were treated as distinct despite being the same token in the script, which is direct evidence that the segmentation can diverge from orthographic reality. This is a correctness risk, not merely a disagreement with prior consensus; the dataset is not released and no expert validation is provided, so the concern cannot be resolved from the manuscript. A concrete expert-segmentation comparison and a re-run of at least one model on the resulting puzzles would settle whether the reported failures are robust to tokenization choices. Because the reader already issued CONDITIONAL, this concern does not move the verdict; it reinforces the condition. No alternative concern is more load-bearing: the small sample size and lack of error bars weaken generality but do not undermine the internal validity of the tokenization as directly.","tokens_in":139,"tokens_out":1604,"duration_ms":27978,"concrete_test":"Have at least two independent linguists or script experts segment the five puzzle images using the paper's rule plus a written definition of 'horizontal extension'; measure inter-annotator agreement on token boundaries (e.g., boundary F1). Then rerun one model (e.g., GPT-4o) on the puzzles using the expert consensus segmentation instead of the authors' segmentation. If inter-annotator agreement is below 0.90, or if the model's correctness on any puzzle changes (e.g., Avoiuli mirror-token handling, Mandombe component-splitting), the segmentation rule is not consistent enough to support the paper's qualitative conclusions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that current LVLMs/LLMs struggle to decipher non-Unicode scripts rests on five puzzles tokenized by the paper's own visual rule (§3): adjacent glyphs are separate tokens if a vertical white line passes through their gap, with a subjectively judged exception for horizontal extensions. This rule is never validated against the scripts' actual orthographic units. Avoiuli is a concrete problem: the paper states that mirrored glyphs should be the same token, yet it tokenized them as distinct and never tested whether models could recover the equivalence. For Mandombe and Ditema tsa Dinoko, where glyph-internal components carry phonological meaning, the rule may split at the wrong level entirely; if the placeholder sequence does not correspond to the script's real grapheme/syllable structure, then every measured failure—poor token description, failed geometric reasoning, no phonological analysis—is an artifact of the constructed representation rather than a finding about decipherment. The paper explicitly concedes the exception criterion 'may vary among individuals' (§3), which means the dataset is not uniquely determined and the published numbers (e.g., below 15% reverse-mapping accuracy for Mandombe, 40.0%/13.4%/31.3% description-pairing accuracy) cannot be reproduced or compared without the exact segmentation. Since the tokenization is the foundation for both the Picture Method and the Description Method, an incorrect segmentation invalidates the central claim that current models fail at deciphering these scripts as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a visual glyph-tokenization method for representing scripts that are not Unicode-encodable, then evaluates GPT-4o, Gemini, and Claude 3.5 Sonnet on five linguistic puzzles drawn from UKLO and NACLO. Two prompting strategies are introduced: the Picture Method for vision-language models and the Description Method for text-only models. Experiments compare token-description generation, puzzle-solving under placeholders/images/descriptions, and the effect of Unicode encoding for a common language (Malayalam) versus a low-resource language (Meroitic). The main reported findings are that current models struggle with non-Unicode scripts, especially in token description and geometric reasoning, and that Unicode encoding helps only for common languages.","tokens_in":6885,"tokens_out":2979,"duration_ms":31758,"significance":"If the method and measurements are reliable, the paper addresses a genuine gap by constructing multimodal puzzles for scripts outside Unicode coverage and by proposing a tokenization scheme that lets both LVLMs and LLMs work with non-encodable scripts. The use of externally sourced UKLO and NACLO puzzles is a strength, as is the explicit acknowledgment of limited scope in Section 5.4. The qualitative observations about model behavior and the demonstration that Unicode support alone does not help low-resource languages are useful for the community. However, the significance is currently constrained by the small number of puzzles, the single-run nature of the experiments, and the lack of validation for the central tokenization step.","major_comments":[{"comment":"The glyph-tokenization rule is not validated against the scripts' actual orthographic units, and the exception for horizontal extensions is explicitly subjective ('The criteria for identifying and applying this exception may vary among individuals'). This rule is the foundation of both the Picture Method and the Description Method, since all placeholders are derived from it. If the segmentation does not match real grapheme or syllable boundaries, then the reported accuracies, including the description-pairing figures in Section 5.1 and the 'below 15%' reverse-mapping rates in Section 5.2, do not measure decipherment of the scripts as written. The authors should provide inter-annotator agreement, linguistic validation of the token boundaries, or an ablation showing that results are robust across reasonable alternative segmentations.","section":"Section 3"},{"comment":"For Avoiuli, the paper states that mirrored glyphs should be considered the same token, yet it explicitly treats them as distinct tokens during tokenization and does not test whether the models can recover the equivalence. Since the paper lists Avoiuli as a case where token identification is the primary challenge, this mismatch means the evaluation does not fully reflect the linguistic property that the authors themselves identify. The reported failure to infer mirror equivalence is therefore not a fair test of the models' decipherment ability. Please either align the tokenization with the stated orthographic property or present mirror handling as a separate controlled experiment.","section":"Section 5.2"},{"comment":"The quantitative evidence is reported inconsistently and at too coarse a grain. The description-pairing accuracies (40.0%, 13.4%, 31.3%) are given without the number of tokens, the number of trials, or any variance measure, and the reverse-mapping success for Mandombe is only described as 'below 15%' without exact values or per-model breakdowns. With only five puzzles and single runs, these numbers cannot be compared across models or across puzzles. The authors should report exact per-puzzle, per-model results with counts and, ideally, multiple runs, and should state which numbers are averages over which items.","section":"Section 5.1 and 5.2"},{"comment":"The paper's own limitation statement acknowledges that the same prompt was used for every model and that only five puzzles were tested, yet the abstract and conclusion make general claims about LVLM and LLM capabilities in deciphering rare scripts. This mismatch between scope and generalization is load-bearing for the paper's framing. A discussion of how the fixed prompt and small puzzle set might bias the results, together with a restriction of the central claims to the five tested puzzles, would make the conclusions more defensible.","section":"Section 5.4"}],"minor_comments":[{"comment":"Figure 1 is not described in enough detail in the text; readers cannot see the exact token boundaries or the horizontal-extension exception being applied. Adding an annotated example with boxes or arrows around the relevant gaps would make the tokenization rule reproducible.","section":"Section 3, Figure 1"},{"comment":"The term 'preliminary experiments' is used in Section 4, but this hedged framing is not carried through the abstract and conclusion, which present the findings as established strengths and limitations. The wording should be made consistent throughout.","section":"Section 4"},{"comment":"Appendix 7.1 references Figure 3 and Table description examples, but the appendix text does not include the actual description table or the full set of model-generated descriptions. Including the complete table would let readers judge the qualitative claims about 'Lack of Details', 'Direction Confusion', and 'Misleading Elaboration'.","section":"Appendix 7.1"},{"comment":"The models are referred to inconsistently as GPT-4o, Gemini, and Claude 3.5 Sonnet, while Figure 2 and the text later use 'Claude-3.5'; please standardize the notation. Also, the paper does not provide the specific prompt templates, model version numbers, or decoding parameters, which are needed for reproducibility.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The central idea is interesting and the qualitative results are plausible, but the contribution currently rests on an unvalidated and partially subjective tokenization rule. The paper is short and reads as a preliminary study; the authors should either substantially expand the validation and quantitative reporting or narrow the claims. I do not see a fatal flaw that would require rejection, provided the tokenization concern is addressed with additional experiments or careful caveats."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version: this paper is the first I know of to test LVLMs and LLMs on linguistic olympiad puzzles written in scripts that can't be rendered in Unicode, by swapping glyphs for placeholders. That's a real contribution, even if the results are expected: current models struggle with visual token description and geometric reasoning, and Unicode encoding only helps when the model already knows the language.\n\nWhat's actually new and useful: the picture method/description method comparison, and the observation that models fail to do phonological analysis on Mandombe and Ditema. The paper is also honest about its limits — it lists the five-puzzle scope and the single-prompt design.\n\nThe soft spots are real but proportionate. Five puzzles, one run per model, no released data, and the quantitative reporting is loose ('below 15%' without exact numbers). The tokenization rule in Section 3 is the biggest worry: it's a heuristic that the authors themselves say may vary across individuals. The stress-test is right that this rule is never validated against the scripts' actual orthographic units. Avoiuli is a concrete problem — the paper admits mirrored glyphs were treated as distinct despite being the same token — and for Mandombe/Ditema the segmentation could split at the wrong level entirely.\n\nThat said, I don't think this invalidates the central qualitative finding. Even with a slightly off segmentation, the models are shown actual glyph images and still fail to map them to placeholders or reason about components. The exact accuracy numbers, though, are not trustworthy as reported and shouldn't be used for model comparison until the tokenization is pinned down and the data released.\n\nSo: this is a promising early exploration with a useful benchmark idea, but it's not a publishable result yet. A serious refereeing process could help the authors fix the reproducibility problems. I'd accept it to peer review, expecting major revisions.\n\nMy recommendation: if the authors release the benchmark, validate the tokenization (could they ask human annotators to segment the scripts?), run multiple seeds, and report exact numbers, the paper would be worth citing. As it stands, it's a discussion piece, not a measurement paper.","headline":"Small but genuinely novel probe of LVLMs on non-Unicode script puzzles; the tokenization is subjective and evidence is thin, but the core difficulty finding is plausible and the paper deserves a chance to grow through review.","tokens_in":7273,"tokens_out":3640,"would_cite":false,"duration_ms":36231,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current vision-language and text models cannot decipher rare scripts that lack Unicode encoding, and Unicode representation improves performance only for common languages.","keywords":["rare scripts","decipherment","glyph tokenization","vision-language models","linguistic puzzles","low-resource languages","Unicode encoding","visual reasoning"],"falsifier":"Give the same five puzzles to expert annotators to segment glyphs according to the orthographies and rerun the models on those versions. If expert-segmented puzzles yield substantially different model accuracies, or if the vertical-white-line rule is shown to split single letters or merge distinct ones in Meroitic or Mandombe, the central performance claims are artifacts of the tokenization rather than a measure of model capability.","tokens_in":6301,"feed_emoji":"🔤","tokens_out":6808,"duration_ms":61313,"temperature":0.7,"pith_summary":"This paper asks whether current vision-language and text-only models can decipher scripts that cannot be typed in Unicode, and answers with a mostly negative result. It builds a small set of linguistic puzzles from olympiad materials, replaces each glyph with a placeholder token, and then feeds the puzzles to GPT-4o, Gemini, and Claude 3.5 Sonnet in three forms: images of the glyphs, written descriptions of the glyphs, or placeholders alone. Across the non-Unicode puzzles, the models rarely reconstruct the scripts' sound systems and frequently misdescribe or misread the geometry of the glyphs. On the encoded side, Unicode helps the models use memorized knowledge for a common language (Malayalam) but does not rescue a low-resource language (Meroitic). The upshot is that current models can do pattern matching but not true decipherment, and that placeholders offer a way to test reasoning without pretraining shortcuts.","feed_headline":"AI models fail to decipher scripts Unicode can't encode","feed_subtitle":"Five puzzles show GPT-4o, Gemini, and Claude 3.5 fail at glyph reasoning; Unicode only helps common languages.","key_machinery":"The central object is the glyph token, defined as the fundamental unit of visual information in an unknown script; two adjacent glyphs are split into separate tokens when a vertical white line can pass through their gap, with a subjective exception for horizontal extensions bridging the gap. After segmentation, each glyph token is replaced by a placeholder such as <token_i>, so reasoning can be tested without exposing the model to real characters. That machinery supports two input methods: the Picture Method, which shows an image of all glyphs labeled with placeholders to a vision-language model, and the Description Method, which supplies a table of verbal descriptions of each glyph's appearance and relations to a text-only model.","core_discovery":"The paper's central claim is that contemporary LVLMs and LLMs cannot reliably decipher rare scripts that lack Unicode encoding, and that their failures follow a regular pattern: token descriptions are imprecise or geometrically wrong, direction is confused, and reasoning from images often misfires. For glyphs made of indivisible tokens (Avoiuli), the models can sometimes start down the right path by counting and matching tokens, but they cannot complete a syllable-token mapping. For glyphs whose subcomponents carry phonological information (Mandombe, Ditema tsa Dinoko), no model performs phonological analysis and progress stalls. The paper further claims that Unicode encoding is not a cure: it unlocks pre-trained knowledge for a common language but, for Meroitic, even correct Unicode characters and directionality leave performance poor because the models lack the underlying language data.","pith_inferences":["A natural next experiment is to feed Unicode-encoded common languages through the same placeholder pipeline; if performance collapses, it would confirm that placeholders remove retrieval without requiring visual decipherment.","The 35-word description limit may itself cause failures; models might succeed with structured or relational description formats such as coordinates or component graphs, so the paper's negative result may overstate the ceiling of description-based decipherment.","Connecting the glyph-description task to existing spatial-reasoning benchmarks would isolate whether the failure is about language, vision, or the mapping between the two.","If the segmentation rule is applied at scale, it could become a curriculum for testing whether a model detects visual token boundaries, since the same image should yield stable splits under expert annotation."],"forward_implications":["Token-description accuracy is a bottleneck: if descriptions cannot reliably distinguish glyphs, then the Description Method cannot support decipherment even when the underlying language rules are learnable.","Unicode encoding is a proxy for pretraining exposure, not a substitute for it; adding Unicode to low-resource scripts does not make models reason better.","Placeholder-based puzzles can serve as a diagnostic that separates genuine linguistic reasoning from retrieval of memorized script knowledge.","For scripts with compositional glyphs, models need phonological and structural analysis of subcomponents; current vision-language models instead fixate on surface geometry.","Mirrored glyphs and other non-orthographic visual conventions remain undetected by models, so future datasets or prompts must either make these conventions explicit or test for their discovery."],"supporting_citations":[{"why":"Supplies the Rosetta Stone linguistic-puzzle format from which the paper's UKLO-style puzzles are drawn.","marker":"Bozhanov and Derzhanski, 2013"},{"why":"ModeLing benchmark for linguistics-olympiad reasoning that the paper extends to visual and rare scripts.","marker":"Chi et al., 2024"},{"why":"Lingoly benchmark of olympiad-level puzzles in low-resource and extinct languages, providing the comparison point for scarce-language reasoning.","marker":"Bean et al., 2024"},{"why":"GPT-4 technical report; the model paper evaluates as GPT-4o.","marker":"Achiam et al., 2023"},{"why":"Gemini model report; the model paper evaluates.","marker":"Team et al., 2024"},{"why":"Claude 3.5 Sonnet release; the model paper evaluates.","marker":"Anthropic, 2024"}],"fun_headline_variants":["AI fails to read rare scripts that Unicode lacks","Unicode doesn't rescue AI on rare script puzzles","GPT-4o, Gemini, Claude all stumble on glyph riddles","Rare script puzzles expose AI reasoning blind spots","No Unicode means no luck: LLMs fail at glyph decipherment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the glyph-token segmentation rule, which treats any gap a vertical white line can pass through as a token boundary and applies a subjective exception for horizontal extensions, produces the correct and consistent units for all five puzzles; if the segmentation does not match the scripts' real orthographic units, the measured accuracy says nothing about deciphering the scripts as written.","fun_headline_variants_meta":{"raw":{"variants":["AI fails to read rare scripts that Unicode lacks","Unicode doesn't rescue AI on rare script puzzles","GPT-4o, Gemini, Claude all stumble on glyph riddles","Rare script puzzles expose AI reasoning blind spots","No Unicode means no luck: LLMs fail at glyph decipherment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000586,"raw_usage":{"total_tokens":2707,"prompt_tokens":850,"completion_tokens":1857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":1776}},"tokens_in":466,"tokens_out":1857,"duration_ms":13535,"temperature":1.0,"reasoning_tokens":1776,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:31:52.499642+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same five puzzles to expert annotators to segment glyphs according to the orthographies and rerun the models on those versions. If expert-segmented puzzles yield substantially different model accuracies, or if the vertical-white-line rule is shown to split single letters or merge distinct ones in Meroitic or Mandombe, the central performance claims are artifacts of the tokenization rather than a measure of model capability.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Rosetta Stone linguistic-puzzle format from which the paper's UKLO-style puzzles are drawn."},{"cited_title":"McCoy, and Dragomir Radev","cited_arxiv_id":null,"evidence_quote":"ModeLing benchmark for linguistics-olympiad reasoning that the paper extends to visual and rare scripts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Claude 3.5 Sonnet release; the model paper evaluates."}],"review_version":1}