{"id":"fc3422c5-3c87-4f70-9e14-9a40d39f8638","arxiv_id":"2506.07896","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"On an unvalidated AI-graded zero-shot test, several commercial LLMs score higher and more consistently than open-source small models on simplified versions of the frame and symbol grounding problems.","lead":"The paper creates two quiz-style prompts inspired by classic AI philosophy problems, the frame problem and the symbol grounding problem, and gives them to 13 chatbots with no training examples. A single AI judge, ChatGPT-4o, then grades all answers, and the study reports that several paid models score higher and more steadily than open-source ones.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on the unvalidated subjective ratings of ChatGPT-4o, which is itself one of the scored models; a human inter-rater check is needed before any capacity conclusion can be drawn.","rationale":"The reader identified the unvalidated LLM rater as the weakest assumption, and my reading converges on the same point. The paper's headline claim is entirely an interpretation of scores generated by ChatGPT-4o, one of the evaluated models. The absence of any correlation with human ratings makes the 'high scores' ambiguous: they might reflect the rater's aesthetic biases, fluency preferences, or even self-preference toward its own output style. Because the abstract advances a substantive claim about cognitive capacities, the measurement instrument must be shown to measure those capacities. The paper's own limitations section (5.5) explicitly flags this. The proposed human-scoring check would directly test whether the rater's scores are reproducible and whether the model ranking changes under independent evaluation. If it fails, the central claim collapses; if it passes, the paper's descriptive results gain some empirical footing. I therefore agree with the REJECT verdict: the manuscript does not currently support its central claim, though the concern is empirical and could be addressed by additional validation.","tokens_in":25891,"tokens_out":3406,"duration_ms":45216,"concrete_test":"Have two independent human raters with relevant expertise blindly score all 130 responses (13 models x 2 tasks x 5 trials) using the same six-criteria rubric, without model names. Compute inter-rater reliability between ChatGPT-4o's scores and the human average (e.g., Krippendorff's alpha or ICC) and compare the ranking of the 13 models. If agreement is below 0.7 or if human rankings place a different model at the top (e.g., ChatGPT-4o not near the top), the rater-based results are not a valid basis for the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that 'select modern LLMs may be acquiring capacities sufficient to produce meaningful and stable responses' to the frame and symbol grounding problems. Every score supporting this claim is produced by a single rater LLM, ChatGPT-4o (Section 3.4), which is also subject model No.1 in Table 1. The rater's six criteria (e.g., creativity, expression) are subjective aesthetic judgments, not calibrated against any external standard. The paper itself concedes in Section 5.5 that no human inter-rater reliability has been established. Without evidence that ChatGPT-4o's scores correspond to human judgments or to the intended theoretical capacities, the observed high scores could simply reflect the rater's stylistic preferences. The anonymization by number does not eliminate this concern, since the rater may still favor outputs similar to its own. Thus the abstract's conclusion about LLMs 'acquiring capacities' is unsupported by the measurement instrument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a zero-shot benchmark intended to test whether LLMs can address the frame problem and the symbol grounding problem. Two prompts are designed as loose reformulations of these problems and administered to 13 model settings (four closed, nine open-source). Each model's five outputs are scored by ChatGPT-4o on six 10-point criteria per task, yielding means and standard deviations. The paper reports that closed models, especially ChatGPT o3, score consistently high, while small open-source models score very low, and it concludes that some modern LLMs may be acquiring capacities relevant to these long-standing challenges.","tokens_in":26111,"tokens_out":5161,"duration_ms":65914,"significance":"If the measurement instrument were valid, this would be a useful empirical bridge between philosophical AI problems and contemporary LLM evaluation. The paper is transparent in several ways: the exact prompts are given, the Python scripts for open-source models are appended, all raw outputs and per-trial scores are included in the appendix and linked GitHub repository, and Section 5.5 candidly acknowledges the main limitations. That transparency is a real strength. However, the significance of the empirical contribution is currently blocked by a fundamental validity problem in the scoring procedure and by the loose linkage between the tasks and the named theoretical problems.","major_comments":[{"comment":"Section 3.4 states that all responses are scored by ChatGPT-4o as the sole rater LLM, and Table 1 lists ChatGPT 4o as subject model No.1. Section 5.5 concedes that no human inter-rater reliability has been established. Every score in Tables 2.a and 2.b is therefore a subjective rating produced by a model that is itself one of the subjects in the same benchmark. The high scores could reflect the rater's stylistic preferences or self-similarity rather than the intended cognitive capacities. This is load-bearing for the abstract's conclusion that several closed models 'may be acquiring capacities sufficient to produce meaningful and stable responses'; without an external anchor, the quantitative results do not support that claim.","section":"3.4 and Table 1"},{"comment":"The prompts are only loose proxies for the frame problem and the symbol grounding problem. The frame prompt is an urban route-planning exercise with ten listed events and an information update; it tests relevance filtering and replanning, but it does not engage the frame problem's core formal issues of representing non-change or the qualification/ramification problems. The kluben prompt asks for creative narrative elaboration around a few sensory-like properties; it does not operationalize Harnad's question of how arbitrary symbols acquire intrinsic meaning through sensorimotor grounding. Because the abstract and conclusion make claims about these specific 'long-standing theoretical challenges,' the construct validity gap is central, not cosmetic.","section":"3.2"},{"comment":"The per-model sample is five trials, with no formal hypothesis tests (Section 3.5 states this explicitly), so the 'consistently achieved high scores' claim rests on descriptive statistics over five subjective ratings. Moreover, several models produced identical outputs across all five trials, yielding SD=0 (e.g., Phi 3 and TinyLlama in both tasks); the paper itself notes in Section 5.1 that this reflects output determinism rather than stable quality. The conclusion should be tempered accordingly, and the current data cannot support statements about models 'acquiring capacities' without larger samples and independent validation.","section":"3.5 and Tables 2.a/2.b"}],"minor_comments":[{"comment":"The opening sentence reads 'a total of 12 different 13-condition models'; this should be rephrased, e.g., '12 distinct model families under 13 conditions.'","section":"3.1"},{"comment":"There is a typo in the abstract: 're vitalized' should be one word, 'revitalized.'","section":"Abstract"},{"comment":"The text refers to the unknown object as 'kluben,' but Appendix B.2.9 uses 'Kurben'; unify the spelling throughout.","section":"Appendix B.2.9"},{"comment":"Reference [7] is incomplete and garbled ('Palminteri S Yax N., Anlló H. ... page 51, 2024'), and reference [10] also contains a formatting error ('M. shanahan, solving the frame problem'); full author, title, venue, and year information should be provided.","section":"References"},{"comment":"Closed models had no output token limit while open-source models were capped at 1000 tokens; this is a potential confound for comparing response depth and should be justified or aligned.","section":"3.3"}],"recommendation":"reject","confidential_remarks":"The paper's transparency is commendable, but the scoring procedure is circular (the rater is itself a subject) and the tasks do not convincingly instantiate the named philosophical problems. These are not local presentation issues; they undermine the central claim as it is stated. I would not recommend resubmission unless the benchmark is redesigned and validated with independent human ratings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Shoko Oka's paper is a small empirical benchmark, not a theoretical contribution. The new thing is a score table: 13 models, two ad-hoc prompts, five zero-shot trials each, scored by ChatGPT-4o on six author-designed criteria. The appendix contains all 130 responses and the per-criterion scores, plus the Python script. That transparency is genuinely good and rare in this kind of write-up. The paper also explicitly says the statistics are exploratory (n=5, no formal tests) and concedes the single rater's validity is unverified. Credit where it's due: the author is honest about the instrument.\n\nThe soft spots are real. The evaluator is ChatGPT-4o, which is also one of the 13 subjects. The stress-test worry is that the ratings reflect the rater's stylistic preferences rather than any external standard. That's a fair concern, but note the rater gives ChatGPT o3 higher scores than itself on both tasks, so crude self-favoritism isn't happening. The deeper problem is the lack of any human inter-rater check: we don't know whether ChatGPT-4o's notion of 'creativity' or 'contextual fit' matches a human reader's. The tasks themselves are also loose proxies—a route-planning scenario and an invented-word metaphor exercise are not the frame problem or symbol grounding as McCarthy or Harnad formulated them. So the result is a descriptive comparison of how some LLMs handle two English-language prompts under zero-shot conditions. That is a fine thing to report. What the paper cannot support is the abstract's concluding sentence about 'acquiring capacities sufficient to address these challenges.' That goes beyond the evidence.\n\nGiven the paper's own limitations section, the author knows this. The text is written in a way that sometimes overstates—'high cognitive abilities,' 'rudimentary evidence of internal construction'—but the load-bearing data is reported in full, so a reader can re-weight the conclusions. I would not call this a fatal circularity because the scores are internally consistent and the rater's preferences are visible in the appendix. Still, the central interpretation is weak.\n\nWho is this for? Someone collecting zero-shot evaluation data points on small versus large models, or someone interested in pitfalls of LLM-as-judge. It won't change your view of the frame problem. I would send it to peer review—the methodology needs scrutiny and an editor could invite a revision that adds a human-rated subsample and tones down the claims. A desk reject would be too harsh.","headline":"A transparent, honestly limited benchmark whose single-LLM judge undermines the grander claims, but not a fatal flaw.","tokens_in":26611,"tokens_out":2841,"would_cite":false,"duration_ms":35212,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChatGPT o3 averages 53.6/60 on the frame problem and 56.6/60 on symbol grounding in a zero-shot benchmark.","keywords":["large language models","frame problem","symbol grounding","zero-shot benchmark","ChatGPT o3","cognitive evaluation","kluben","instruction tuning"],"falsifier":"Have human experts in cognitive science and philosophy blind-score the same 130 responses on the same six criteria; if their rankings do not reproduce the reported gap (o3 near 55, 1B open models near zero), or if human scores diverge from ChatGPT-4o's by more than a few points per model, the central claim that select LLMs have these capacities is unsupported.","tokens_in":25689,"feed_emoji":"🧠","tokens_out":5751,"duration_ms":65844,"temperature":0.7,"pith_summary":"The paper converts two classical AI challenges—the frame problem (deciding which facts matter as a situation changes) and the symbol grounding problem (how arbitrary symbols acquire meaning)—into two zero-shot English prompts and runs them five times each on 13 large language models. Its central claim is that several closed models, especially ChatGPT o3, return meaningful and stable responses to both challenges, with o3 averaging 53.6/60 on the frame task and 56.6/60 on the grounding task, while most small open-source models score near zero (Llama 3.2 1B averages 2.6 and 3.8). If true, this matters because it turns long-standing philosophical debates into measurable behaviour and suggests that some modern LLMs have started to acquire the capacities these problems were thought to demand.","feed_headline":"ChatGPT o3 tops zero-shot frame and symbol grounding tests","feed_subtitle":"Closed models score up to 56.6/60 on classic AI challenges; small open models score near zero.","key_machinery":"The carrying mechanism is a matched pair of zero-shot prompts: one compresses the frame problem into ten simultaneous urban events plus an initial route request followed by closure and multi-destination updates, and the other compresses symbol grounding into three questions about an invented object 'kluben' with the properties warm, soft, elastic, and light-absorbing. Each model answers each prompt five times from fresh sessions; a single rater LLM (ChatGPT-4o) then scores every response on six 10-point criteria (information selection, situation modeling, logic, adaptability, optimization, clarity for the frame task; understanding, introspection, creativity, logic, applicability, expression for the grounding task), and the paper uses the resulting means and standard deviations for all comparisons.","core_discovery":"The paper claims that modern LLMs can engage productively with the two problems that classical symbolic AI could not solve. When the frame problem is posed as a route-choice task with ten competing urban events plus road-closure and destination updates, and the symbol grounding problem is posed as meaning-construction for an invented object called 'kluben', ChatGPT o3 scores 53.6 and 56.6 out of 60 respectively, with small standard deviations across five trials; ChatGPT 4o, Claude 3.7 Sonnet, and Gemini 2.0 Flash also score in the 41–54 range. The same prompts reduce most small open-source models to near-zero scores, non-responses, or repetitive loops. The author reads this gap as evidence that select closed models may be acquiring capacities sufficient for meaningful and stable responses to these long-standing theoretical challenges, not as proof that either problem is solved.","pith_inferences":["Because ChatGPT-4o both rates and is rated, a human-blind re-scoring study is the natural next test; if the o3 lead shrinks under human raters, the reported effect is partly a rater-family effect rather than a task ability.","The qualitative finding that all capable models do better on 'kluben' than on route updating suggests a testable extension: adding physically inconsistent updates to the frame prompt (e.g., a road that blocks both travel and light) should disproportionately lower closed-model scores if their strengths are linguistic association rather than causal simulation.","A further extension would validate each 10-point scale item by item against expert human judgments, turning the six original criteria into a reusable instrument for future LLM cognition benchmarks."],"forward_implications":["ChatGPT o3's 53.6/60 and 56.6/60 give a concrete, reproducible reference point for measuring future models against classical AI problems.","Instruction tuning, not parameter count alone, is what lifts small models: Llama 3B Instruct reaches 27.8/60 on frame and 50.2/60 on grounding while the 3B base stays at 5.4 and 10.4.","8-bit quantization of Phi-3 costs a few points (frame 28 to 19, grounding 46 to 41) while preserving exactly deterministic outputs across trials.","Across all model tiers, symbol-grounding style meaning construction scores higher than frame-style situational updating, suggesting the two capacities are unevenly distributed."],"supporting_citations":[{"why":"Defines the frame problem as the challenge of efficiently handling facts that do not change in a dynamic world, which the frame-task prompt is built to test.","marker":"[1]"},{"why":"Reframes the frame problem as common-sense selection of what to ignore in a robot, motivating the ten-event urban scenario and information-filtering criteria.","marker":"[2]"},{"why":"Defines the symbol grounding problem that the 'kluben' prompt operationalizes as meaning construction for an unknown symbol.","marker":"[3]"},{"why":"Provides prior evidence that frontier LLMs show broad adaptive abilities, setting the expectation that closed models might score high.","marker":"[4]"},{"why":"Supports the paper's premise that advanced capacities such as theory of mind can emerge from language training, making meaningful responses to these problems plausible.","marker":"[6]"}],"fun_headline_variants":["o3 scores 56.6/60 on symbol grounding; open models near zero","Closed LLMs tackle frame and grounding benchmarks; open ones collapse","Zero-shot: o3 aces frame and symbol grounding, small open models fail","LLMs surprise: o3 high on frame and grounding; open models near zero","Frame and symbol grounding: o3 leads, open-source LLMs lag"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that ChatGPT-4o's subjective ratings actually measure the cognitive capacities under study; that premise is fragile because the rater is itself one of the rated models and no human inter-rater reliability has been established.","fun_headline_variants_meta":{"raw":{"variants":["o3 scores 56.6/60 on symbol grounding; open models near zero","Closed LLMs tackle frame and grounding benchmarks; open ones collapse","Zero-shot: o3 aces frame and symbol grounding, small open models fail","LLMs surprise: o3 high on frame and grounding; open models near zero","Frame and symbol grounding: o3 leads, open-source LLMs lag"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001099,"raw_usage":{"total_tokens":4562,"prompt_tokens":896,"completion_tokens":3666,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":3565}},"tokens_in":512,"tokens_out":3666,"duration_ms":29548,"temperature":1.0,"reasoning_tokens":3565,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:22:57.724932+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human experts in cognitive science and philosophy blind-score the same 130 responses on the same six criteria; if their rankings do not reproduce the reported gap (o3 near 55, 1B open models near zero), or if human scores diverge from ChatGPT-4o's by more than a few points per model, the central claim that select LLMs have these capacities is unsupported.","supporting_citations":[{"cited_title":"Some philosophical problems f rom the standpoint of artiﬁcial intelligence","cited_arxiv_id":null,"evidence_quote":"Defines the frame problem as the challenge of efficiently handling facts that do not change in a dynamic world, which the frame-task prompt is built to test."},{"cited_title":"Cognitive wheels: The frame problem of ai","cited_arxiv_id":null,"evidence_quote":"Reframes the frame problem as common-sense selection of what to ignore in a robot, motivating the ten-event urban scenario and information-filtering criteria."},{"cited_title":"The symbol grounding problem","cited_arxiv_id":null,"evidence_quote":"Defines the symbol grounding problem that the 'kluben' prompt operationalizes as meaning construction for an unknown symbol."},{"cited_title":"Sparks of artiﬁ cial general intelligence: Early experiments with gpt-4","cited_arxiv_id":null,"evidence_quote":"Provides prior evidence that frontier LLMs show broad adaptive abilities, setting the expectation that closed models might score high."},{"cited_title":"Evaluating large language models in theory of mind t asks","cited_arxiv_id":null,"evidence_quote":"Supports the paper's premise that advanced capacities such as theory of mind can emerge from language training, making meaningful responses to these problems plausible."}],"review_version":1}