{"id":"ba388b62-8e9b-41c7-8696-5482c2297807","arxiv_id":"2608.09888","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A compact recurrent latent-reasoning model achieves 29.5% pass@2 on ARC-AGI-1 at a claimed cost of $0.0007 per task, purportedly the cheapest such result to date.","lead":"BDH-CQ is a 150M-parameter model that learns visual transformations from examples at inference time and solves them through iterative computation in a continuous latent space. The paper reports 29.5% pass@2 on ARC-AGI-1 at an estimated $0.0007 per task, claiming a new cost-efficiency frontier.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Paper's own §6.6 cost figures contradict the §5 headline: STANDARD cost for the 29.5% result is $0.00265/task, not $0.00070/task, so the claimed cost-efficiency breakthrough is unsupported until reconciled.","rationale":"The central claim is quantitative: a particular operating point on the ARC-AGI-1 cost-accuracy Pareto frontier at $0.00070/task. For this claim to be true, the cost number must be well-defined, measured consistently, and the same number must be used wherever the headline point appears. The paper violates this: §5 computes $0.00070 from 0.85 H200 GPU-seconds, but §6.6 reports $0.00265246 for the STANDARD setting that produces the same 118/400 result on the same full evaluation. Since both numbers are in the same paper and are meant to describe the same system, one of them is wrong or they measure different cost objects (e.g., model-only compute vs. full service). In either case the headline figure cannot be used as the basis for a Pareto-frontier claim without a reconciliation. The reader's concern about external comparability with leaderboard API prices is real, but the §5/§6.6 mismatch is stronger because it does not depend on external data. I still credit the paper for an extensive behavioral study and for reporting repeatability checks (identical re-runs) and the absence of a semantic-ID effect; those are useful and provide some internal consistency for the accuracy claim. I also note that the 'independent audit' is performed by co-authors, so it does not constitute fully independent verification, and that no code or weights are released. None of these are needed to establish the cost contradiction, which stands alone. The correct disposition remains conditional: the paper could be accepted after a unified cost accounting and a re-plotted frontier. If the true cost is $0.00265, the specific abstract claim is quantitatively wrong; if the authors can show $0.00070 is the correct full-pipeline cost and reconcile §6.6, the central claim could survive. Thus verdict CONDITIONAL, with partial agreement with the reader.","tokens_in":13612,"tokens_out":11281,"duration_ms":95430,"concrete_test":"Ask the authors to supply the complete cost ledger for the 400-task STANDARD run that produced 118/400 (29.5%): total H200 GPU-seconds, number of candidates generated, and all pipeline overhead (input encoding, candidate construction, ranking, two-attempt selection). Convert the metered total to dollars at $3/H200-hour and compare with §5's $0.00070 and §6.6's $0.00265246. If the true metered cost is $0.00265, recompute the '57× cheaper than GPT 5.6 Luna' comparison and re-plot Figure 2; if the point falls inside the existing frontier, the headline cost-efficiency claim is false as stated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing problem is an internal contradiction in the paper's own cost accounting, not just cross-system comparability. Section 5 reports the default operating point as 'approximately 0.85 H200 GPU-seconds per task', which at $3/H200-hour gives $0.00070/task, and this is the point plotted for the 29.5% pass@2 result. Section 6.6, describing the same full public ARC-AGI-1 evaluation, states that the MIN effort setting cost one third of STANDARD ($0.00088399 versus $0.00265246 per task) and scored 111/400 versus 118/400 pass@2. Since 118/400 is exactly the 29.5% headline result, the STANDARD setting in §6.6 is the same operating point, yet its reported cost is about 3.8× larger than §5's $0.00070. At $3/H200-hour, $0.00265 corresponds to about 3.18 H200 GPU-seconds per task, not 0.85. If §6.6's number is the true metered system cost, the abstract's 'less than one-tenth of a cent' is wrong, the stated 'approximately 57× cheaper than GPT 5.6 Luna (Low)' should be about 15×, and the Pareto-frontier plot must be redrawn at the corrected point. No line-item reconciliation is provided. The paper's caveat that other leaderboard costs 'may represent hardware estimates or API prices' makes the external comparison fragile, but the §5/§6.6 conflict is decisive internally: the quantity that defines the central claim is not stable within the manuscript. This concern is addressable—the authors could disclose a full cost model—but until then the cost-efficiency breakthrough is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BDH-CQ, a 150M-parameter system that combines in-context learning with recurrent latent reasoning for ARC-style visual tasks. Demonstrations are ingested sequentially into a recurrent memory, and the query is solved by iterating a latent workspace without decoding intermediate tokens. The authors report 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed $0.00070 per task, claim this point breaks the cost-accuracy Pareto frontier, and support the result with ConceptARC profiling, controlled post-freeze generalization experiments, and generated ladders. The paper is unusually transparent about many limitations, including a corrected 240-task rerun, contradictory generated tasks, and the statistical weakness of the effort-tier comparison.","tokens_in":14025,"tokens_out":7673,"duration_ms":64000,"significance":"If the reported cost figure can be reconciled and the system made independently verifiable, the result would be significant: a compact, cheap ARC-AGI-1 system with a high score-to-cost ratio, plus a behavioral methodology that distinguishes isolated correct outputs from consistent rule application. The controlled experiments use deterministic oracles and exact-output metrics, and they deliver falsifiable findings such as the demonstration-coverage effect on nesting and ordering. However, the central cost-efficiency claim is currently undermined by an internal contradiction between §5 and §6.6, and the claimed independent audit is conducted by co-authors. These issues must be resolved before the headline result can be accepted.","major_comments":[{"comment":"Section 5 states that the default 29.5% pass@2 result costs approximately 0.85 H200 GPU-seconds per task, giving $0.00070 at $3/hour. Section 6.6 describes the same full public ARC-AGI-1 evaluation under STANDARD effort, scoring 118/400 pass@2 (exactly 29.5%) at $0.00265246 per task, while MIN effort costs $0.00088399 and scores 111/400. These two cost figures for the same operating point differ by a factor of about 3.8; at $3/hour, $0.00265246 corresponds to 3.18 H200 GPU-seconds per task, not 0.85. Because the Pareto-frontier plot, the 'less than one-tenth of a cent' claim, and the 'approximately 57x cheaper than GPT 5.6 Luna (Low)' comparison are all based on the $0.00070 figure, the manuscript's central cost-efficiency result is not internally consistent. The authors must provide a line-item cost model and reconcile the two figures before the claim can be evaluated.","section":"§5 vs §6.6"},{"comment":"The paper's verification section is labeled 'Independent evaluation' but identifies the auditors as Remigiusz Kinas and Richard Zhong, both of whom are listed as co-authors of this manuscript on the title page. A verification performed by co-authors is not independent in the sense needed to support the claim that the deployed system's 29.5% score was reproduced by outside parties. This matters because the paper contains no publicly accessible audit report or URL, so the manuscript's only external check is an internal one. Please either have the audit performed by genuinely independent researchers who are not authors, or remove the 'independent' wording and describe the procedure as an internal black-box check.","section":"§5 (Independent evaluation) and title page"},{"comment":"Section 3.3 states that dimensions, exact update rules, and implementation details remain proprietary, Section 4.1 states that the complete internal training recipe remains proprietary, and no weights or evaluation code are provided. For an empirical systems claim with a measured cost value, this would be acceptable if every reported number were internally consistent and auditable; the unresolved §5/§6.6 discrepancy shows that this precondition is not met. Please disclose the cost model (GPU-seconds per candidate, number of candidates, batch effects, hardware assumption) and enough of the inference procedure to allow a third party to reproduce the measurement, or explicitly reframe the cost-efficiency claim as a self-reported figure.","section":"§3.3, §4.1, §5"},{"comment":"The Pareto-frontier claim assumes that BDH-CQ's computed hardware cost is commensurable with the leaderboard costs of other systems, which the manuscript itself notes may represent hardware estimates or API prices. If other points are API prices that include provider margins or different hardware, an apple-to-apples comparison is not established. Please provide a sensitivity analysis over the $3/hour assumption and, if possible, a comparison using a uniform cost metric (all hardware estimates or all API prices) before claiming a frontier breakthrough.","section":"§5, Figure 2"}],"minor_comments":[{"comment":"The text says propagation and copying remain correct on 48/48 held-out outputs, but Figure 5's caption says these families use 12 outputs per point; please state the number of plotted points explicitly so the reader can verify the denominator.","section":"§6.2, Figure 5"},{"comment":"The 'approximately 57x cheaper' and 'approximately 11x cheaper' comparisons to GPT 5.6 Luna (Low) do not show their arithmetic; in particular, the 11x figure appears to assume the 80% price reduction, but the revised Luna price is not stated. Please include the calculation explicitly.","section":"§5"},{"comment":"The passage notes that all 75 single-candidate records were correct at rank one, which is important context for pass@2; please also state how many tasks received two candidates and whether the reported cost per task includes the compute for one or two candidate generations.","section":"§6.6"},{"comment":"The corrected 240-task rerun is described in the text, but Table 8 and the surrounding discussion do not annotate which entries come from the corrected set versus the original generator run; please add explicit markers.","section":"Appendix A.3"}],"recommendation":"major_revision","confidential_remarks":"The internal cost contradiction in §5 versus §6.6 is the main technical blocker; if the authors reconcile the numbers and provide a cost model, the paper could be publishable. The 'independent audit' by co-authors should be reworded, and the lack of code or weights makes the empirical claims unusually hard to verify. There is also a scope question: the paper's headline is a benchmark cost-efficiency record, but the proprietary nature of the system means the record cannot be independently confirmed from the manuscript alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about BDH-CQ. The architectural idea is genuinely new: demonstrations write into a recurrent memory and the query is solved by latent-space iteration, with no verbalized steps, and it reaches 29.5% pass@2 on ARC-AGI-1 with a 150M model. The controlled behavioral study is careful and unusually honest. But the headline cost-efficiency claim does not survive contact with the paper's own numbers. Section 5 says the default point costs $0.00070 per task (0.85 H200 GPU-seconds at $3/hour). Section 6.6, describing the same full public ARC-AGI-1 evaluation, reports the STANDARD effort setting at $0.00265246 per task with 118/400 pass@2—exactly the headline result. That is roughly 3.8 times the Section 5 number. No reconciliation is given. The abstract's \"one-tenth of a cent\" is wrong at the reported standard cost, and the \"57x cheaper than GPT 5.6 Luna\" claim becomes about 15x. The Pareto-frontier point must be redrawn. This is not a minor footnote; the central practical claim rests on a cost number that is unstable within the manuscript.\n\nWhat does the paper do well? The evaluation design is a cut above typical ARC papers. The authors report Wilson intervals, distinguish test-pair from strict-task accuracy, rerun a corrected 240-task set, acknowledge contradictory generated tasks and overlapping intervals, and flag that final grids cannot distinguish an incomplete rule from a narrower rule. The composition and ordering experiments are interesting, and the intervention results (e.g., matched demonstrations eliminating depth-5 nesting failures) are credible. The related work is well placed, and the distinction from task-trained recursive solvers like HRM/TRM is legitimate.\n\nSoft spots beyond the cost problem: the \"independent audit\" is performed by listed co-authors, which is not independence. No weights or code are released, and the training recipe and architecture details are proprietary. The cost comparison uses a computed hardware estimate versus leaderboard costs that may be API prices; that alone would make the Pareto-frontier claim fragile, but the internal contradiction makes it unsupported.\n\nThis paper deserves a serious referee—the architecture and behavioral methodology are worth engaging with—but the cost claim needs a full line-item reconciliation and ideally an external audit before the headline can be accepted. If the standard cost is genuinely $0.00265, the result is still cheap but not the claimed frontier breakthrough. The authors can likely fix this; the rest of the paper is worth reading in the meantime.","headline":"Solid behavioral study undermined by an internal cost-accounting contradiction that invalidates the headline Pareto-frontier claim until reconciled.","tokens_in":14554,"tokens_out":2395,"would_cite":false,"duration_ms":19892,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BDH-CQ claims that in-context learning and recurrent latent reasoning combine to reach 29.5% pass@2 on ARC-AGI-1 at $0.00070 per task, establishing a new state of the art in benchmark cost efficiency.","keywords":["in-context learning","recurrent latent reasoning","ARC-AGI-1","cost-accuracy frontier","continuous thought","BDH architecture","visual reasoning","skill acquisition"],"falsifier":"Re-measure BDH-CQ's per-task cost on a metered cloud GPU with all overheads (startup, batching, energy, candidate ranking) and compare it with the reported costs of the systems on the leaderboard; if the measured cost at 29.5% pass@2 is not below the cost of every system at equal or higher accuracy, the claimed Pareto-frontier breakthrough is unsupported. A second decisive check would be an independent run that verifies the 118/400 pass@2 score under the stated protocol.","tokens_in":13444,"feed_emoji":"🧩","tokens_out":4686,"duration_ms":36380,"temperature":0.7,"pith_summary":"The paper introduces BDH-CQ, a reasoning system that performs in-context learning through recurrent memory and solves tasks by iterating over a continuous latent workspace instead of verbalizing intermediate steps. The central claim is that this combination lets a 150M-parameter model reach 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.00070 per task, a point the paper argues breaks the previously reported cost-accuracy Pareto frontier. The authors also use ConceptARC and controlled ARC-like tasks to map what the model can bind from demonstrations, how consistently it applies an inferred transformation, and where it fails. A sympathetic reader should care because the result suggests that nontrivial abstract reasoning can be achieved without token-by-token narration, and at a small fraction of the cost of verbal-reasoning systems.","feed_headline":"150M-parameter model breaks ARC-AGI-1 cost-accuracy frontier","feed_subtitle":"Recurrent latent reasoning solves visual tasks from demonstrations alone at under one-tenth of a cent each.","key_machinery":"The central mechanism is the separation of contextual memory from the reasoning workspace: demonstrations continuously update a recurrent latent state, and the query is answered by iterative computation in a high-dimensional latent workspace using BDH layers (ReLU-low-rank transformations combined with linear attention). The update equations $S_t = U_\\theta(S_{t-1}, D_t)$ and $H_{r+1} = F_\\theta(H_r, S_K)$ define the system-level interface that the paper studies; the memory carries what the demonstrations specify, while the workspace carries the ongoing computation for the current query.","core_discovery":"BDH-CQ processes each demonstration sequentially, updating a recurrent memory $S_t = U_\\theta(S_{t-1}, D_t)$, then encodes the query and iterates a latent workspace $H_{r+1} = F_\\theta(H_r, S_K)$ before decoding an answer. The paper's central discovery is that this design acquires unseen visual transformations purely from demonstrations, with no parameter updates and no task-specific identity, and does so efficiently enough to establish a new state of the art in ARC-AGI-1 cost efficiency: 29.5% pass@2 at $0.00070 per task, about 57x cheaper than the comparable leaderboard system GPT 5.6 Luna (Low). The behavioral studies show the model binds dense color mappings, extrapolates boundary propagation and copying across the tested ranges, but shows structured limits on ordering long sequences and deep nesting, with composition success depending on the operation.","pith_inferences":["If the computed cost is genuinely commensurable with leaderboard hardware and API prices, then ARC-AGI-1's score-cost plane becomes accessible to very small recurrent models, which would shift the competition toward cost-aware benchmarks rather than raw accuracy.","The strong in-context binding of dense mappings suggests that the same recurrent memory could be trained to condition on textual demonstrations for language and math tasks, a direction the paper only outlines.","The consistency gap between pair and task accuracy implies that pass@2 as usually reported may overstate the rate of rule induction; reporting whole-task solve rates would make capability claims more comparable across systems.","The composition results (rotation composes with relocation 72/72, reflection 47/72, color swap 0/72) indicate that compositionality is not a single ability but is mediated by the representation of the operations being composed, which can be tested by varying motif families and operation pairs."],"forward_implications":["ARC-AGI-1 cost efficiency gains a new frontier point: a 150M-parameter system at 29.5% pass@2 and $0.00070 per task, roughly 57x cheaper than GPT 5.6 Luna (Low) at 34.2% and $0.040.","Increasing the latent reasoning effort during inference raises pass@2 from 21% (LOW) to 27% (MEDIUM) to 29.5% (HIGH), with cost reductions of 22% and 11% respectively.","Demonstrations alone can bind dense task-specific mappings: a fresh color permutation is applied to all 96 held-out outputs at rank one, even with eight simultaneous bindings.","The model's consistency gap on ConceptARC (77.92% test-pair pass@2 vs 59.38% strict-task pass@2) shows that correct individual outputs do not always transfer to all test inputs of a task.","Ordering eight bars and nesting five containment relations expose distinct bottlenecks: ordering failures break the whole output structure, while nesting failures preserve structure and differ in a single containment decision; adding a matched demonstration largely removes the nesting cliff."],"supporting_citations":[{"why":"Defines ARC and the skill-acquisition framing the paper evaluates against.","marker":"Chollet, 2019"},{"why":"Provides the ARC-AGI-1 evaluation protocol and leaderboard conventions (pass@2).","marker":"Chollet et al., 2024"},{"why":"Introduces the BDH architecture whose layers BDH-CQ builds on.","marker":"Kosowski et al., 2025"},{"why":"Serves as the prior continuous-thought mechanism the paper contrasts with, providing hidden-state feedback.","marker":"Hao et al., 2024"},{"why":"Supplies the ConceptARC benchmark used for the concept-organized capability profile.","marker":"Moskvichev et al., 2023"},{"why":"Source of leaderboard cost and accuracy points that define the Pareto frontier being broken.","marker":"ARC Prize Foundation, 2026"},{"why":"Represents the task-trained recursive solver baseline whose transductive pipeline BDH-CQ targets.","marker":"Wang et al., 2025"}],"fun_headline_variants":["BDH-CQ: 29.5% on ARC-AGI-1 for $0.0007","Latent reasoning hits 29.5% on ARC-AGI-1 at $0.0007","150M model sets new ARC-AGI-1 cost-accuracy low","Recurrent latent reasoning: 29.5% ARC-AGI-1, $0.0007/task","BDH-CQ breaks ARC-AGI-1 cost frontier at $0.0007"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central cost-efficiency claim assumes that BDH-CQ's computed $0.00070 per task (0.85 H200 GPU-seconds at $3/hour) is directly comparable to the hardware estimates and API prices reported for other leaderboard systems, which the paper itself notes may not be commensurable.","fun_headline_variants_meta":{"raw":{"variants":["BDH-CQ: 29.5% on ARC-AGI-1 for $0.0007","Latent reasoning hits 29.5% on ARC-AGI-1 at $0.0007","150M model sets new ARC-AGI-1 cost-accuracy low","Recurrent latent reasoning: 29.5% ARC-AGI-1, $0.0007/task","BDH-CQ breaks ARC-AGI-1 cost frontier at $0.0007"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001662,"raw_usage":{"total_tokens":6559,"prompt_tokens":869,"completion_tokens":5690,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":5571}},"tokens_in":485,"tokens_out":5690,"duration_ms":34561,"temperature":1.0,"reasoning_tokens":5571,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:51:54.253330+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-measure BDH-CQ's per-task cost on a metered cloud GPU with all overheads (startup, batching, energy, candidate ranking) and compare it with the reported costs of the systems on the leaderboard; if the measured cost at 29.5% pass@2 is not below the cost of every system at equal or higher accuracy, the claimed Pareto-frontier breakthrough is unsupported. A second decisive check would be an independent run that verifies the 118/400 pass@2 score under the stated protocol.","supporting_citations":[],"review_version":1}