{"id":"a0c2740a-71a7-497f-b81c-316630d8003e","arxiv_id":"2607.04926","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"In information-matched tiny transformers, zero-shot compositional binding fails for every route, while few-shot efficiency is governed by input-pathway sharing and code readability.","lead":"Tiny transformers never zero-shot bind held-out attributes even when every input route fully determines the answer; few-shot binding instead tracks pathway parameter sharing and code readability. The fully enumerable setup removes information and sampling confounds that usually muddy grounding and compositionality claims.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified to the central two-factor claim as scoped.","rationale":"The reader correctly flags the three-cell partial factorial and limited attribute generality as the softest point, yet the paper already scopes the claim to “within the tested cells,” “two pairwise dissociations,” and “not a full interaction estimate,” while treating the non-replicating weak-perceptual edge as testbed-specific. Those disclosures, together with the exact ceilings, exhaustive zero-variance evaluations, graded readability sweep (Spearman ρ=1.0 at fixed d), parameter-matched sharing control, causal interventions, and confirmatory + three-object replications, keep the two-factor account and the endpoint-invariance diagnostic secure as stated. No hidden inconsistency or untested assumption is required for the claim to hold inside its stated bounds; therefore the ACCEPT verdict and low correctness risk stand without adjustment. The concrete check simply verifies the already-released confirmatory numbers.","tokens_in":17474,"tokens_out":529,"duration_ms":21432,"concrete_test":"Re-aggregate the released confirmatory (bind:1) JSONL manifests and recompute the n=20 pooled Wilcoxon signed-rank p-values and Cliff’s δ for one-hot-shared vs. oracle and strong-entangled vs. oracle across the four doses; if they match Table 6 within bootstrap noise, the logged two-factor effects are confirmed as reported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is explicitly limited to pairwise dissociations within the tested cells (oracle vs. one-hot-shared isolating sharing at matched readability/params; weak vs. strong plus dimension-matched/graded sweeps isolating readability at fixed shared pathway and d), not a full factorial interaction. The missing fourth cell is acknowledged as having no natural construction (§9), the weak-perceptual ranking is flagged as testbed-specific (fails to carry to the three-object stress check), and replications cover two shape holdouts at n=20 plus the larger-world cores. Exact Bayes ceilings of 1.0, exhaustive evaluation, parameter-matched 16→24 pathways, causal input interventions, and full JSONL manifests make the within-scope evidence for sharing + practical readability (and for zero-shot failure under a lookup-sufficient objective) internally consistent and well-controlled. Disclosed limits on scale, synthetic codes, and attribute generality do not invert the reported effects or create an unaddressed confound.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies how information-matched input pathways (symbolic tokens, per-factor oracle embeddings, shared one-hot, and weak/strong entangled synthetic codes) affect compositional binding in ~6–10K-parameter transformers on fully enumerable factored worlds (128–729 states). Using exhaustive evaluation (zero sampling variance), exact Bayes ceilings of 1.0 on informative routes, 10–20 seeds with paired Wilcoxon/Holm–Bonferroni tests, and an exploratory/confirmatory split, it reports four findings: (1) endpoint invariance—no informative route achieves converged zero-shot binding on held-out query types despite ceiling 1.0; (2) a two-factor account of few-shot efficiency driven by input-pathway parameter sharing and practical code readability (supported by a three-cell partial factorial, dimension-matched and graded readability controls); (3) a double dissociation in which early-training transient zero-shot transfer tracks code format while few-shot efficiency tracks pathway sharing; (4) failure anatomy separating representation death (symbolic), systematic mis-binding with intact residual decodability (index routes, confirmed by causal input intervention), and inherited input readability (strong-entangled). The central positive claim is the two-factor few-shot account; endpoint and anatomy results are diagnostic constraints. Full code, manifests, and per-seed logs are released.","tokens_in":17792,"tokens_out":1326,"duration_ms":31989,"significance":"If the results hold as scoped, the paper makes a high-value methodological and mechanistic contribution. Information-matched routes with exact injectivity/Bayes ceilings, exhaustive evaluation, and parameter-matched shared pathways cleanly separate pathway sharing from code format and readability from input dimension—confounds that modality and grounding comparisons ordinarily leave entangled. The work isolates known ingredients (parameter sharing aids systematic generalization; disentangled codes are not sufficient) under unusually tight controls and adds a graded readability result, a double dissociation between trajectory and few-shot behavior, and causal input-intervention evidence for mis-binding. Full JSONL manifests and exact reproduction are genuine strengths. The scale is deliberately tiny; the paper treats this as a feature for classification of failures (information vs inductive bias) rather than a claim about large models. Within that regime the two-factor account is a clear, falsifiable isolation that future work on adapters, prompts, and learned perception can test.","major_comments":[{"comment":"Abstract and §1 Contributions vs §9: the abstract states that few-shot sample efficiency is “best explained by” pathway sharing and readability, while §9 carefully frames a three-cell partial factorial with two pairwise dissociations (“within the tested cells”), notes that the fourth cell has no natural construction, and flags the weak-perceptual edge over the oracle as testbed-specific. Align the abstract and contribution bullets with the body’s three-cell language so the central claim is not over-read as a full interaction estimate or as ranking among all readable shared codes.","section":"Abstract / §9"},{"comment":"§9 replications and “portable factors”: both confirmatory held-out types are shape queries (bind:3 exploratory, bind:1 confirmatory); the three-object stress check corroborates sharing and strong-entangled collapse but not the weak-perceptual ranking. Limitations already list attribute generality as future work, but the Discussion’s “portable factors” phrasing and the unqualified two-factor framing in the abstract should explicitly bound portability to shape-binding under the tested routes until attribute- and modality-general replications exist. This is a claim-scope issue, not a request for new experiments before acceptance.","section":"§9 / Discussion / Limitations"}],"minor_comments":[{"comment":"Table 2 and Figure 1: the strong-entangled peak CI overlaps chance; the text already notes this, but the figure caption and the double-dissociation summary in §7 could state more prominently that the clear above-chance transient is for symbolic and weak-perceptual only, so the format grouping is not uniform across all distributed codes.","section":"§7 / Table 2 / Figure 1"},{"comment":"Table 6: confirmatory text_only vs oracle is p=0.053 (directional after Holm). The body is careful; ensure the abstract and ordering language (“shared readable pathways > modular oracle”) do not imply a confirmed symbolic advantage on the second holdout.","section":"§9 / Table 6"},{"comment":"§6 transition holdout: the authors correctly caution that the tiny world may not reward the successor rule and treat the result as weaker supporting evidence. Consider moving the transition numbers fully to an appendix or a single sentence so they do not dilute the binding-focused endpoint claim.","section":"§6"},{"comment":"Table 1 / §3: “perceptual” is used descriptively for fixed synthetic tanh mixes. A one-sentence reminder in the table caption that these are not learned or naturalistic perception would reduce misreading by readers skimming only the routes table.","section":"Table 1"},{"comment":"§8 / Table 4: the causal intervention is a strong addition. Briefly note in the table caption that each entry is a fraction of (background, value) cases so the metric is self-contained without the main text.","section":"§8 / Table 4"},{"comment":"Reproducibility Statement and Appendix B: the one-page experimental specification (Table 7) is excellent. Consider adding the exact AdamW β, batch size, and evaluation cadence already in Appendix D into Table 7 so the single-page spec is fully self-contained.","section":"Appendix B / Table 7"}],"recommendation":"minor_revision","confidential_remarks":"Strong methods paper with unusually tight controls and full release artifacts; good fit for a journal that values careful small-scale mechanistic work. The central claim is scoped honestly enough that minor wording alignment should suffice—no load-bearing experimental gap. I would not require new attribute holdouts for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: in these fully enumerable MicroGround worlds, no information-matched route (symbolic, oracle, shared one-hot, weak or strong perceptual) reaches converged zero-shot compositional binding—all sit at or below chance despite Bayes ceilings of 1.0—while few-shot sample efficiency tracks two factors the authors isolate with parameter-matched controls: input-pathway sharing and practical readability of the code. The clean per-factor oracle is not the best readable route; shared projections transfer better.\n\nWhat is actually new is the confound removal. Parameter sharing aiding systematic generalization and disentangled codes not being sufficient are already known (Csordás, Montero). Here the routes are information-matched with exact injectivity ceilings, evaluation is exhaustive (zero sampling variance), pathways are parameter-matched (identical 16→24 linear for one-hot-shared and weak-perceptual), and they run a graded readability sweep at fixed dimension plus a dimension-matched low-readability control. The three-cell partial factorial cleanly separates sharing from format and readability from dimension. Failure anatomy is also useful: symbolic loses the answer at readout; index routes keep it decodable but causally track the wrong slot under input intervention; entangled routes inherit input readability. Full JSONL manifests and code are released.\n\nSoft spots are real but disclosed and proportionate. Worlds are tiny (128–729 states), models ~6–10K params, codes are fixed synthetic mixes not learned perception, and the fourth factorial cell (distributed through modular pathway) has no natural construction. The weak-perceptual edge over the oracle is testbed-specific and fails to carry to the three-object stress check; only the two core effects (sharing helps, low readability hurts) do. Endpoint invariance is correctly scoped as inductive bias under a lookup-sufficient objective, not a universal claim. Citation pattern is honest about priors.\n\nThis is for people who care about compositional generalization, binding, or how to control grounding-style claims. The math is ordinary experimental ML; the data and stats (10–20 seeds, paired Wilcoxon, Holm–Bonferroni, confirmatory holdout type) look solid. I would send it to peer review. Engage with it if you work on these questions; the methodological standard is worth adopting even if the scale limits the reach of the claims.","headline":"Careful tiny-model study that cleanly isolates pathway sharing and code readability as drivers of few-shot binding, with exhaustive eval and exact ceilings; zero-shot fails everywhere under a lookup-sufficient objective.","tokens_in":18372,"tokens_out":578,"would_cite":true,"duration_ms":5315,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"No input route yields zero-shot compositional binding in tiny transformers; few-shot efficiency tracks pathway sharing and code readability.","keywords":["transformers","compositional binding","few-shot learning","input pathways","parameter sharing","code readability","symbol grounding","enumerable testbeds"],"falsifier":"Construct a natural distributed code routed through a modular pathway that outperforms the shared one-hot route on few-shot binding, or increase the number of held-out query types until pure lookup is costly and observe any informative route escape chance-level zero-shot composition within the same capacity and training sweep.","tokens_in":18338,"feed_emoji":"🧩","tokens_out":1052,"duration_ms":18064,"temperature":0.7,"pith_summary":"This paper asks how the form of an input pathway—symbolic tokens, a clean per-factor oracle code, or an entangled perceptual vector—affects whether a small transformer can bind that information compositionally. It studies ~6–10K-parameter models on fully enumerable factored worlds so every measurement covers the whole input space and every informative route is information-matched to an exact Bayes ceiling of 1.0. The central result is endpoint invariance: no informative route reaches converged zero-shot composition on held-out binding queries; all end at or below chance under a training objective for which lookup suffices. Once a few held-out examples are leaked, sample efficiency is best predicted by two factors: whether the input pathway shares parameters across query types, and how practically readable the code is. Distributed codes show a transient above-chance phase early in training while index-like codes do not, yet that format effect is dissociated from the sharing effect that governs few-shot gains. A sympathetic reader cares because the work supplies a controlled, fully measurable account of which mundane pathway properties actually move the needle once pure memorization is no longer enough.","feed_headline":"Tiny transformers never bind zero-shot on any route","feed_subtitle":"Few-shot gains track pathway sharing and code readability, not oracle cleanliness or entanglement","key_machinery":"The MicroGround testbed: finite factored worlds (128–729 states) realized under five information-matched input routes (symbolic, factored-oracle, one-hot-shared, weak- and strong-entangled perceptual), exact per-route Bayes ceilings of 1.0, and exhaustive evaluation of every query so behavioral measurements have zero sampling variance. This isolates pathway sharing and readability while classifying failures as inductive-bias rather than information-limited.","core_discovery":"Within information-matched routes into tiny transformers on exhaustively enumerable factored worlds, no informative route achieves converged zero-shot compositional binding—all end at or below chance despite exact Bayes ceilings of 1.0. Few-shot binding efficiency is best explained by a two-factor account: input-pathway parameter sharing (a shared projection transfers to unseen query types; private per-factor tables do not) and practical readability of the code (poorly readable entangled codes saturate far below readable alternatives). The clean per-factor oracle is not the most sample-efficient readable route; shared readable pathways transfer better.","pith_inferences":["Making pure lookup costly by holding out more query types or enlarging the world may break endpoint invariance and let some routes escape memorization.","The same two factors—shared versus modular adapters and readability of continuous codes—may predict transfer of prompt- or adapter-style interventions in larger models.","A high-information code that is linearly unreadable yet nonlinearly recoverable would cleanly test whether practical readability or raw information content is the true gate.","Free-order symbolic descriptions that force genuine tag-based binding would tighten the comparison against absolute-position shortcuts."],"forward_implications":["Any claim that a non-symbolic or grounded side-channel induces zero-shot composition must survive exact information matching; at this scale and objective it does not.","Shared input projections should transfer better to unseen query types than modular per-factor embeddings, even when both codes are fully readable.","Practical readability of an input code, not mere injectivity or entanglement strength, gates how well a model binds through it.","Early-training above-chance transients track distributed versus index-like code format and do not predict few-shot efficiency.","Failure modes split by route: symbolic loses the answer at readout, index routes mis-bind while keeping the answer decodable, entangled routes inherit input readability."],"fun_headline_variants":["Tiny transformers fail zero-shot binding on every input route","No informative path yields zero-shot composition in tiny transformers","Few-shot binding tracks pathway sharing and code readability","Endpoint: all routes stay at or below chance for zero-shot binding","Input pathways drive few-shot—not zero-shot—binding in tiny models"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The three-cell partial factorial plus replications on two shape query types and a three-object stress check are enough to treat pathway sharing and readability as the dominant portable predictors, even though a full four-cell factorial cannot be built and one ranking is testbed-specific.","fun_headline_variants_meta":{"raw":{"variants":["Tiny transformers fail zero-shot binding on every input route","No informative path yields zero-shot composition in tiny transformers","Few-shot binding tracks pathway sharing and code readability","Endpoint: all routes stay at or below chance for zero-shot binding","Input pathways drive few-shot—not zero-shot—binding in tiny models"]},"model":"grok-4.5","effort":"low","cost_usd":0.00291,"raw_usage":{"total_tokens":1121,"prompt_tokens":906,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":29100000,"prompt_tokens_details":{"text_tokens":906,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":125,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":906,"tokens_out":90,"duration_ms":1661,"temperature":1.0,"reasoning_tokens":125,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T11:21:24.052081+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Construct a natural distributed code routed through a modular pathway that outperforms the shared one-hot route on few-shot binding, or increase the number of held-out query types until pure lookup is costly and observe any informative route escape chance-level zero-shot composition within the same capacity and training sweep.","supporting_citations":[],"review_version":1}