{"id":"e8c1f4e0-7bf7-4504-942a-eb0ac9f25ebf","arxiv_id":"2607.26937","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Grouping multi-view code evidence by object identity plus one task-aware companion and budgeted rendering nearly closes the RepoBench retrieval-to-context survival gap while cutting tokens ~36%.","lead":"VITAL-RAG is a context-allocation layer for coding agents that groups retrieved code fragments by canonical object, keeps at most one query-relevant local companion, and renders under token budgets. It raises evidence survival under a 4K budget on RepoBench and improves executable Pass@1 on RepoExec across three model backends.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The paired RepoBench gain is load-bearing but may over-attribute survival to token authority if the fixed portfolio already embeds Entity-First 4+1 structure.","rationale":"The reader correctly flags provenance fragility and artifact/baseline gaps and lands on CONDITIONAL with moderate confidence; that verdict direction is right. The more load-bearing threat to the strongest claim as written, however, is internal to the isolation design: Base in Table 3 is Entity-First 4+1, not unstructured multi-view retrieval, so the +24.08 Recall@4K / −35.63% tokens result primarily identifies budget-constrained rendering (Table 6a: No rendering vs Full) after discrete selective invariance is already applied. That does not invalidate the systems result—ablations still show both quotient and rendering matter when structure is removed—but it means the abstract/claim sentence can over-sell “VITAL-RAG vs conventional allocation” unless the fixed portfolio is truly method-agnostic. Downstream Pass@1/Token-F1 ordering across three backends remains supportive but softer because of “-style” baselines and no error bars. I therefore keep CONDITIONAL rather than REJECT or ACCEPT: accept-shaped if the authors restate the paired claim as rendering-under-Entity-First (or show the Raw-Surface freeze still yields a large jump) and release enough artifacts to re-run Table 3/6; the reader’s adapter concern stays a necessary stress test but is not the single hinge of the central quantitative claim.","tokens_in":13249,"tokens_out":753,"duration_ms":15525,"concrete_test":"Re-run the exact Table 3 paired protocol with the discrete portfolio frozen to Raw Surface / Exact Dedup top-5 (Table 2: R@4K ≈25%, no canonical quotient) and apply only VITAL-RAG’s budget rendering (Eq. 4–5). If Recall@4K stays near ~25–40% rather than approaching 63%, the headline survival claim depends on Entity-First already being in the portfolio and should be restated as a rendering gain on pre-quotiented evidence.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim rests on Table 3’s paired isolation: under a fixed discrete portfolio, VITAL-RAG lifts Recall@4K 39.59→63.67% while cutting tokens 35.63%. That base is Entity-First 4+1 (§3.1, §6.1–6.2), which already enforces four canonical objects plus one companion—the same discrete skeleton as VITAL-RAG (§5.2–5.3). Table 6(a) then shows “No rendering” = 39.59% and “Full” = 63.67% on that skeleton, so the headline delta is almost entirely continuous rendering (Eq. 4–5: per-object cap, conditional floor, query-centered windows). The paper’s narrative frames the contribution as selective invariance (object quotient + companion + budgets). If Entity-First already removed most authority multiplication, the large paired gain mainly validates budget-aware rendering of an already-quotiented list, not the full invariance-race story versus ordinary multi-view RAG. Downstream tables use “-style” baselines without the same paired freeze, so they cannot independently re-anchor the isolation. The path+name adapter concern is real but secondary: even with perfect keys, the causal attribution of the 24-point survival jump is only as clean as the claim that the fixed portfolio does not already bake in the discrete half of the method.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper identifies a retrieval-to-context failure in repository-level coding agents: multi-view fragments of the same code object can multiply context authority and crowd out independent evidence under a bounded input budget, even when retrieval ranking is correct. It frames the resulting tension as an invariance race (stable under redundant renderings, sensitive to task-relevant local semantics) and proposes VITAL-RAG, a retriever-agnostic allocation layer that (i) quotients views by canonical object identity from a provenance adapter, (ii) reserves at most one query-relevant provenance-local companion, and (iii) renders the portfolio under per-object and global token budgets, with an optional post-generation program-semantic transfer gate. Empirically, on 16,490 RepoBench tasks it reports a ~24.6-point gap between Recall@5 and Recall@4K under Entity-First 4+1; under a fixed discrete portfolio it raises Recall@4K from 39.59% to 63.67% while cutting mean evidence tokens by 35.63%; and across Gpt-5.4, Claude Sonnet 4.6, and Qwen3-8B it matches or beats recent baselines on RepoClassBench and obtains the highest raw Pass@1 on RepoExec.","tokens_in":13671,"tokens_out":1633,"duration_ms":51505,"significance":"If the result holds, the paper cleanly isolates a stage that repository RAG and coding-agent work have under-specified: not candidate discovery, but how retrieved multi-view evidence is converted into a bounded model input. The large-scale RepoBench survival gap (Table 1), the demonstration that exact dedup, canonical keys, file quotas, and MMR leave Recall@4K near 39% (Table 2), and the component ablations (Table 6a) are concrete contributions that other systems can measure against. Strengths include the paired isolation of token authority, cross-language consistency, multi-backend downstream checks, and explicit separation of optional transfer from context construction. The work is significant as systems evidence and as a design principle (selective invariance), not as a new learning method; its lasting value is the problem formulation and the measurable survival metric under a fixed budget.","major_comments":[{"comment":"Abstract and §6.2 headline the +24.08 Recall@4K gain (39.59→63.67) and 35.63% token cut against Entity-First 4+1. That baseline already enforces four canonical objects plus one companion (§3.1, Table 2), i.e., the same discrete skeleton as §§5.2–5.3. Table 6(a) confirms the paired jump is almost entirely continuous rendering (No rendering 39.59 vs Full 63.67), while quotienting mainly restores portfolio entry (No quotient R@5 46.54 vs Full 64.16). The body discloses the isolation (§6.1), but the abstract/intro read as if selective invariance as a whole closes a conventional multi-view RAG gap. Please restate the primary claim so Table 3 is explicitly “budgeted rendering on a fixed quotiented portfolio,” and give equal billing to the full pipeline vs raw multi-view (Table 2: 25.08→63.67) and to the discrete vs continuous split in Table 6(a).","section":"Abstract; §6.1–6.2; Table 3; Table 6(a)"},{"comment":"The invariance guarantee and measured gains rest on g_I(v)=(normpath(p(v)), casefold(n(v))) (§5.2). Limitations §7 notes richer indexes are possible, but the manuscript does not quantify mis-grouping: path+name collisions, overloaded names, renames, generated code, or multi-symbol regions. If distinct helpers share a key, companion selection and authority multiplicity (Eq. 1) are ill-defined; if true multi-views split keys, authority multiplication returns. A short error analysis or sensitivity check on the RepoBench subset (collision rate, manual audit of failures where R@5 holds but R@4K fails under Full) would make the load-bearing provenance axiom falsifiable rather than assumed.","section":"§5.2 Eq. (1)–(2); §7"},{"comment":"Downstream tables use CodeRAG/GraphCoder/RepoScope “-style” settings (§6.1) without a paired freeze of the candidate pool comparable to Table 3. Pass@1 and Token-F1 gains are therefore consistent with better allocation but do not independently re-isolate quotienting vs rendering vs retrieval differences. Either (i) report one paired downstream condition with a shared retrieved pool, or (ii) clearly limit causal claims for Tables 4–5 to “end-to-end competitiveness under a common evaluation harness,” and keep causal attribution of the survival jump to RepoBench Tables 3 and 6.","section":"§6.1; Tables 4–5"}],"minor_comments":[{"comment":"Figure 1 is helpful but the two nearly identical multi-view panels are hard to parse in text form; label the failure mode (eviction of Helper C) more explicitly in the caption.","section":"Figure 1"},{"comment":"Eq. (3)’s query relevance test T(q)∩T(v*_e)≠∅ is very coarse (identifier token overlap). A sentence on failure modes (polysemous tokens, comments-only matches) would help readers judge companion precision.","section":"§5.3 Eq. (3)"},{"comment":"Table 6(b) is labeled “on RepoExec” but reports recall-style metrics (All R, Strict R) that match the RepoBench survival language; clarify the task subset and metric definitions.","section":"Table 6(b)"},{"comment":"Panel 6(d) Pass@1 (48.36 Full) is not aligned with Table 5’s per-model Pass@1 (e.g., 57.75 / 65.63 / 21.69). State which backend and whether transfer is included in Table 5.","section":"Table 5; Table 6(d)"},{"comment":"Minor typos/spacing: “VITALRAG” vs “VITAL-RAG” in the abstract; “aninvariance race”, “redundantrenderings”, and similar missing spaces appear in the front matter and §1.","section":"Abstract; §1"},{"comment":"Cite or briefly contrast chunk-level dedup and reranking in agent frameworks (e.g., SWE-bench agent stacks) so the “allocation as a distinct stage” claim is positioned against practice, not only academic RAG selectors.","section":"§2.2"}],"recommendation":"minor_revision","confidential_remarks":"The core empirical story (survival gap + rendering isolation + multi-backend competitiveness) is publishable and above the bar for a solid systems contribution. The main risk is over-attribution in the abstract: the +24 point number is real but is rendering-on-Entity-First, not full method vs naive RAG. If the authors fix framing and add a light provenance sanity check, I would accept; I would not demand new benchmarks. Fit is appropriate for cs.SE / empirical SE venues that take coding agents seriously."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the measurement, not the branding. On ~16.5k RepoBench tasks they show a ~25-point drop from portfolio entry (Recall@5) to full-snippet survival under 4K, it holds in both Java and Python, and the usual fixes—exact dedup, canonical-5, file quota, MMR, Entity-First—barely move Recall@4K. That isolation is clean and useful. Anyone building repo RAG should internalize it.\n\nWhat is actually new is treating allocation after retrieval as its own stage: object-level authority (no multi-view slot multiplication), one query-local companion, then per-object/global token budgets with query-centered windows. Ablations separate the pieces honestly—quotient helps entry, rendering closes the survival gap, 4+1 beats 5+0 and 3+2 on the companion boundary. Downstream, across Gpt-5.4, Claude Sonnet 4.6, and Qwen3-8B, they match or beat recent baselines on RepoClassBench and take raw Pass@1 on RepoExec. Math is light bookkeeping, not theater; citations cover the right retrieval/agent lines; circularity is low.\n\nSoft spot, in proportion: the stress note is partly right. Table 3 freezes a discrete portfolio that is already Entity-First 4+1, so the 39.59→63.67 jump with 36% fewer tokens is almost entirely continuous rendering (caps, floors, windows), not a full bake-off of selective invariance against raw multi-view RAG. Table 6 still shows you need both halves; the narrative just oversells unity. Secondary: path+name provenance can collide; baselines are “-style”; no error bars or released stack; K/B/caps are hand-set; the optional program-transfer gate is a side quest.\n\nFor people who ship coding agents or repo completion, this is worth a careful read—especially the gap tables and the rendering rule. I would send it to referees. Not a theory paper; a clear empirical systems note with a real failure mode and a practical fix.","headline":"Solid systems paper: large, well-measured post-retrieval survival gap and a simple allocator; the headline paired gain is mostly budgeted rendering on an already-quotiented portfolio.","tokens_in":14289,"tokens_out":543,"would_cite":true,"duration_ms":20429,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Coding agents often retrieve the right code and still lose it: redundant views of one object crowd out independent evidence under the token budget.","keywords":["retrieval-augmented generation","coding agents","context allocation","repository-level code generation","selective invariance","evidence survival","token budget","multi-view retrieval"],"falsifier":"Hold the discrete candidate portfolio fixed and re-measure Recall@4K: if VITAL-RAG’s rendering no longer lifts complete labeled-snippet survival from about 40% toward the portfolio-entry rate, or if swapping the path-name identity key for random or colliding keys erases the gain, the central allocation claim fails.","tokens_in":14100,"feed_emoji":"🧩","tokens_out":953,"duration_ms":24300,"temperature":0.7,"pith_summary":"When a coding agent pulls evidence from a whole repository, only a small slice fits in the model input. The paper shows that a large share of failure happens after retrieval succeeds: several renderings of the same function—body, signature, call site—each claim a context slot, so needed helpers get clipped even when they were ranked correctly. The authors call the design tension an invariance race: allocation should ignore redundant views of one object, yet still admit a nearby fragment when it adds task-relevant meaning. VITAL-RAG sits between retriever and generator, groups fragments by canonical code object, keeps at most one query-relevant companion from the same source region, and renders the portfolio under per-object and global token caps. On a large repository completion benchmark it roughly closes the gap between “found in the top five” and “still fully visible at 4K tokens,” while using substantially fewer evidence tokens, and the same allocation improves class reconstruction and end-to-end test pass rates across three model backends.","feed_headline":"Retrieved code still vanishes: redundant views steal the budget","feed_subtitle":"Grouping by code object and capping tokens nearly closes the 4K survival gap while cutting context size.","key_machinery":"Selective invariance, implemented as VITAL-RAG: provenance quotienting of multi-view fragments into canonical code objects, query-conditioned refinement that admits one distinct companion only when it is provenance-local and token-overlap relevant to the query, and budget-constrained rendering that assigns each selected object a capped token share before ordered concatenation into the model input.","core_discovery":"After correct retrieval, repository coding agents still lose labeled evidence during bounded input construction because multiple surface views of one code object multiply its context authority and displace independent objects. Selective invariance—one external slot per canonical object, plus at most one query-relevant local companion, under per-object and global token budgets—preserves far more complete evidence at 4K tokens with fewer tokens, and that survival gain carries through to class-level reconstruction and executable Pass@1.","pith_inferences":["Any multi-view retrieval stack—graphs, sliding windows, signatures plus bodies—likely needs the same object-level authority rule, not only code RAG.","Richer symbol or AST identity should strengthen the method where path-name collisions or renames are common, without changing the allocation principle.","Agent loops that re-retrieve after edits will keep reintroducing view multiplicity unless allocation stays between every retrieve and generate step.","The same invariance-versus-local-detail race may appear in long-context document agents that store many excerpts of one section."],"forward_implications":["Raising retrieval recall alone is not enough for repository agents; the pipeline needs an explicit post-retrieval allocation stage.","Exact text dedup, file quotas, and generic diversity reranking do not solve multi-view authority multiplication of the same code object.","Under a fixed top-k object set, smarter token shares can raise usable evidence while cutting mean context length.","One reserved companion slot is the practical balance: zero misses local helpers; two starts to trade away global coverage.","Shorter, object-structured context can improve executable Pass@1 rather than only similarity metrics."],"fun_headline_variants":["Redundant code views crowd out evidence after retrieval","One slot per object stops views from stealing the 4K budget","Selective invariance lifts Recall@4K and cuts evidence tokens","Canonical objects plus one companion preserve task-relevant code","VITAL-RAG keeps more labeled evidence inside tight context limits"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The system’s object identity key—basically normalized file path plus identifier name—is stable and fine-grained enough that true multi-view fragments share one slot while distinct helpers stay separate.","fun_headline_variants_meta":{"raw":{"variants":["Redundant code views crowd out evidence after retrieval","One slot per object stops views from stealing the 4K budget","Selective invariance lifts Recall@4K and cuts evidence tokens","Canonical objects plus one companion preserve task-relevant code","VITAL-RAG keeps more labeled evidence inside tight context limits"]},"model":"grok-4.5","effort":"low","cost_usd":0.001869,"raw_usage":{"total_tokens":888,"prompt_tokens":756,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":18688000,"prompt_tokens_details":{"text_tokens":756,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":67,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":756,"tokens_out":65,"duration_ms":3789,"temperature":1.0,"reasoning_tokens":67,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T18:44:48.966608+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold the discrete candidate portfolio fixed and re-measure Recall@4K: if VITAL-RAG’s rendering no longer lifts complete labeled-snippet survival from about 40% toward the portfolio-entry rate, or if swapping the path-name identity key for random or colliding keys erases the gain, the central allocation claim fails.","supporting_citations":[],"review_version":2}