{"id":"9cb70442-48f9-43f9-883c-fd7db669778c","arxiv_id":"2605.28047","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Knot estimates QA model sensitivity to candidate knowledge via subset counterfactual training and latent factor coverage, yielding unit rankings that outperform baselines without extra model calls.","lead":"The paper introduces Knot, a structured rank-aware estimator that learns subset sensitivity from counterfactual supervision to score how much a black-box QA model depends on different knowledge units. A smart generalist might read it to see a practical way to flag unreliable LLM answers by tracking knowledge dependencies in noisy retrieval settings.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Subset-level counterfactual supervision generation may still require costly or exhaustive QA calls during training","rationale":"The reader's weakest_assumption directly identifies the same point about obtaining subset-level counterfactual supervision without exhaustive cost. Because the full text was not supplied in the initial read, the concern remains the primary open assumption; confirming or refuting the training-data generation procedure is the single check that would move the UNVERDICTED status.","tokens_in":1681,"tokens_out":376,"duration_ms":27547,"concrete_test":"In the methods/experiments section, locate the paragraph describing training-data construction for subset counterfactuals; count the typical number of knowledge units per instance and the sampling or oracle procedure used. If units > 8 and no non-exhaustive strategy (e.g., structured sampling or synthetic labels) is specified, regenerate the training set for one benchmark using exhaustive enumeration on a 6-unit subset and retrain; if Knot's subset-sensitivity correlation drops below the reported margin over baselines, the supervision assumption is load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Knot learning accurate subset sensitivity from counterfactual supervision that captures redundancy/substitutability/complementarity without exhaustive perturbation. The abstract states the model 'learns from subset-level counterfactual supervision' and delivers gains 'without extra QA-model calls,' but the latter qualifier applies only at inference. If training data construction itself enumerates or heavily samples subsets (exponential in candidate units) or relies on repeated black-box QA evaluations to label sensitivity, the practical separation from perturbation-based baselines disappears. Latent factor coverage is offered as mitigation, yet without an explicit sub-exponential supervision procedure the reported outperformance on subset-sensitivity prediction and faithful rankings may be an artifact of how the training distribution was built rather than a general property of the estimator.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes Knot, a structured rank-aware knowledge dependency estimator for reliable QA. It learns subset sensitivity from counterfactual supervision, models interactions (redundancy, substitutability, complementarity) via coverage over latent dependency factors, and derives unit rankings to identify influential knowledge without extra QA-model calls at inference. The abstract claims outperformance over baselines on multiple-choice and generative QA benchmarks for subset-sensitivity prediction and faithful rankings, plus utility for early flagging of error-prone predictions.","tokens_in":1834,"tokens_out":401,"duration_ms":30136,"significance":"If the empirical claims hold and the supervision procedure is efficient, Knot could offer a deployable method for dependency estimation that improves reliability screening in LLM-based QA systems. The latent-factor approach to capturing knowledge interactions without exhaustive perturbation is a potentially useful modeling contribution.","major_comments":[{"comment":"Method section: the description of subset-level counterfactual supervision does not specify a sub-exponential generation procedure. If labeling requires enumerating or heavily sampling subsets with repeated black-box QA evaluations, the claimed separation from perturbation-based baselines disappears at training time, undermining the practical advantage stated in the abstract.","section":"Method"},{"comment":"Experiments section: no quantitative results, baseline details, or experimental setup (e.g., number of units, subset sampling strategy, or exact metrics) are supplied in the abstract or summary description, preventing assessment of whether the reported outperformance on sensitivity prediction and rankings is robust or an artifact of supervision construction.","section":"Experiments"}],"minor_comments":[{"comment":"Abstract: include at least one key quantitative result (e.g., improvement in subset-sensitivity AUC or ranking correlation) to support the outperformance claim.","section":"Abstract"},{"comment":"Notation: clarify how latent dependency factors are defined and how coverage is computed to derive unit scores.","section":"Method"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment point by point below.","responses":[{"response":"We agree that the method section requires an explicit description of the subset generation procedure to clarify efficiency. Knot uses a latent-factor-guided sampling approach that selects subsets based on coverage of dependency factors rather than exhaustive enumeration, resulting in a number of black-box QA calls that scales linearly with the number of units (with a small constant factor from repeated sampling per factor). We will revise the manuscript to include the precise sampling algorithm, the bound on evaluations, and empirical training costs. This preserves the inference-time advantage while making the training procedure transparent.","revision_made":"yes","referee_comment":"[Method] Method section: the description of subset-level counterfactual supervision does not specify a sub-exponential generation procedure. If labeling requires enumerating or heavily sampling subsets with repeated black-box QA evaluations, the claimed separation from perturbation-based baselines disappears at training time, undermining the practical advantage stated in the abstract."},{"response":"The full manuscript's Experiments section contains the quantitative results on subset-sensitivity prediction and unit ranking fidelity, along with baseline implementations, the number of knowledge units per instance, the subset sampling strategy, and the exact metrics used. The abstract provides only a high-level summary due to length constraints. To improve clarity, we will add a concise experimental setup paragraph early in the paper and ensure all details are cross-referenced from the abstract claims.","revision_made":"partial","referee_comment":"[Experiments] Experiments section: no quantitative results, baseline details, or experimental setup (e.g., number of units, subset sampling strategy, or exact metrics) are supplied in the abstract or summary description, preventing assessment of whether the reported outperformance on sensitivity prediction and rankings is robust or an artifact of supervision construction."}],"tokens_in":1280,"tokens_out":405,"duration_ms":34303,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"Knot tries to estimate how much a fixed black-box QA model depends on each candidate knowledge unit in noisy, redundant settings. It learns from subset-level counterfactual supervision, uses latent factors to cover interactions like redundancy and complementarity, and produces rank-aware scores without extra model calls at inference.\n\nThe framing around realistic knowledge sources (retrieval, decomposition, reasoning) is reasonable, and the structured estimator looks distinct from plain attention or single-unit perturbation. If the experiments actually deliver better subset-sensitivity prediction and more faithful rankings than deployable baselines, plus useful error flagging, that would be a practical step.\n\nThe abstract states outperformance on multiple-choice and generative benchmarks but gives zero quantitative results, baseline names, or experimental details. That makes it impossible to judge whether the data or method supports the claims. The stress-test concern about supervision cost is on point: the abstract does not show a sub-exponential way to generate the counterfactual labels, so any claimed separation from perturbation methods could disappear once training cost is counted.\n\nThis is aimed at people working on reliable LLM QA who need dependency scores for risk screening. A reader focused on that application might find the problem setup useful, but without the numbers the work does not look ready for serious refereeing.","headline":"Knot proposes a rank-aware dependency estimator trained on subset counterfactuals but the abstract supplies no numbers or setup details at all.","tokens_in":2293,"tokens_out":324,"would_cite":false,"duration_ms":26285,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Knot estimates how sensitive a black-box QA model is to each knowledge unit by learning from subset counterfactuals and latent factors.","keywords":["knowledge dependency estimation","question answering","LLM reliability","counterfactual supervision","sensitivity estimation","risk screening","black-box models"],"falsifier":"Measure actual QA output changes under exhaustive single-unit and subset perturbations on a held-out benchmark and check whether Knot's predicted dependency scores correlate with those measured changes.","tokens_in":2595,"feed_emoji":"🔍","tokens_out":614,"duration_ms":25630,"temperature":0.7,"pith_summary":"The paper introduces a method called Knot to estimate the sensitivity of fixed QA models to different pieces of knowledge in noisy, redundant candidate sets drawn from context or retrieval. It trains on subset-level counterfactual supervision to model how entire subsets affect predictions, then represents interactions among units through coverage over latent dependency factors before producing ranked unit scores. This approach avoids exhaustive test-time perturbations while capturing redundancy, substitutability, and complementarity. A sympathetic reader would care because reliable QA requires knowing not only whether an answer is correct but which knowledge actually supports it, enabling early identification of fragile predictions.","feed_headline":"Knot ranks knowledge influence for QA from subset signals","feed_subtitle":"Dependency scores flag error-prone predictions early without extra model calls across benchmarks","key_machinery":"Knot, the structured rank-aware knowledge dependency estimator, which trains on subset counterfactuals and computes unit influence via coverage over latent dependency factors.","core_discovery":"Knot learns from subset-level counterfactual supervision, models subset sensitivity through coverage over latent dependency factors, and derives rank-aware unit scores to identify influential candidates, outperforming baselines in subset-sensitivity prediction and producing more faithful rankings without extra QA-model calls.","pith_inferences":["The same subset-counterfactual training pattern could be adapted to estimate dependencies in other black-box generation tasks such as summarization or code completion.","If latent factors prove stable across domains, Knot-style estimators might serve as lightweight add-ons for any retrieval-augmented pipeline without retraining the underlying model.","Scaling the approach to very large candidate sets would require testing whether the latent-factor coverage remains computationally tractable."],"forward_implications":["Knot produces higher-accuracy subset-sensitivity predictions than compared baselines on multiple-choice and generative QA tasks.","Its unit rankings are more faithful to true influence than those from deployable baselines that avoid extra QA calls.","The resulting dependency scores can be used at inference time to screen and flag error-prone QA predictions before deployment.","The estimator captures redundancy and complementarity among knowledge units through its latent factor coverage mechanism."],"fun_headline_variants":["Knot estimates QA knowledge dependencies from subsets","Rank-aware Knot identifies key QA knowledge units","Knot learns dependency factors via counterfactual subsets","Subset signals yield Knot's influence scores for QA"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Subset-level counterfactual supervision can be generated at training time in a form that teaches accurate sensitivity without requiring exhaustive perturbation of every candidate combination.","fun_headline_variants_meta":{"raw":{"variants":["Knot estimates QA knowledge dependencies from subsets","Rank-aware Knot identifies key QA knowledge units","Knot learns dependency factors via counterfactual subsets","Subset signals yield Knot's influence scores for QA"]},"model":"grok-4.3","cost_usd":0.00458,"raw_usage":{"total_tokens":2234,"prompt_tokens":589,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":45799500,"prompt_tokens_details":{"text_tokens":589,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1591,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":589,"tokens_out":54,"duration_ms":19067,"temperature":1.0,"reasoning_tokens":1591,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T12:41:39.165696+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Measure actual QA output changes under exhaustive single-unit and subset perturbations on a held-out benchmark and check whether Knot's predicted dependency scores correlate with those measured changes.","supporting_citations":[],"review_version":1}