{"id":"258f6e17-6f00-41a6-9842-c20a01b5aad9","arxiv_id":"2604.18309","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Targeted, evidence-rich context partitions improve causal clarity and actionability of LLM failure explanations while large undifferentiated contexts produce vaguer outputs, with higher-quality explanations correlating to better downstream repair rates.","lead":"The study systematically tests 93 different ways of feeding debugging information to LLMs and measures how those choices change the quality of the explanations the models produce for why a program failed. Developers and tool builders may read it to decide what context to include when building AI assistants that explain bugs instead of just suggesting fixes.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of 93 configs and real bugs for causal generalization of context effects","rationale":"The reader's weakest assumption correctly isolates the external-validity premise required to move from the controlled empirical results to the broader causal claim. Within the tested scope the design (systematic variation + human validation of LLM judge) supplies direct support; no internal inconsistency in the reported logic was identified that would override this.","tokens_in":1859,"tokens_out":282,"duration_ms":28317,"concrete_test":"Re-run the full 93-configuration experiment on a disjoint bug set drawn from a second benchmark (e.g., Defects4J if the original uses a different corpus); compare direction and effect size of context-type differences on the six quality criteria and on repair pass rates. If the pattern fails to replicate, the causal generalization weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that context composition causally affects explanation quality (and downstream repair) rests on the assumption that the 93 systematically varied configurations and the selected real bugs adequately sample the space of failure mechanisms, artifact types, and economically viable models. If the bugs cluster in particular domains or the partitions omit key failure-specific signals, the observed quality differences (evidence-rich vs. large contexts) and quartile-repair correlations may reflect dataset idiosyncrasies rather than a general causal relationship.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that LLM-generated failure explanations for debugging are causally affected by context composition. Using 93 systematically varied context configurations on real bugs and three models (gpt-5-mini, DeepSeek-V3.2, Grok-4.1-fast), it evaluates explanations on six quality criteria, validates LLM-as-a-judge scores via human ratings, and links higher explanation scores to improved downstream repair pass rates (and, for some models, closer-to-minimal fixes). Evidence-rich, failure-specific artifacts outperform overly large contexts, which produce vaguer explanations; low-score explanations can underperform a no-explanation baseline.","tokens_in":1925,"tokens_out":400,"duration_ms":43934,"significance":"If the results hold, the work supplies concrete, actionable guidance on context design for LLM debugging systems, moving beyond ad-hoc prompting. Credit is due for the systematic variation across 93 configurations, evaluation on three economically viable models, human validation of the LLM judge, explicit downstream repair-rate measurements, and the public reproduction package.","major_comments":[{"comment":"The central causal claim—that context composition affects explanation quality and repair outcomes across economically viable models—rests on the assumption that the 93 configurations and chosen real bugs adequately sample failure mechanisms and artifact types. The abstract (and presumably the experimental sections) does not detail exact bug-selection criteria or statistical controls for representativeness; without this, the observed quality differences and quartile-repair correlations risk being dataset-specific rather than general.","section":"Abstract and experimental results sections"}],"minor_comments":[{"comment":"Clarify the exact model identifiers (e.g., whether 'gpt-5-mini' is a typo or specific variant) and ensure all six evaluation criteria are defined with explicit rubrics or examples in the main text.","section":"Abstract and §4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and for acknowledging the systematic variation across 93 configurations, the use of multiple models, human validation of the LLM judge, and the downstream repair measurements. We address the major comment below.","responses":[{"response":"We agree that the abstract provides limited detail on bug selection and that the experimental sections would benefit from greater transparency on this point. The manuscript describes the bugs as real bugs drawn from open-source projects with failing tests and ground-truth fixes, and the 93 configurations systematically vary context elements such as code slices, test cases, error messages, and stack traces. To strengthen the presentation, we will revise the experimental results section to include a dedicated paragraph on dataset construction: bugs were selected from established benchmarks to include a range of failure mechanisms (e.g., null dereferences, incorrect conditionals, resource management errors) and artifact types, with explicit criteria for inclusion (reproducible failures, availability of minimal patches). We will also add a limitations paragraph noting that, while the design supports causal claims about context composition within the sampled space and across three models, full statistical representativeness of all possible software failures would require a substantially larger corpus. These changes will clarify the scope of the claims without altering the core results.","revision_made":"yes","referee_comment":"[Abstract and experimental results sections] The central causal claim—that context composition affects explanation quality and repair outcomes across economically viable models—rests on the assumption that the 93 configurations and chosen real bugs adequately sample failure mechanisms and artifact types. The abstract (and presumably the experimental sections) does not detail exact bug-selection criteria or statistical controls for representativeness; without this, the observed quality differences and quartile-repair correlations risk being dataset-specific rather than general."}],"tokens_in":1436,"tokens_out":379,"duration_ms":44291,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper shows context composition has a real effect on how well LLMs explain failures, backed by 93 varied setups, human-checked scores, and downstream repair measurements on actual bugs. Evidence-rich partitions beat large undifferentiated ones for causal and actionable quality, and higher scores tie to better patch rates while low ones can fall below a no-explanation baseline. That mapping is more direct than most prior work that treats explanations as a side effect of repair prompts. The public reproduction package and the three-model comparison add practical value. The design is systematic enough to support the claim that certain artifacts improve explanation quality over generic prompting. The soft spot is representativeness. The causal generalization rests on whether the chosen real bugs and the 93 partitions sample failure mechanisms and artifact types broadly enough. If the bugs cluster in particular domains or miss key signals, the quality differences and quartile correlations could be narrower than presented. The abstract is light on exact bug-selection criteria and full statistical controls, which makes it harder to judge how far the patterns extend to other models or codebases. This is for researchers building or evaluating LLM tools for software maintenance and debugging. A reader who needs concrete data on what context elements help or hurt explanations will get usable findings here. It has enough empirical grounding and external validation to deserve a serious referee rather than a desk reject, though the generalizability section will likely need tightening.","headline":"Context composition affects LLM failure explanation quality with measurable repair links, but the 93 configs and bug set leave generalizability open.","tokens_in":2427,"tokens_out":350,"would_cite":true,"duration_ms":28123,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical SE study of LLM context partitioning for failure explanations has no structural overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper's machinery (context modules from CODE/TEST/ERROR/slices, 6 binary quality criteria C1–C6, LLM-as-a-judge scoring, quartile-to-repair correlation) operates entirely in the domain of prompt engineering and empirical debugging evaluation. It invokes no J-cost, reciprocal symmetry, golden-ratio ladder, 8-tick periodicity, or parameter-free derivation of constants. No RS theorem (e.g., reality_from_one_distinction, Jcost uniqueness, AlexanderDuality D=3 forcing, or ArithmeticFromLogic) is paralleled or contradicted.","tokens_in":53145,"confidence":"high","tokens_out":166,"duration_ms":12350,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The quality of LLM-generated failure explanations depends causally on the composition of the debugging context provided.","keywords":["LLM-generated explanations","failure explanations","context partitioning","program slices","LLM-as-a-judge","software debugging","bug repair","explanation quality"],"falsifier":"A replication on a fresh collection of bugs or with additional LLMs that fails to reproduce the reported correlation between context type and both explanation scores and repair pass rates.","tokens_in":2747,"feed_emoji":"🐛","tokens_out":634,"duration_ms":44859,"temperature":0.7,"pith_summary":"The paper tests how different assemblies of debugging information affect the causal accuracy and usefulness of explanations that large language models produce for software failures. It runs experiments across 93 context configurations built from real bugs, varying which artifacts such as program slices, tests, and error messages are included. Focused, failure-specific evidence improves the explanations while very large undifferentiated contexts tend to produce vague ones. These quality differences also track with success rates on later repair tasks, and the automated scores receive validation from human raters. The work therefore treats explanation quality itself as a measurable first-class output rather than an incidental byproduct of debugging workflows.","feed_headline":"Context composition causally shapes LLM bug explanation quality","feed_subtitle":"Focused failure-specific artifacts improve causal and actionable clarity while large contexts produce vague results linked to lower repair成功","key_machinery":"Context partitioning, the systematic construction of 93 distinct debugging contexts from program slices and other artifacts, combined with LLM-as-a-judge scoring against human-validated criteria.","core_discovery":"By partitioning the available debugging information into distinct context compositions and scoring the resulting LLM outputs with an LLM-as-a-judge on six criteria for faithfulness and actionability, the study shows that explanation quality is causally affected by context composition: evidence-rich, failure-specific artifacts improve causal and action-oriented quality, whereas overly large contexts tend to yield vague explanations, with higher explanation-score quartiles associated with higher downstream repair pass rates.","pith_inferences":["Debugging assistants could benefit from automated selection of relevant slices rather than feeding all available artifacts to the model.","Treating explanation quality as an explicit optimization target may improve reliability in other LLM-assisted software engineering workflows.","The partitioning technique offers a way to isolate which artifacts drive diagnostic performance when applying LLMs to fault localization or root-cause analysis."],"forward_implications":["Higher explanation scores correlate with higher success rates on subsequent bug repair tasks.","Overly large contexts produce vague explanations that lack causal specificity.","Evidence-rich failure-specific artifacts improve the action-oriented usefulness of the explanations.","Low-scoring explanations can reduce repair performance below the level achieved with no explanation at all."],"fun_headline_variants":["Context composition affects LLM failure explanation quality","Failure-specific contexts aid causal LLM explanations","Large contexts produce vague LLM bug reports","Targeted debug inputs raise LLM explanation faithfulness","Context choice links to repair pass rates in LLMs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 93 context configurations and the selected real bugs are representative enough for the observed quality differences to generalize beyond the tested dataset and models.","fun_headline_variants_meta":{"raw":{"variants":["Context composition affects LLM failure explanation quality","Failure-specific contexts aid causal LLM explanations","Large contexts produce vague LLM bug reports","Targeted debug inputs raise LLM explanation faithfulness","Context choice links to repair pass rates in LLMs"]},"model":"grok-4.3","cost_usd":0.006542,"raw_usage":{"total_tokens":3108,"prompt_tokens":767,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":65424500,"prompt_tokens_details":{"text_tokens":767,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2278,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":767,"tokens_out":63,"duration_ms":28456,"temperature":1.0,"reasoning_tokens":2278,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T08:43:49.173538+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication on a fresh collection of bugs or with additional LLMs that fails to reproduce the reported correlation between context type and both explanation scores and repair pass rates.","supporting_citations":[],"review_version":2}