{"id":"b8548aea-c695-4ada-a50b-4c3b07025cd1","arxiv_id":"2608.11727","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Harness-IF scores 256 rules across coding-agent runs and finds every model performs 3.6 to 7.4 points worse on rules that oppose unprompted defaults, so aggregate compliance scores overstate true instruction following.","lead":"Harness-IF is a new benchmark that scores coding agents' compliance with individual rules placed across five instruction surfaces, and it reports that all 12 tested models follow rules that oppose their unprompted defaults less often than aggregate scores suggest. The work gives evaluators a way to separate genuine instruction following from behavior the model would have produced anyway.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Judge-swap instability is the load-bearing risk: with 86.8% of verdicts from GPT-5.2 and only κ=0.163 against Claude Opus 4.7, the 3.6–7.4-point gap cannot be separated from judge-content bias until an alternate-judge AP-Acc is computed.","rationale":"The reader's weakest assumption names exactly the right load-bearing premise. I agree with the diagnosis and sharpen it: the judge may contaminate both the pass/fail labels and the prior labels, so the instrument error is not merely random noise. The paper's own reliability section concedes the swap sensitivity and the absence of row-level swap verdicts; the deterministic-only and common-support analyses are real supporting evidence and prevent a stronger verdict, but they do not reproduce the headline margin or the model-specific rank changes. The headline gap and the rank-pair exchanges are conditional on the GPT-5.2 instrument, and publishing the row-level swap panel plus computing the alternate-judge gap would settle the concern. The paper is transparent about its limitations and provides substantial robustness checks, but the central quantitative claim remains instrument-dependent until that check is done. This does not change the reader's CONDITIONAL verdict.","tokens_in":23341,"tokens_out":10648,"duration_ms":123182,"concrete_test":"Rerun the Claude Opus 4.7 judge on a powered stratified sample (e.g., 2,000 verdicts balanced by family×modality×prior, retaining row-level labels) using the same 3-vote protocol, join the released prior labels, and recompute per-model Acc, AP-Acc, and Δ for both judges on the common rows. Accept the central claim only if the alternate-judge Δ remains positive for all 12 models and the mean Δ stays within roughly 1.5 points of 5.81; additionally, recompute the 287 zero-injection prior labels with deterministic-only checks to verify that the against-prior denominator is not judge-defined.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is the model-specific Acc−AP-Acc gap (3.6–7.4 points, mean 5.81) computed over 37,616 eligible verdicts, of which 86.8% are labeled by a GPT-5.2 three-vote judge. Appendix E.1's judge-swap calibration undermines the assumption that this instrument preserves the sign and magnitude of the gap: on a balanced 200-row subset, swapped labels agree only 62.1% with κ=0.163 on 116 paired clean verdicts, and per-agent pass-rate deltas range from −40 to +33 points—a range far larger than the 5.81-point mean gap. The row-level swap verdicts are not retained, so the alternate-judge Acc−AP-Acc gap cannot be computed. Compounding the issue, the against-prior denominator is itself defined from zero-injection probe evaluations whose scoring instrument is not reported; if the same judge family scored those probes, then a content-sensitive judge can inflate the gap twice: once when labeling a rule against-prior in the withheld condition, and again when scoring it as failed in the instructed condition. The deterministic-only subset does show a larger positive gap (+13.09 over 5,013 verdicts), which supports the direction, but it covers a different, pattern-heavy subpopulation and does not validate the displayed model-specific margins. The headline magnitude and the claim that prior control changes rank pairs therefore rest on an instrument whose error is uncharacterized with respect to the outcome variable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Harness-IF is a benchmark for rule-level instruction following in multi-turn coding agents. It curates a 642-rule library, instantiates 302 rules in 60 multi-turn coding items, and scores 256 rules from execution evidence across five configurable instruction surfaces (system prompt, tool description, skill description, project file, user instruction) plus a fixed harness default. The main methodological contribution is Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, using zero-injection probe observations and curated prior annotations. Across 12 frontier models the paper reports that AP-Acc is 3.6–7.4 points (mean 5.81) below overall accuracy for every model, with the direction surviving a common-support analysis and item-clustered intervals. A separate counterbalanced conflict pilot (E0) reports a pooled surface precedence ordering with system prompts, project files, and user instructions ahead of tool and skill descriptions. The paper also reports failure decompositions by rule family, modality, and failure type, and an exploratory non-coding extension.","tokens_in":23539,"tokens_out":12029,"duration_ms":111198,"significance":"If the headline findings hold, Harness-IF fills a genuine gap: existing instruction-following benchmarks concentrate rules in the user turn and coding-agent benchmarks emphasize final task success, so neither separates compliance from coincidence. The AP-Acc prior control is a simple and transferable idea, and the benchmark is unusually reproducible: a single script regenerates every displayed number from the released 2,160-record verdict panel. The authors are also transparent about selection optimism, judge-swap sensitivity, prior-label overlap with the evaluated cohort, and the non-auditability of the human calibration study, and the E0 pilot is properly scoped as a pooled tendency rather than a universal hierarchy. These strengths are substantial. The unresolved judge-swap issue, however, puts the model-specific magnitude and rank-pair claims at risk, so the central quantitative result is not yet established at the precision claimed.","major_comments":[{"comment":"The model-specific Δ values are not yet established as measurements of compliance rather than judge artifacts. Replacing the GPT-5.2 judge with Claude Opus 4.7 on the retained 200-row stratum gives 62.1% raw agreement and κ=0.163 on the 116 paired clean verdicts, with per-agent pass-rate deltas of −40 to +33 percentage points. Because 86.8% of eligible verdicts are produced by the GPT-5.2 judge, and because row-level swap verdicts are not retained, the alternate-judge AP-Acc and Acc−AP-Acc cannot be computed. The reported mean gap of 5.81 points is far smaller than the observed judge-swap pass-rate swings, so a judge bias correlated with against-prior content could plausibly produce or invert the obtained gaps. The deterministic-only subset in §4.5 (+13.09 points) is reassuring for the sign but is drawn from a different, pattern-heavy subpopulation and cannot validate the displayed per-model margins. I request an alternate-judge computation of the AP-Acc gap, or an analysis demonstrating that judge disagreements are independent of the against-prior label; absent that, the Δ column and the 'exchanges three adjacent rank pairs' statement should be relabeled as directional.","section":"§4.1, Table 2, Appendix E.1"},{"comment":"The AP-Acc denominator is partly defined by the very models being scored. The paper states that a 5/9 zero-injection consensus necessarily includes at least one build whose identifier overlaps the evaluated panel, so no zero-injection label is independent of the scored cohort. The mitigation—that the seven evaluated models absent from the probe cohort still show positive gaps (+5.63 vs. +6.06) and that only 44.1% of against-prior verdicts carry zero-injection labels—is an appropriate disclosure but not a test. I request a robustness analysis that recomputes AP-Acc and the gap using only curated prior labels, or only zero-injection labels determined by non-overlapping probe builds, and reports whether the 3.6–7.4-point range and the per-model ordering survive. If they do not, the claim that AP-Acc controls for unprompted defaults is materially weakened.","section":"§3.4, Appendix E.4"},{"comment":"The paper claims that prior control 'leaves the top build unchanged and exchanges three adjacent rank pairs (2–3, 4–5, and 11–12)' while also stating that common-support item-clustered intervals leave every adjacent model comparison unresolved and that the displayed order is a point ranking. Given the test-retest cell-level range of 15.6 points (Appendix E.2) and the judge-swap variability documented in Appendix E.1, the rank-pair exchanges are not statistically supported and should be reported as point-estimate observations, not as a finding of the paper. This is especially important because the abstract and conclusion repeat the rank-exchange claim without the caveat that adjacent comparisons are unresolved.","section":"Abstract, §4.1, §4.5"}],"minor_comments":[{"comment":"The paragraph beginning 'The pooled ordering survives both direction-specific fits...' is duplicated verbatim; one copy should be removed.","section":"Appendix D.4"},{"comment":"The paper states that four probe builds fall outside the evaluated panel's identifier set, but the D.4 build list appears to contain five such builds (Claude Opus 4.6, GPT 5.4, DeepSeek V3.2, Kimi K2.5, and Qwen 3.6 Plus) under the Table 3 identifiers; please reconcile the count.","section":"Appendix D.4 and Appendix E.4"},{"comment":"Section 4.1 says 86.8% of rows involve the judge, while Table 7 in Appendix E.3 reports judge plus hybrid covering 86.7% of eligible verdicts; standardize this number.","section":"§4.1 and Appendix E.3"},{"comment":"The judge-swap stratum was balanced across family×modality×prior cells, but the paired 116-row subset may have lost that balance; report the prior-label distribution of the paired subset so readers can assess how directly the swap result applies to the AP-Acc denominator.","section":"Appendix E.1"},{"comment":"Reference [30] is the GPT-5 System Card; a specific citation for the GPT-5.2 judge model would be useful.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper is honest and unusually complete in its disclosure, and the benchmark is a useful community asset. The deciding factor is that the judge-swap issue is not a presentation nit; it is the main measurement's validation. I would not reject, because the direction of the effect appears robust and the authors have the tools to fix it, but the revision must either produce an alternate-judge AP-Acc analysis or explicitly demote the Δ values and rank-pair exchanges to directional claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Harness-IF is worth serious engagement. The core idea is genuinely new: treat instruction delivery surface as an experimental variable, score rules one at a time from execution evidence, and use AP-Acc to separate real compliance from behavior the model would have produced anyway. That last piece is the real contribution. Existing IF benchmarks concentrate rules in the user turn; Harness-IF spreads them across system prompts, project files, tool descriptions, and so on, and the zero-injection probe is a clean way to label whether a rule opposes a default. The E0 conflict pilot, with its three-way tie among SP, PF, and UI, is a nice sanity check that prompt depth is not the whole story.\n\nThe paper is also unusually transparent. The authors disclose selection optimism, report a retired keyword taxonomy because it was not reproducible, provide denominator-level detail, and run common-support, threshold, and cluster analyses. That discipline earns credit. The central direction is probably right: every model does worse on against-prior rules, and the deterministic-only subset shows an even larger gap, which supports the sign of the effect.\n\nThe soft spot is the judge, and it is the load-bearing one. 86.8% of verdicts come from a GPT-5.2 three-vote judge, and the judge-swap calibration against Claude Opus 4.7 gives kappa 0.163 on 116 paired verdicts, with per-agent pass-rate deltas ranging from -40 to +33 points. That range is far larger than the 5.81-point mean gap the paper headlines. The authors concede absolute levels are instrument-specific, but they do not compute an alternate-judge AP-Acc gap, and row-level swap verdicts are not released, so the reader cannot check whether the gap survives a judge change. This is fixable—re-run the judge on the against-prior subset and report the gap—but until then the exact margins and the rank-pair swaps should be read as tentative.\n\nThe prior-label provenance is a second, smaller worry. A 5/9 zero-injection consensus necessarily includes at least one probe build that overlaps the evaluated cohort, so no zero-injection label is fully independent of the scored models. The authors bound this with a 0.43-point difference between overlapping and non-overlapping builds, which is fair, but the label lineage could be cleaner.\n\nWho benefits: anyone building coding-agent benchmarks or working on instruction-following evaluation. The artifacts are promised but not yet public, so independent verification is currently blocked. I would send this to peer review, ask for the artifacts and an alternate-judge AP-Acc as conditions, and cite it once the release ships. The measurement idea is solid enough that the field should see it.","headline":"A new benchmark axis—instruction surface placement—and a smart prior-control metric, but the headline gap rests on an LLM judge whose swap agreement is too weak to trust exact margins.","tokens_in":24240,"tokens_out":1609,"would_cite":true,"duration_ms":18422,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Coding agents pass rules they would obey anyway, and a new metric exposes the gap.","keywords":["instruction following","coding agents","benchmark evaluation","instruction surfaces","against-prior accuracy","LLM judge","execution evidence","rule-level metrics"],"falsifier":"Re-run the released 200-verdict subset through a second judge with the same three-vote protocol while retaining row-level verdicts, then recompute AP-Acc with the alternate labels; if the per-model accuracy-minus-AP-Acc gap becomes negative for several models, or if swapping the judge on the 116 paired clean verdicts removes the sign for a majority of the 12 models, the central prior-alignment claim fails.","tokens_in":22998,"feed_emoji":"🤖","tokens_out":4130,"duration_ms":39725,"temperature":0.7,"pith_summary":"If a coding agent obeys a rule, the agent may well have been going to do that anyway; existing instruction-following benchmarks cannot tell compliance from coincidence. Harness-IF scores individual rules from execution evidence across the five surfaces a deployed agent reads: system prompt, tool description, skill description, project file, and user instruction. To separate genuine compliance from default behavior, the paper introduces Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing what the agent does when the rule is withheld. Across 12 frontier models on 60 multi-turn coding items, every model scores lower on against-prior rules by 3.6 to 7.4 points, with a mean gap of 5.81 points, so aggregate accuracy overstates compliance by a model-specific margin. A separate counterbalanced pilot finds that when surfaces conflict, system prompts, project files, and user instructions tie ahead of tool and skill descriptions, so surface precedence does not follow prompt depth.","feed_headline":"Coding agents pass rules they'd already obey; new metric exposes it","feed_subtitle":"Across 12 models, obeying against-default rules is 3.6 to 7.4 points harder, a gap aggregate accuracy hides.","key_machinery":"The central object is Against-Prior Accuracy (AP-Acc), a metric that restricts scoring to rules labeled as opposing the agent's unprompted default, with defaults observed by zero-injection probes that rerun each task with the rule withheld across nine builds and curated prior annotations. It is carried by a rule-level verdict pipeline that produces one pass/fail/not-applicable judgment per applicable rule per run from traces, diffs, tests, and artifacts rather than a single task-level outcome, and by five configurable instruction surfaces plus a fixed harness default. AP-Acc is explicitly a behavioral stratification, not a training-provenance claim, and it is compared with accuracy over the same eligible verdicts so the gap is like-for-like.","core_discovery":"The paper's central claim is that instruction-following evaluation must separate prompted compliance from unprompted default behavior. Across twelve frontier coding agents on 60 realistic multi-turn items, accuracy spans 72.1-85.9% while AP-Acc spans only 66.1-78.6%, and the accuracy-minus-AP-Acc gap is positive for every model, both on the full panel and under a common-support analysis with item-clustered intervals. Aggregate instruction-following scores therefore overstate compliance, and the overstatement is model-specific: it ranges two-fold across the cohort, leaves the top-ranked build unchanged, and swaps three adjacent rank pairs. The paper also claims, from its E0 conflict pilot, that pooled surface precedence is not explained by prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.","pith_inferences":["A natural generalization of the prior-control idea is to compute AP-Acc for non-coding agents by withholding each rule from a probe run; the paper's 40-case extension uses a different metric, so this transfer remains untested.","The low judge-swap agreement on the 200-verdict subset means absolute AP-Acc levels are instrument-specific, even though shared-instrument comparisons among models can still carry the sign of the gap.","If surfaces differ in effective authority as E0 suggests, future benchmark construction should counterbalance surface placement rather than assign rules to whichever surface is easiest to phrase.","The selection of items partly by observed discriminativeness implies the reported accuracy spread may be optimistic for unselected coding tasks, so AP-Acc should be re-estimated on a held-out item set."],"forward_implications":["Aggregate instruction-following scores should be reported alongside a prior-controlled denominator such as AP-Acc, since the gap between them varies two-fold across models.","Leaderboard comparisons between adjacent models are not reliable; only the prior-alignment direction, broad difficulty patterns, and failure signatures carry weight.","Benchmark designers should deliberately inject rules that oppose known defaults if they want to measure instruction following rather than default-aligned behavior.","Because shortfall rules absorb 77.1% of failure mass, verifiers tuned to detect excess output address only about a fifth of the failures in realistic coding work.","The E0 surface ordering implies that critical constraints may be more reliably honored through system prompts, project files, and user instructions than through tool or skill descriptions."],"supporting_citations":[{"why":"Supplies the verifiable-constraint checking protocol that Harness-IF extends to rule-level scoring across multiple surfaces.","marker":"[48]"},{"why":"Represents the coding-agent benchmark family that emphasizes final task success, the contrast for Harness-IF's rule-level outcome.","marker":"[14]"},{"why":"Adds a long-horizon software-engineering benchmark whose final-success focus motivates measuring operational rules separately.","marker":"[6]"},{"why":"Defines the instruction-hierarchy view of privileged sources that E0's surface-precedence result complements without assuming.","marker":"[35]"},{"why":"Closest prior agentic instruction-following benchmark, which Harness-IF distinguishes by adding more surfaces and prior control.","marker":"[25]"},{"why":"Provides a dual-control agent benchmark used in related work for multi-turn tool use and user interaction.","marker":"[4]"},{"why":"Identifies the GPT-5.2 judge model that produces the majority of verdicts, whose reliability is the paper's dominant measurement uncertainty.","marker":"[30]"}],"fun_headline_variants":["New benchmark reveals coding agents often follow rules by coincidence","AP-Acc metric exposes models' compliance gap on against-default rules","When rules contradict defaults, coding agents obey less by up to 7.4 points","Instruction-following scores overstate compliance; new metric corrects it","Prompt depth doesn't decide rule precedence in coding agents, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the GPT-5.2 three-vote LLM judge, which decides 86.8% of eligible verdicts, produces pass/fail labels accurate enough to preserve the sign and size of the accuracy-minus-AP-Acc gap; if judge errors correlate with rule content, the model-specific gap could be an artifact of the instrument.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark reveals coding agents often follow rules by coincidence","AP-Acc metric exposes models' compliance gap on against-default rules","When rules contradict defaults, coding agents obey less by up to 7.4 points","Instruction-following scores overstate compliance; new metric corrects it","Prompt depth doesn't decide rule precedence in coding agents, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000451,"raw_usage":{"total_tokens":2292,"prompt_tokens":983,"completion_tokens":1309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":1216}},"tokens_in":599,"tokens_out":1309,"duration_ms":14135,"temperature":1.0,"reasoning_tokens":1216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:29:02.791695+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the released 200-verdict subset through a second judge with the same three-vote protocol while retaining row-level verdicts, then recompute AP-Acc with the alternate labels; if the per-model accuracy-minus-AP-Acc gap becomes negative for several models, or if swapping the judge on the 116 paired clean verdicts removes the sign for a majority of the 12 models, the central prior-alignment claim fails.","supporting_citations":[{"cited_title":"Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan","cited_arxiv_id":null,"evidence_quote":"Represents the coding-agent benchmark family that emphasizes final task success, the contrast for Harness-IF's rule-level outcome."},{"cited_title":"AgentIF: Bench- marking instruction following of large language models in agentic scenarios","cited_arxiv_id":null,"evidence_quote":"Closest prior agentic instruction-following benchmark, which Harness-IF distinguishes by adding more surfaces and prior control."}],"review_version":1}