Pith. sign in

REVIEW 3 major objections 6 minor 17 references

In a fixed-player honesty evaluation, changing the verdict grammar and disclosing the success criterion swung measured outcomes from 38/40 strong claims to 7/40 and false verdicts from 18/59 to 0/58.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Holding one player model fixed, changing the verdict grammar, criterion disclosure, and budget rendering moved measured honesty verdicts dramatically, so eval findings can reflect the instrument.

T0 review reviewed 2026-08-02 challenge →

load-bearing objection A credible, unusually self-auditing demonstration that instrument choices can dominate measured 'honesty' in LLM evals; the core warning holds, but the most quotable numbers rest on disclosed confounds and should be treated as provisional. the 3 major comments →

arxiv 2607.14399 v1 pith:VQTHEA36 submitted 2026-07-15 cs.AI cs.CL

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

classification cs.AI cs.CL
keywords evaluation instrumenthonesty evaluationoutcome taxonomycriterion disclosureverdict grammardistribution stabilitynarrative channelpreregistration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper demonstrates that, in one tightly controlled text-adventure evaluation, the measurement instrument—not the model—drove the outcomes. With the player model held fixed, expanding the verdict grammar from two to three options reduced strong claims from 38/40 to 7/40 on byte-identical instances; one sentence disclosing the success criterion cut matched-instance false verdicts from 18/59 to 0/58. Repeated runs of a fixed configuration produced non-stable verdict distributions on 3 of 4 instances, so single-run evals report samples as dispositions. The paper concludes that any eval that does not control these knobs cannot attribute its findings to the model, and it proposes a four-check integrity protocol for eval instruments.

Core claim

The central discovery is an instrument-effect demonstration: with the exact same player model, four measurement choices—how many terminal verdicts exist, whether the success criterion is disclosed, how the resource budget is rendered, and whether a narrative voice is present—substantially changed what an honesty eval would report. On four byte-identical anchor instances, moving from a binary to a three-verdict grammar moved strong claims from 38/40 to 7/40, while the newly expressible 'incomplete' verdict absorbed 28/40 outcomes; across the series, 93/158 valid games ended incomplete. Disclosing the win mechanism in one sentence took matched-instance false verdicts from 18/59 to 0/58, mainly

What carries the argument

The carrying object is a text-adventure world with engine-held ground truth: the game engine, not any model, knows whether the quest is completable and scores every terminal verdict without a judge. The player model acts in natural language through a fixed operation set; a narrator costumes engine deltas without access to ground truth, and an interpreter maps player prose to operations. Because the world is deterministic and auditable, instrument changes (verdict grammar, criterion disclosure, budget rendering, register presence) can be varied with the player fixed, making the measurement channel itself the experimental variable.

Load-bearing premise

The headline grammar-effect numbers assume that on the four byte-identical anchor instances the only effective difference between the binary and three-verdict runs was the verdict list, with time and provider-side drift negligible; the paper itself flags this confound.

What would settle it

Run the four anchor instances under both verdict grammars in an interleaved, contemporaneous design with the same player model and repeated epochs; if the strong-claim rate does not move from roughly 38/40 to 7/40—or if within-grammar outcome distributions are as variable as across-grammar ones—the taxonomy effect may be sampling noise or drift rather than a grammar effect.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Single-run LLM honesty scores should be reported as samples from a distribution, not fixed traits.
  • Adding an explicitly calibrated 'incomplete' verdict to an eval will absorb many strong claims, so binary grammars overstate false-positive behavior.
  • Disclosing the success criterion can eliminate false verdicts on matched instances, implying hidden criteria manufacture apparent dishonesty.
  • Budget rendering and narrative presence move verdicts enough to matter, so evals must either hold them constant or vary them deliberately.
  • The four-check integrity protocol (taxonomy saturation, criterion disclosure, censoring analysis, distribution replication) is cheap enough to run before any character claim.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same instrument-sensitivity likely affects other LLM evaluations that score terminal verdicts, such as tool-use, question-answering, and refusal benchmarks—not just honesty evals.
  • If narrator scarcity-compression replicates across model families, any eval where one model narrates resource state to another will inherit state-dependent distortion, so resource narration should be engine-rendered rather than model-narrated.
  • The distributional instability observed at n=10 suggests that many published single-run eval numbers have confidence intervals wide enough to change conclusions; the protocol's distribution-replication check could be adopted as a reporting standard.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports a single-system, auditable demonstration that choices of evaluation instrument—outcome taxonomy, disclosure of the success criterion, budget rendering, and narrator register—substantially change measured LM 'honesty' behavior while the player model is held fixed. The instrument is a text-adventure world in which the engine, not any model, holds ground truth; verdicts are scored against engine state. The main reported effects are: expanding a two-verdict grammar to three verdicts moved strong claims from 38/40 to 7/40 on four byte-identical anchors; a one-sentence disclosure of the win mechanism moved false verdicts from 18/59 to 0/58 on matched instances; lantern versus meter budget rendering moved strong-claim rates from .150 to .383; and repeated runs of one fixed configuration produced non-stable verdict distributions on 3 of 4 instances. The paper concludes that an evaluation that does not control these knobs cannot attribute its findings to the model, and it proposes a four-check integrity protocol. The claims are carefully scoped to one player-narrator pairing, and the paper explicitly labels some positive patterns as hypothesis-generating.

Significance. If the central demonstration holds, the paper makes a valuable contribution to the evaluation-validity literature: it provides an end-to-end, ground-truth-checkable case in which instrument choices, not model-level traits, drive headline verdict statistics. The strongest assets are the design itself (engine-held truth with a fixed player), the unusually complete audit trail (gated preregistration with run artifacts binding git revisions, a hash-pinned release manifest, recomputable tables, cross-family re-coding), and the willingness to report null and disconfirming results (the falsified P-G1 gradient and the mediation null). The proposed four-check protocol is concrete and portable. The paper does not claim external transportability of effect sizes, and it explicitly limits the scope to a single system, which is appropriate. However, the most striking causal claim—the grammar effect—currently rests on a time- and provider-confounded comparison, and several headline contrasts are censored or based on small conditional samples. These issues are addressable and do not, in my assessment, undermine the broader warning, but they need to be fixed before the central claim can be accepted

major comments (3)
  1. [§4.2, Table 2] The anchor-replication comparison is the direct support for the first and most quotable instrument effect, but the table's own note concedes that the grammar change is confounded with time and provider-side drift. The abstract and §5.1 nonetheless present 'expanding a two-verdict grammar to three verdicts' as the cause of the 38/40-to-7/40 migration. A same-time, same-provider binary-grammar arm on the same four anchors is needed before the outcome-migration numbers can be attributed to the grammar knob rather than to total configuration change. At minimum, the causal phrasing in the abstract, §1, and §5.1 should be revised to match the table note, and the paper should state explicitly that this particular contrast is a configurational demonstration, not a pure grammar counterfactual.
  2. [§4.3, Table 3] The headline disclosure contrast (18/59 to 0/58 false verdicts) conflates two channels: the intervention reduced the number of games reaching a halt (43/59 to 10/58) and, conditional on halting, produced 0 false verdicts in only 10 disclosed games. The paper does report this breakdown, but the abstract and §5.1 still lead with the unconditional contrast. The 'zero' should be accompanied by an exact binomial or Wilson confidence interval, and the conditional analysis should be made the primary statement of the disclosure effect. Otherwise readers will take a heavily censored 0/10 as evidence that disclosure eliminates false verdicts, which the paper itself disclaims.
  3. [§3.5 / §4.4] The Gate 2 design does not state whether the five cells were run in a randomized or interleaved order, nor whether a cell-order or time-trend analysis was performed. If the hero, incident, mundane, none/lantern, and none/meter cells were run sequentially, the same provider-drift concern the paper raises against Table 2 applies to the rendering and register contrasts (e.g., meter .383 versus lantern .150). Since the rendering effect is one of the four headline instrument effects, the paper should report the cell execution order, any randomization, and a time-trend check, or explicitly acknowledge this as an additional limitation of the Gate 2 comparisons.
minor comments (6)
  1. [Abstract and §5.1] The phrase 'single runs report samples as dispositions' is a useful slogan but could be misread as a claim about the model's stable traits. Consider replacing 'dispositions' with 'estimates' or adding a one-clause definition.
  2. [Table 3] The caption states that the zero is an observed count, not proof of zero probability; the same caveat should appear in the main text where the 0/58 figure is quoted.
  3. [§2 and references] Minor typo: 'Cote' should be 'Côté' in the TextWorld reference.
  4. [§4.6] The sentence 'The 11 exceptions are 10/11 verdict=complete' is slightly awkward; '10 of the 11 exceptions were verdict=complete' would be clearer.
  5. [§3.3] The distinction between 'ratified in substance' and 'ratified in full ceremony' is important and well disclosed, but the phrase 'rules-before-results' is used loosely. Consider adding a one-line summary of which gates meet which evidentiary standard in the introduction, since the reader must otherwise reconstruct it from §3.3.
  6. [§5.2] It would help to state explicitly at the top of the discussion that the instrument measures false-positive behavior only, since the true-positive cell is unpopulatable by construction. The current statement in §5.2 is clear but appears late.

Circularity Check

0 steps flagged

No significant circularity: the headline instrument effects are empirical contrasts scored against engine ground truth, not derivations from fitted or self-cited inputs.

full rationale

The paper's central claims are measurements, not reductions. With the player model fixed, it compares configurations (two- vs three-verdict grammar, hidden vs disclosed criterion, register skins, budget rendering) and reports outcome distributions (e.g., Table 2, §4.2; Table 3, §4.3; Table 4, §4.4). These contrasts are scored by a deterministic game engine, not by a model judge; no parameter is fitted to a target result and then renamed as a prediction. The one registered directional prediction (P-G1) was falsified, and the registered mediation test returned a null, which is the opposite of a circular 'prediction' forced by the design. Where effects are confounded or under-powered, the paper says so explicitly: the Table 2 grammar contrast is labeled 'total configuration change rather than a pure grammar effect' (§4.2); the register-presence effect is labeled hypothesis-generating and confounded with prompt composition (§4.4); and the incomplete verdict is described as unfalsifiable by construction. These disclosures remove the patterns that would constitute circularity: no promised effect is defined in terms of the measurement, no fitted input is promoted to a prediction, and no load-bearing self-citation chain appears—the reference list contains no self-citations, and no uniqueness or ansatz is imported from the authors' prior work. The study is a single-system demonstration with limited external validity, but that is a scope limitation, not circularity. Therefore the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 6 axioms · 0 invented entities

No physical constants or fitted coefficients; the free parameters are hand-set audit thresholds that define the headline counts. The axioms are design guarantees and sampling assumptions stated in the paper but not independently proven; they define the scope within which the instrument-effect warning holds.

free parameters (2)
  • Gate 0 stability threshold = modal category >= 8/10
    Hand-chosen preregistered cutoff; defines 'non-stable' on 3/4 instances. A different cutoff (e.g., >=7/10) changes the headline distributional-stability number.
  • Scarcity recount tolerance = max(1.5, D/12)
    Hand-set scale-aware tolerance for adjudicating budget contradictions; determines which of the 1,006 flagged mentions count as confirmed contradictions (407). The count is sensitive to this choice.
axioms (6)
  • domain assumption The engine holds exact ground truth and the narrator/interpreter never receive target location, token, NPC reliability, or instance solvability; any actionable fact not traceable to the engine delta is a leak.
    Section 3.1; if ground truth leaks through narration, verdicts no longer measure the player's epistemic state and all four effects could be leakage artifacts.
  • domain assumption Byte-identical anchor instances across Gate 0 and series 2 carry no content difference, and provider-side drift or other time-varying changes are small enough not to account for the outcome migration.
    Section 3.4 and Table 2 note; the paper discloses the confound but still uses the 38/40 vs 7/40 anchor contrast as the taxonomy counterfactual.
  • domain assumption The fixed configuration (GLM-5.2 player, Haiku-4.5 narrator and interpreter, temperature 0.7) is behaviorally identical across cells within a series; remaining variability is sampling noise quantified by epochs.
    Section 3.5; cross-cell comparisons attribute all differences to the instrument knob only if the stack is stable.
  • domain assumption A HALT with verdict 'complete' is false by construction because any genuine completion terminates the game as WIN before a HALT can be asserted.
    Sections 3.2 and 5.2; this definition underwrites all false-verdict accounting and means the measurement only captures false-positive assertion.
  • domain assumption The interpreter's mapping from player prose to operations is faithful except for the enumerated seam (11/256 interpreter-mapped HALTs with provenance laundering).
    Section 4.6; if the mapping were systematically unfaithful, verdict and confidence stamps would be measuring the interpreter rather than the player.
  • domain assumption The preregistration gate ordering (rules before results) is authentic because eval logs bind the git revision they executed and are pinned by hash; commit timestamps are author-controlled but artifacts witness the chain.
    Section 3.3 and Section 9; the evidentiary strength of 'rules-before-results' depends on this artifact-binding assumption.

reviewed 2026-08-02 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration." pith.science (2026). https://pith.science/paper/VQTHEA36

@misc{pith2026260714399,
  author       = {Pith},
  title        = {Pith review of: Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VQTHEA36}},
  note         = {Machine review of arXiv:2607.14399}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Evaluations of language-model honesty read the model's verdicts as evidence about the model. We test the instrument instead. We built a text-adventure world where the game engine, not any model, knows whether the quest can be completed. A language model plays under a budget and must eventually declare its quest complete, unreachable, or not yet decidable; the engine scores every verdict. Decision rules were recorded before results were read, and run artifacts bind the revisions they executed; the strength of preregistration varies by series and is disclosed. With the player held fixed, instrument choices substantially changed measured behavior. On four byte-identical anchors, expanding a two-verdict grammar to three verdicts moved strong claims from 38/40 to 7/40, while the new incomplete verdict took 28/40 outcomes; across series 2, 93/158 valid games ended incomplete. One sentence disclosing the success criterion took matched-instance false verdicts from 18/59 to 0/58, through fewer decision points and cleaner decisions. Repeated runs of one fixed configuration produced non-stable verdict distributions on 3 of 4 instances: single runs report samples as dispositions. A formally preregistered narrative-register gradient was falsified; two post-hoc, hypothesis-generating patterns remain: register presence roughly doubled strong claims, and budget rendering moved verdicts more than register content (.383 meter vs .150 lantern). The narrator compressed abundant budgets toward scarcity landmarks, yet the registered mediation test returned a null. We propose a four-check integrity protocol for eval instruments.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

17 extracted references · 2 linked inside Pith

  1. [1]

    But I do not halt yet

    Introduction When an evaluation reports that a language model confabulated a completion, refused honestly, or miscalibrated its confidence, the report is read as evidence about the model. That reading assumes the instrument is neutral: that the outcome taxonomy, the visibility of the success criterion, the resource budget and its rendering, and the narrat...

  2. [2]

    (i) Outcome grammar: the pre-specified anchor replication supplies the two-grammar counterfactual directly

    Four measured instrument effects, with the player model held fixed. (i) Outcome grammar: the pre-specified anchor replication supplies the two-grammar counterfactual directly. Four instances carried unchanged from Gate 0 (binary grammar) into series 2 (three-verdict gram- mar), byte-identical, 10 epochs each. Complete verdicts moved from 22/40 to 7/40, un...

  3. [3]

    Averification-culturemethod: gatedpre-resultruleswithdiscloseddifferencesinformalityand loggeddeviations, atwo-stagecodingprotocolwithfrozenconventionsandasingleaccountable judge, independent audits that twice falsified the analysis lead’s claims before results were banked, and cross-family re-extraction and re-coding of every headline count

  4. [4]

    A release bundle in which the rules-before-results chain is witnessed by the run artifacts themselves (each eval log binds the git revision it executed, with a clean/dirty flag), commit ordering provides the supporting record, and every table recomputes from released logs via a hash manifest

  5. [5]

    Model behavior moves with the surface form of the question

    Related work Surface-form sensitivity. Model behavior moves with the surface form of the question. Sclar et al. (2024) report accuracy spreads up to 76 points across equivalent prompt formats. Mizrahi et al. (2024) show instruction paraphrases reorder model rankings and call for multi-prompt evaluation. Multiple-choice answers track option-identifier prio...

  6. [6]

    The Latent Underground

    Methods 3.1 The world The instrument is a text-adventure world (“The Latent Underground”) run as an Inspect AI task (UK AI Security Institute 2024). Each instance is a generated site graph with planted ground truth: a target site and win token known to the engine and withheld from every model in the loop. All world-state changes gate through a determinist...

  7. [7]

    One mechanism is worth knowing: the quest is completed only by a commit that pins the target site with its token; a halt declaring completion does not itself complete anything

    Results 4.1 Distributional ground (Gate 0) Run the same evaluation twice and it can tell two different stories. Repeated runs of a fixed configuration (4 instances, 10 epochs each, GLM-5.2 player, Haiku-4.5 DM roles) produced non- stable verdict distributions on 3 of 4 instances under the preregistered rule (a modal category at or above 8/10 counts as sta...

  8. [8]

    Discussion 5.1 Four knobs, one fixed model Holding one player model fixed, four instrument choices substantially changed what a character eval would have reported. The outcome grammar decided whether weaker refusal was expressible at all: on the four direct anchors, adding incomplete moved strong verdicts from 38/40 to 7/40 while incomplete took 28/40 out...

  9. [9]

    Verify the outcome grammar can express the weaker claim (a calibrated incomplete or equivalent)

    Taxonomy saturation. Verify the outcome grammar can express the weaker claim (a calibrated incomplete or equivalent). Add the missing verdict and measure what it absorbs

  10. [10]

    Run a disclosed-criterion arm

    Criterion disclosure. Run a disclosed-criterion arm. If false verdicts collapse, the hidden criterion was manufacturing them, and the eval was measuring criterion opacity rather than honesty

  11. [11]

    Condition verdict rates on reaching a decision point, and check whether resource budgets censor which verdicts can be observed at all

    Censoring analysis. Condition verdict rates on reaching a decision point, and check whether resource budgets censor which verdicts can be observed at all

  12. [12]

    Re-run fixed configurations (n of at least 10 per instance here) and report verdict distributions, not single draws

    Distribution replication. Re-run fixed configurations (n of at least 10 per instance here) and report verdict distributions, not single draws. The checks are cheap and portable. On this system’s evidence they are worth running before any character claim; their portability is the claim series 3 tests. 5.4 Channel-law candidates (hypothesis-generating) Two ...

  13. [13]

    One player-narrator pairing (GLM-5.2 with Haiku-4.5 in both DM roles), one constructed world, one instrument revision per series

    Limitations External validity. One player-narrator pairing (GLM-5.2 with Haiku-4.5 in both DM roles), one constructed world, one instrument revision per series. Effect sizes are not expected to transport; the demonstrated-warning framing in section 1 is the claim’s outer boundary. Construction. The true-positive cell is unpopulatable (auto-win), so all co...

  14. [14]

    Two independent audits falsified the lead analyst’s claims before they could bank

    Process integrity as method The verification culture is reported as method because the failure class it guards against is the one the instrument measures: confident assertion outrunning ground truth. Two independent audits falsified the lead analyst’s claims before they could bank. The pre- ratificationauditshowedtheGate2controlcellwasnotclean: theleadhad...

  15. [15]

    AI-use and contribution statement The analyses, code, and substantial portions of this paper’s text were produced by a Claude Fable 5 (Anthropic)contextactingasanalysisleadundergateddecisionruleswhoseorderingandratification status are disclosed in section 3.3. Before ratification of the final experiment, an independent Fable 5 context audited the lead’s a...

  16. [16]

    (1) arXiv carries the paper source only, with the two preregistration PDFs optionally attached as ancillary files; the repository copies are canonical

    Data availability Release follows a three-artifact architecture. (1) arXiv carries the paper source only, with the two preregistration PDFs optionally attached as ancillary files; the repository copies are canonical. (2) The GitHub repository, public from submission and tagged v1.0-arxiv (https://github.com/ Aargau/latent-underground/releases/tag/v1.0-arx...

  17. [17]

    A Helpful Assistant

    References N. Alzahrani, H. A. Alyahya, Y. Alnumay, S. Alrashed, S. Alsubaie, Y. Almushaykeh, F. Mirza, N. Alotaibi, N. Altwairesh, A. Alowisheq, M. S. Bari, and H. Khan. When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. InACL 2024, 2024. arXiv:2402.01781. A. M. Bean, R. O. Kearns, A. Romanou, et al. Measuring wh...

This paper was first reviewed by deepseek-v4-flash on August 2, 2026.