REVIEW 3 major objections 6 minor 17 references
In a fixed-player honesty evaluation, changing the verdict grammar and disclosing the success criterion swung measured outcomes from 38/40 strong claims to 7/40 and false verdicts from 18/59 to 0/58.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Holding one player model fixed, changing the verdict grammar, criterion disclosure, and budget rendering moved measured honesty verdicts dramatically, so eval findings can reflect the instrument.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection A credible, unusually self-auditing demonstration that instrument choices can dominate measured 'honesty' in LLM evals; the core warning holds, but the most quotable numbers rest on disclosed confounds and should be treated as provisional. the 3 major comments →
Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central discovery is an instrument-effect demonstration: with the exact same player model, four measurement choices—how many terminal verdicts exist, whether the success criterion is disclosed, how the resource budget is rendered, and whether a narrative voice is present—substantially changed what an honesty eval would report. On four byte-identical anchor instances, moving from a binary to a three-verdict grammar moved strong claims from 38/40 to 7/40, while the newly expressible 'incomplete' verdict absorbed 28/40 outcomes; across the series, 93/158 valid games ended incomplete. Disclosing the win mechanism in one sentence took matched-instance false verdicts from 18/59 to 0/58, mainly
What carries the argument
The carrying object is a text-adventure world with engine-held ground truth: the game engine, not any model, knows whether the quest is completable and scores every terminal verdict without a judge. The player model acts in natural language through a fixed operation set; a narrator costumes engine deltas without access to ground truth, and an interpreter maps player prose to operations. Because the world is deterministic and auditable, instrument changes (verdict grammar, criterion disclosure, budget rendering, register presence) can be varied with the player fixed, making the measurement channel itself the experimental variable.
Load-bearing premise
The headline grammar-effect numbers assume that on the four byte-identical anchor instances the only effective difference between the binary and three-verdict runs was the verdict list, with time and provider-side drift negligible; the paper itself flags this confound.
What would settle it
Run the four anchor instances under both verdict grammars in an interleaved, contemporaneous design with the same player model and repeated epochs; if the strong-claim rate does not move from roughly 38/40 to 7/40—or if within-grammar outcome distributions are as variable as across-grammar ones—the taxonomy effect may be sampling noise or drift rather than a grammar effect.
If this is right
- Single-run LLM honesty scores should be reported as samples from a distribution, not fixed traits.
- Adding an explicitly calibrated 'incomplete' verdict to an eval will absorb many strong claims, so binary grammars overstate false-positive behavior.
- Disclosing the success criterion can eliminate false verdicts on matched instances, implying hidden criteria manufacture apparent dishonesty.
- Budget rendering and narrative presence move verdicts enough to matter, so evals must either hold them constant or vary them deliberately.
- The four-check integrity protocol (taxonomy saturation, criterion disclosure, censoring analysis, distribution replication) is cheap enough to run before any character claim.
Where Pith is reading between the lines
- The same instrument-sensitivity likely affects other LLM evaluations that score terminal verdicts, such as tool-use, question-answering, and refusal benchmarks—not just honesty evals.
- If narrator scarcity-compression replicates across model families, any eval where one model narrates resource state to another will inherit state-dependent distortion, so resource narration should be engine-rendered rather than model-narrated.
- The distributional instability observed at n=10 suggests that many published single-run eval numbers have confidence intervals wide enough to change conclusions; the protocol's distribution-replication check could be adopted as a reporting standard.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a single-system, auditable demonstration that choices of evaluation instrument—outcome taxonomy, disclosure of the success criterion, budget rendering, and narrator register—substantially change measured LM 'honesty' behavior while the player model is held fixed. The instrument is a text-adventure world in which the engine, not any model, holds ground truth; verdicts are scored against engine state. The main reported effects are: expanding a two-verdict grammar to three verdicts moved strong claims from 38/40 to 7/40 on four byte-identical anchors; a one-sentence disclosure of the win mechanism moved false verdicts from 18/59 to 0/58 on matched instances; lantern versus meter budget rendering moved strong-claim rates from .150 to .383; and repeated runs of one fixed configuration produced non-stable verdict distributions on 3 of 4 instances. The paper concludes that an evaluation that does not control these knobs cannot attribute its findings to the model, and it proposes a four-check integrity protocol. The claims are carefully scoped to one player-narrator pairing, and the paper explicitly labels some positive patterns as hypothesis-generating.
Significance. If the central demonstration holds, the paper makes a valuable contribution to the evaluation-validity literature: it provides an end-to-end, ground-truth-checkable case in which instrument choices, not model-level traits, drive headline verdict statistics. The strongest assets are the design itself (engine-held truth with a fixed player), the unusually complete audit trail (gated preregistration with run artifacts binding git revisions, a hash-pinned release manifest, recomputable tables, cross-family re-coding), and the willingness to report null and disconfirming results (the falsified P-G1 gradient and the mediation null). The proposed four-check protocol is concrete and portable. The paper does not claim external transportability of effect sizes, and it explicitly limits the scope to a single system, which is appropriate. However, the most striking causal claim—the grammar effect—currently rests on a time- and provider-confounded comparison, and several headline contrasts are censored or based on small conditional samples. These issues are addressable and do not, in my assessment, undermine the broader warning, but they need to be fixed before the central claim can be accepted
major comments (3)
- [§4.2, Table 2] The anchor-replication comparison is the direct support for the first and most quotable instrument effect, but the table's own note concedes that the grammar change is confounded with time and provider-side drift. The abstract and §5.1 nonetheless present 'expanding a two-verdict grammar to three verdicts' as the cause of the 38/40-to-7/40 migration. A same-time, same-provider binary-grammar arm on the same four anchors is needed before the outcome-migration numbers can be attributed to the grammar knob rather than to total configuration change. At minimum, the causal phrasing in the abstract, §1, and §5.1 should be revised to match the table note, and the paper should state explicitly that this particular contrast is a configurational demonstration, not a pure grammar counterfactual.
- [§4.3, Table 3] The headline disclosure contrast (18/59 to 0/58 false verdicts) conflates two channels: the intervention reduced the number of games reaching a halt (43/59 to 10/58) and, conditional on halting, produced 0 false verdicts in only 10 disclosed games. The paper does report this breakdown, but the abstract and §5.1 still lead with the unconditional contrast. The 'zero' should be accompanied by an exact binomial or Wilson confidence interval, and the conditional analysis should be made the primary statement of the disclosure effect. Otherwise readers will take a heavily censored 0/10 as evidence that disclosure eliminates false verdicts, which the paper itself disclaims.
- [§3.5 / §4.4] The Gate 2 design does not state whether the five cells were run in a randomized or interleaved order, nor whether a cell-order or time-trend analysis was performed. If the hero, incident, mundane, none/lantern, and none/meter cells were run sequentially, the same provider-drift concern the paper raises against Table 2 applies to the rendering and register contrasts (e.g., meter .383 versus lantern .150). Since the rendering effect is one of the four headline instrument effects, the paper should report the cell execution order, any randomization, and a time-trend check, or explicitly acknowledge this as an additional limitation of the Gate 2 comparisons.
minor comments (6)
- [Abstract and §5.1] The phrase 'single runs report samples as dispositions' is a useful slogan but could be misread as a claim about the model's stable traits. Consider replacing 'dispositions' with 'estimates' or adding a one-clause definition.
- [Table 3] The caption states that the zero is an observed count, not proof of zero probability; the same caveat should appear in the main text where the 0/58 figure is quoted.
- [§2 and references] Minor typo: 'Cote' should be 'Côté' in the TextWorld reference.
- [§4.6] The sentence 'The 11 exceptions are 10/11 verdict=complete' is slightly awkward; '10 of the 11 exceptions were verdict=complete' would be clearer.
- [§3.3] The distinction between 'ratified in substance' and 'ratified in full ceremony' is important and well disclosed, but the phrase 'rules-before-results' is used loosely. Consider adding a one-line summary of which gates meet which evidentiary standard in the introduction, since the reader must otherwise reconstruct it from §3.3.
- [§5.2] It would help to state explicitly at the top of the discussion that the instrument measures false-positive behavior only, since the true-positive cell is unpopulatable by construction. The current statement in §5.2 is clear but appears late.
Circularity Check
No significant circularity: the headline instrument effects are empirical contrasts scored against engine ground truth, not derivations from fitted or self-cited inputs.
full rationale
The paper's central claims are measurements, not reductions. With the player model fixed, it compares configurations (two- vs three-verdict grammar, hidden vs disclosed criterion, register skins, budget rendering) and reports outcome distributions (e.g., Table 2, §4.2; Table 3, §4.3; Table 4, §4.4). These contrasts are scored by a deterministic game engine, not by a model judge; no parameter is fitted to a target result and then renamed as a prediction. The one registered directional prediction (P-G1) was falsified, and the registered mediation test returned a null, which is the opposite of a circular 'prediction' forced by the design. Where effects are confounded or under-powered, the paper says so explicitly: the Table 2 grammar contrast is labeled 'total configuration change rather than a pure grammar effect' (§4.2); the register-presence effect is labeled hypothesis-generating and confounded with prompt composition (§4.4); and the incomplete verdict is described as unfalsifiable by construction. These disclosures remove the patterns that would constitute circularity: no promised effect is defined in terms of the measurement, no fitted input is promoted to a prediction, and no load-bearing self-citation chain appears—the reference list contains no self-citations, and no uniqueness or ansatz is imported from the authors' prior work. The study is a single-system demonstration with limited external validity, but that is a scope limitation, not circularity. Therefore the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- Gate 0 stability threshold =
modal category >= 8/10
- Scarcity recount tolerance =
max(1.5, D/12)
axioms (6)
- domain assumption The engine holds exact ground truth and the narrator/interpreter never receive target location, token, NPC reliability, or instance solvability; any actionable fact not traceable to the engine delta is a leak.
- domain assumption Byte-identical anchor instances across Gate 0 and series 2 carry no content difference, and provider-side drift or other time-varying changes are small enough not to account for the outcome migration.
- domain assumption The fixed configuration (GLM-5.2 player, Haiku-4.5 narrator and interpreter, temperature 0.7) is behaviorally identical across cells within a series; remaining variability is sampling noise quantified by epochs.
- domain assumption A HALT with verdict 'complete' is false by construction because any genuine completion terminates the game as WIN before a HALT can be asserted.
- domain assumption The interpreter's mapping from player prose to operations is faithful except for the enumerated seam (11/256 interpreter-mapped HALTs with provenance laundering).
- domain assumption The preregistration gate ordering (rules before results) is authentic because eval logs bind the git revision they executed and are pinned by hash; commit timestamps are author-controlled but artifacts witness the chain.
Cite this review
Pith. "Pith review of Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration." pith.science (2026). https://pith.science/paper/VQTHEA36
@misc{pith2026260714399,
author = {Pith},
title = {Pith review of: Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration},
year = {2026},
howpublished = {\url{https://pith.science/paper/VQTHEA36}},
note = {Machine review of arXiv:2607.14399}
}
read the original abstract
Evaluations of language-model honesty read the model's verdicts as evidence about the model. We test the instrument instead. We built a text-adventure world where the game engine, not any model, knows whether the quest can be completed. A language model plays under a budget and must eventually declare its quest complete, unreachable, or not yet decidable; the engine scores every verdict. Decision rules were recorded before results were read, and run artifacts bind the revisions they executed; the strength of preregistration varies by series and is disclosed. With the player held fixed, instrument choices substantially changed measured behavior. On four byte-identical anchors, expanding a two-verdict grammar to three verdicts moved strong claims from 38/40 to 7/40, while the new incomplete verdict took 28/40 outcomes; across series 2, 93/158 valid games ended incomplete. One sentence disclosing the success criterion took matched-instance false verdicts from 18/59 to 0/58, through fewer decision points and cleaner decisions. Repeated runs of one fixed configuration produced non-stable verdict distributions on 3 of 4 instances: single runs report samples as dispositions. A formally preregistered narrative-register gradient was falsified; two post-hoc, hypothesis-generating patterns remain: register presence roughly doubled strong claims, and budget rendering moved verdicts more than register content (.383 meter vs .150 lantern). The narrator compressed abundant budgets toward scarcity landmarks, yet the registered mediation test returned a null. We propose a four-check integrity protocol for eval instruments.
Reference graph
Works this paper leans on
-
[1]
Introduction When an evaluation reports that a language model confabulated a completion, refused honestly, or miscalibrated its confidence, the report is read as evidence about the model. That reading assumes the instrument is neutral: that the outcome taxonomy, the visibility of the success criterion, the resource budget and its rendering, and the narrat...
Pith/arXiv arXiv 2026
-
[2]
(i) Outcome grammar: the pre-specified anchor replication supplies the two-grammar counterfactual directly
Four measured instrument effects, with the player model held fixed. (i) Outcome grammar: the pre-specified anchor replication supplies the two-grammar counterfactual directly. Four instances carried unchanged from Gate 0 (binary grammar) into series 2 (three-verdict gram- mar), byte-identical, 10 epochs each. Complete verdicts moved from 22/40 to 7/40, un...
-
[3]
Averification-culturemethod: gatedpre-resultruleswithdiscloseddifferencesinformalityand loggeddeviations, atwo-stagecodingprotocolwithfrozenconventionsandasingleaccountable judge, independent audits that twice falsified the analysis lead’s claims before results were banked, and cross-family re-extraction and re-coding of every headline count
-
[4]
A release bundle in which the rules-before-results chain is witnessed by the run artifacts themselves (each eval log binds the git revision it executed, with a clean/dirty flag), commit ordering provides the supporting record, and every table recomputes from released logs via a hash manifest
-
[5]
Model behavior moves with the surface form of the question
Related work Surface-form sensitivity. Model behavior moves with the surface form of the question. Sclar et al. (2024) report accuracy spreads up to 76 points across equivalent prompt formats. Mizrahi et al. (2024) show instruction paraphrases reorder model rankings and call for multi-prompt evaluation. Multiple-choice answers track option-identifier prio...
2024
-
[6]
The Latent Underground
Methods 3.1 The world The instrument is a text-adventure world (“The Latent Underground”) run as an Inspect AI task (UK AI Security Institute 2024). Each instance is a generated site graph with planted ground truth: a target site and win token known to the engine and withheld from every model in the loop. All world-state changes gate through a determinist...
2024
-
[7]
One mechanism is worth knowing: the quest is completed only by a commit that pins the target site with its token; a halt declaring completion does not itself complete anything
Results 4.1 Distributional ground (Gate 0) Run the same evaluation twice and it can tell two different stories. Repeated runs of a fixed configuration (4 instances, 10 epochs each, GLM-5.2 player, Haiku-4.5 DM roles) produced non- stable verdict distributions on 3 of 4 instances under the preregistered rule (a modal category at or above 8/10 counts as sta...
2026
-
[8]
Discussion 5.1 Four knobs, one fixed model Holding one player model fixed, four instrument choices substantially changed what a character eval would have reported. The outcome grammar decided whether weaker refusal was expressible at all: on the four direct anchors, adding incomplete moved strong verdicts from 38/40 to 7/40 while incomplete took 28/40 out...
-
[9]
Verify the outcome grammar can express the weaker claim (a calibrated incomplete or equivalent)
Taxonomy saturation. Verify the outcome grammar can express the weaker claim (a calibrated incomplete or equivalent). Add the missing verdict and measure what it absorbs
-
[10]
Run a disclosed-criterion arm
Criterion disclosure. Run a disclosed-criterion arm. If false verdicts collapse, the hidden criterion was manufacturing them, and the eval was measuring criterion opacity rather than honesty
-
[11]
Condition verdict rates on reaching a decision point, and check whether resource budgets censor which verdicts can be observed at all
Censoring analysis. Condition verdict rates on reaching a decision point, and check whether resource budgets censor which verdicts can be observed at all
-
[12]
Re-run fixed configurations (n of at least 10 per instance here) and report verdict distributions, not single draws
Distribution replication. Re-run fixed configurations (n of at least 10 per instance here) and report verdict distributions, not single draws. The checks are cheap and portable. On this system’s evidence they are worth running before any character claim; their portability is the claim series 3 tests. 5.4 Channel-law candidates (hypothesis-generating) Two ...
-
[13]
One player-narrator pairing (GLM-5.2 with Haiku-4.5 in both DM roles), one constructed world, one instrument revision per series
Limitations External validity. One player-narrator pairing (GLM-5.2 with Haiku-4.5 in both DM roles), one constructed world, one instrument revision per series. Effect sizes are not expected to transport; the demonstrated-warning framing in section 1 is the claim’s outer boundary. Construction. The true-positive cell is unpopulatable (auto-win), so all co...
-
[14]
Two independent audits falsified the lead analyst’s claims before they could bank
Process integrity as method The verification culture is reported as method because the failure class it guards against is the one the instrument measures: confident assertion outrunning ground truth. Two independent audits falsified the lead analyst’s claims before they could bank. The pre- ratificationauditshowedtheGate2controlcellwasnotclean: theleadhad...
-
[15]
AI-use and contribution statement The analyses, code, and substantial portions of this paper’s text were produced by a Claude Fable 5 (Anthropic)contextactingasanalysisleadundergateddecisionruleswhoseorderingandratification status are disclosed in section 3.3. Before ratification of the final experiment, an independent Fable 5 context audited the lead’s a...
2026
-
[16]
(1) arXiv carries the paper source only, with the two preregistration PDFs optionally attached as ancillary files; the repository copies are canonical
Data availability Release follows a three-artifact architecture. (1) arXiv carries the paper source only, with the two preregistration PDFs optionally attached as ancillary files; the repository copies are canonical. (2) The GitHub repository, public from submission and tagged v1.0-arxiv (https://github.com/ Aargau/latent-underground/releases/tag/v1.0-arx...
2026
-
[17]
References N. Alzahrani, H. A. Alyahya, Y. Alnumay, S. Alrashed, S. Alsubaie, Y. Almushaykeh, F. Mirza, N. Alotaibi, N. Altwairesh, A. Alowisheq, M. S. Bari, and H. Khan. When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. InACL 2024, 2024. arXiv:2402.01781. A. M. Bean, R. O. Kearns, A. Romanou, et al. Measuring wh...
Pith/arXiv arXiv 2024
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.