Pith. sign in

REVIEW 2 major objections 22 references

Typed hallucination profiles and a risk index let calibrated multi-agent debate cut legal AI fabrications by 45%.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-27 00:31 UTC pith:FKTKCDHV

load-bearing objection Typed hallucination categories and RDI calibration add a practical layer beyond aggregate rates, but the reported gaps and 45% gains can't be verified without the missing methods. the 2 major comments →

arxiv 2606.18021 v1 pith:FKTKCDHV submitted 2026-06-16 cs.AI cs.CLcs.LGcs.MA

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

classification cs.AI cs.CLcs.LGcs.MA
keywords hallucination auditinglegal AImulti-agent debaterisk direction indextyped profilesCUAD datasetfabrication reductioncontract analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to demonstrate that overall hallucination rates around 52 percent in legal AI conceal substantial differences across claim types and bias directions that affect real deployment safety. It defines four legally relevant categories—numeric, temporal, obligation or entitlement, and factual—then measures performance gaps of 38 to 40 percentage points within the same model on contract clauses. A single scalar called the Risk Direction Index collapses the tendency toward omission versus invention into a comparable number, revealing that two systems can share the same average error rate yet point in opposite risky directions. These measurements are then fed into a multi-agent debate pipeline whose skeptic challenges and asymmetric gates are tuned to the diagnosed failure modes rather than applied generically. The resulting system reduces fabricated detections by 45 percent on 249,000 clause instances while allowing a smaller backbone to match commercial API performance.

Core claim

Across 510 contracts and 249,252 clause-level instances, within-model gaps of approximately 38-40 percentage points appear between obligation or numeric claims and temporal claims that aggregate reporting conceals, and two systems with matched 52 percent rates can carry opposite RDIs; a typed debate pipeline calibrated to both magnitudes and directions reduces fabricated detections by 45 percent with per-category gains tracking the diagnosis.

What carries the argument

LegalHalluLens framework that combines typed hallucination profiles over four claim categories, the Risk Direction Index (RDI) that condenses omission-versus-invention bias into one deployment-comparable scalar, and a multi-agent debate pipeline whose skeptic challenges and asymmetric gates are calibrated to the measured failure modes.

Load-bearing premise

The four legally-motivated claim categories capture the relevant distribution of hallucinations that matter for trustworthy legal AI deployment.

What would settle it

A controlled test on held-out legal contracts in which debate calibrated to the typed profiles and RDI shows no reduction in fabricated detections relative to a generically tuned debate pipeline, or a survey of real legal errors showing that most hallucinations fall outside the numeric-temporal-obligation-factual categories.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Two AI systems reporting identical aggregate hallucination rates can still exhibit opposite risk directions that affect procurement and accountability decisions.
  • Performance gains from the calibrated debate track the specific failure modes identified in the audit rather than appearing uniformly.
  • A model with 4 billion active parameters can reach parity with commercial APIs on legal tasks once the debate is tuned to the diagnosed error profile.
  • Direction-aware diagnostics support more precise agent design and ongoing monitoring in deployed legal workflows.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same typed-auditing approach could be applied to other regulated domains such as medical records or financial filings to expose analogous hidden bias patterns.
  • The RDI scalar might function as a procurement benchmark that lets organizations select models according to acceptable directions of risk rather than average error alone.
  • Periodic re-auditing on new contract corpora would be needed if the distribution of claim types shifts in practice.
  • Live deployment logs could be used to refine the asymmetric gates dynamically rather than relying solely on benchmark-derived profiles.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper presents LegalHalluLens, an auditing framework for hallucinations in legal AI. It defines typed profiles over four claim categories (numeric, temporal, obligation/entitlement, factual) on the CUAD dataset across 510 contracts and 249,252 clauses. It introduces a Risk Direction Index (RDI) scalar to capture omission-versus-invention bias direction. It describes a calibrated multi-agent debate pipeline using Skeptic challenges and asymmetric gates that reduces fabricated detections by 45%, tracks per-category gains, and matches commercial APIs using a 4B-parameter model. The work claims aggregate ~52% rates conceal 38-40pp category gaps and that matched-rate systems can exhibit opposite RDIs.

Significance. If the empirical mappings from typed diagnosis to pipeline gains hold and are reproducible, the framework would supply actionable, direction-aware diagnostics for legal AI procurement and agent design beyond aggregate metrics. The scale of the evaluation (249k instances) and the demonstration that RDI can differentiate systems with identical aggregate error rates are strengths. The calibrated debate result, if verified, offers a concrete path to reduce specific failure modes while using smaller backbones.

major comments (2)
  1. [Abstract] Abstract: the reported 38-40pp within-model gap between obligation/numeric and temporal claims, the opposite RDIs at matched 52% rates, and the 45% fabricated-detection reduction are presented without any description of annotation rules for the four categories, the exact RDI formula and aggregation, the asymmetric gate logic, Skeptic prompt templates, baseline configurations, or statistical tests; these omissions are load-bearing for the central claim that typed profiles and RDI serve as effective calibration inputs.
  2. [Abstract] Abstract: the claim that the four legally-motivated categories capture the relevant hallucination distribution for trustworthy legal deployment rests on an unstated assumption about coverage of CUAD clauses; no justification or sensitivity analysis is supplied to show that other hallucination types (e.g., citation or reasoning errors) do not dominate real-world risk.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback highlighting areas where the abstract and framing could be strengthened for clarity. We address each major comment below and indicate planned revisions to the manuscript.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the reported 38-40pp within-model gap between obligation/numeric and temporal claims, the opposite RDIs at matched 52% rates, and the 45% fabricated-detection reduction are presented without any description of annotation rules for the four categories, the exact RDI formula and aggregation, the asymmetric gate logic, Skeptic prompt templates, baseline configurations, or statistical tests; these omissions are load-bearing for the central claim that typed profiles and RDI serve as effective calibration inputs.

    Authors: We agree the abstract is dense and omits these details due to space limits. The full manuscript defines annotation rules in Section 3.1, the RDI formula and aggregation in Section 3.2, asymmetric gate logic and Skeptic templates in Section 4.2, baselines in Section 5.1, and statistical tests in Section 5.3. We will revise the abstract to include concise references to these elements and a high-level description of the RDI computation, making the central claims more self-contained while preserving length. revision: yes

  2. Referee: [Abstract] Abstract: the claim that the four legally-motivated categories capture the relevant hallucination distribution for trustworthy legal deployment rests on an unstated assumption about coverage of CUAD clauses; no justification or sensitivity analysis is supplied to show that other hallucination types (e.g., citation or reasoning errors) do not dominate real-world risk.

    Authors: The four categories are chosen to align with CUAD's clause-level legal annotations (obligations, entitlements, facts, numeric/temporal elements), which directly map to compliance risks in contract review. We will add explicit justification in the introduction citing legal AI literature on these risk dimensions and a limitations paragraph noting that citation/reasoning errors fall outside CUAD's scope. A full sensitivity analysis would require new datasets, but we can expand the discussion of scope without additional experiments. revision: partial

Circularity Check

0 steps flagged

No circularity: empirical measurements on external CUAD dataset with independent definitions

full rationale

The paper defines typed profiles over four claim categories on the external CUAD dataset (Hendrycks et al. 2021), introduces RDI as a reduction of omission-versus-invention bias to a scalar, and describes a calibrated debate pipeline. No equations, self-citations, or fitted parameters are shown reducing the reported gaps, 45% reduction, or opposite RDIs to the inputs by construction. The framework is self-contained against external benchmarks and does not invoke uniqueness theorems or ansatzes from prior author work.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 1 invented entities

The framework rests on the domain assumption that the four claim categories adequately represent legal hallucinations and that RDI provides a useful scalar reduction; no free parameters or invented physical entities are described.

axioms (1)
  • domain assumption Hallucinations in legal contract analysis can be meaningfully partitioned into numeric, temporal, obligation/entitlement, and factual claim categories.
    Invoked to define the typed profiles and per-category measurements.
invented entities (1)
  • Risk Direction Index (RDI) no independent evidence
    purpose: Reduces omission-versus-invention bias to a single deployment-comparable scalar.
    New metric introduced by the framework; no independent evidence outside the paper is provided.

pith-pipeline@v0.9.1-grok · 5800 in / 1490 out tokens · 46380 ms · 2026-06-27T00:31:58.944631+00:00 · methodology

0 comments
read the original abstract

AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment. We present LegalHalluLens, an auditing framework with three components: typed hallucination profiles across four legally-motivated claim categories (numeric, temporal, obligation/entitlement, factual) over CUAD (Hendrycks et al., 2021); a Risk Direction Index (RDI) that reduces omission-versus-invention bias to a single deployment-comparable scalar; and a typed debate pipeline calibrated to both magnitudes and directions. Across 510 contracts and 249,252 clause-level instances we measure a within-model gap of approximately 38-40 pp between obligation/numeric and temporal claims that aggregate reporting hides, and show that two systems with matched 52% rates can carry opposite RDIs. The debate pipeline reduces fabricated detections by 45% with per-category gains tracking the diagnosis, matching commercial APIs with a substantially smaller backbone (4B active parameters). Typed profiles and RDI surface failure modes that aggregate metrics hide; we further show these diagnostics serve as calibration inputs for multi-agent debate pipelines, where Skeptic challenges and asymmetric gates targeted at measured failure modes outperform generically-tuned debate. The framework supports direction-aware procurement, accountability, and agent design for legal AI deployed in the wild.

Figures

Figures reproduced from arXiv: 2606.18021 by Akshaj Gurugubelli, Lalit Yadav.

Figure 1
Figure 1. Figure 1: Typed hallucination rates on the 510-contract benchmark. The grey band marks the aggregate HalTP cluster (50.9–56.5%). Numeric and obligation claims hallucinate at 64.8–74.3% across every tested model; temporal claims remain at 29.0–35.1%. The resulting within-model gap (approximately 38–41 pp) is not ob￾servable under aggregate reporting [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Error direction across benchmark models (percentage of contradicted TP findings). Scope errors dominate universally (62–71%), but the residual signal reveals a deployment-critical distinction: qwen3-32b predominantly omits conditions (23.7% missing-condition errors), whereas gpt-5.2 predominantly invents them (21.0% extra-condition errors). Both systems report 52% aggregate HalTP. Only the directional deco… view at source ↗
Figure 3
Figure 3. Figure 3: Typed debate pipeline, organised into three phases. (1) Debate: a Skeptic issues claim-type-specific challenges (Appendix C); a Supporter defends with verbatim contract quotes; a Route node directs traffic. If the Skeptic flags a structural error in Round 1, the Re-extractor fires once and the loop restarts. If agents disagree with rounds remaining, the loop continues; on deadlock, the Arbiter tie-breaks c… view at source ↗
Figure 4
Figure 4. Figure 4: Per-type deltas from Experiment 2. Gains concen￾trate on obligation (∆FAR = −8.2, ∆HalGen = −6.3) and factual (∆FAR = −5.8). Temporal HalGen is essentially un￾changed (+0.6 pp), consistent with temporal being the lowest￾hallucination type at baseline. The calibrated intervention pro￾duces the per-type pattern predicted by Experiment 1. ∆Hal in the legend denotes ∆HalGen −0.4 −0.3 −0.2 −0.1 0.0 RDI (← omits… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 5 canonical work pages · 2 internal anchors

  1. [1]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics , year =

    Bang, Yejin and Ji, Ziwei and Schelten, Alan and Hartshorn, Anthony and Fowler, Tara and Zhang, Cheng and Cancedda, Nicola and Fung, Pascale , title =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics , year =

  2. [2]

    Proceedings of NeurIPS Datasets and Benchmarks , year =

    Ji, Lanlan and Seyler, Dominic and Kaur, Gunkirat and Hegde, Manjunath and Dasgupta, Koustuv and Xiang, Bing , title =. Proceedings of NeurIPS Datasets and Benchmarks , year =

  3. [3]

    Proceedings of EMNLP , year =

    Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang Wei and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , title =. Proceedings of EMNLP , year =

  4. [4]

    arXiv preprint arXiv:2407.08488 , year =

    Ravi, Selvan Sunitha and Mielczarek, Bartosz and Kannappan, Anand and Kiela, Douwe and Qian, Rebecca , title =. arXiv preprint arXiv:2407.08488 , year =

  5. [5]

    Guha, Neel and Nyarko, Julian and Ho, Daniel E. and R. arXiv preprint arXiv:2308.11462 , year =

  6. [6]

    Proceedings of the Natural Legal Language Processing Workshop , pages =

    Blair-Stanek, Andrew and Holzenberger, Nils and Van Durme, Benjamin , title =. Proceedings of the Natural Legal Language Processing Workshop , pages =

  7. [7]

    Proceedings of the Natural Legal Language Processing Workshop , year =

    Liu, Shuang and Li, Zelong and Ma, Ruoyun and Zhao, Haiyan and Du, Mengnan , title =. Proceedings of the Natural Legal Language Processing Workshop , year =

  8. [8]

    Proceedings of the Natural Legal Language Processing Workshop , year =

    Hou, Abe Bohan and Jurayj, William and Holzenberger, Nils and Blair-Stanek, Andrew and Van Durme, Benjamin , title =. Proceedings of the Natural Legal Language Processing Workshop , year =

  9. [9]

    , title =

    Dahl, Matthew and Magesh, Varun and Suzgun, Mirac and Ho, Daniel E. , title =. Journal of Legal Analysis , volume =

  10. [10]

    and Ho, Daniel E

    Magesh, Varun and Surani, Faiz and Dahl, Matthew and Suzgun, Mirac and Manning, Christopher D. and Ho, Daniel E. , title =. Journal of Empirical Legal Studies , volume =. 2025 , doi =

  11. [11]

    Mikail and Canbaz, M

    Demir, M. Mikail and Canbaz, M. Abdullah , title =. Proceedings of the Natural Legal Language Processing Workshop , year =

  12. [12]

    and Negreanu, Carina Suzana and Boxall, Kitty and Mincu, Diana , title =

    Enguehard, Joseph and Van Ermengem, Morgane and Atkinson, Kate and Cha, Sujeong and Ghosh Chowdhury, Arijit and Kallur Ramaswamy, Prashanth and Roghair, Jeremy and Marlowe, Hannah R. and Negreanu, Carina Suzana and Boxall, Kitty and Mincu, Diana , title =. Proceedings of the Natural Legal Language Processing Workshop , year =

  13. [13]

    Proceedings of NeurIPS , year =

    Hendrycks, Dan and Burns, Collin and Chen, Anya and Ball, Spencer , title =. Proceedings of NeurIPS , year =

  14. [14]

    and Mordatch, Igor , title =

    Du, Yilun and Li, Shuang and Torralba, Antonio and Tenenbaum, Joshua B. and Mordatch, Igor , title =. Proceedings of ICML , pages =

  15. [15]

    Proceedings of COLING , pages =

    Fang, Yi and Li, Moxin and Wang, Wenjie and Hui, Lin and Feng, Fuli , title =. Proceedings of COLING , pages =

  16. [16]

    Findings of EMNLP , pages =

    Li, Miaoran and Chen, Jiangning and Xu, Minghua and Wang, Xiaolong , title =. Findings of EMNLP , pages =

  17. [17]

    Proceedings of ACL , pages =

    Hu, Wentao and Zhang, Wengyu and Jiang, Yiyang and Zhang, Chen Jason and Wei, Xiaoyong and Li, Qing , title =. Proceedings of ACL , pages =

  18. [18]

    Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

    Snell, Charlie and Lee, Jaehoon and Xu, Kelvin and Kumar, Aviral , title =. arXiv preprint arXiv:2408.03314 , year =

  19. [19]

    Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

    Wu, Yangzhen and Sun, Zhiqing and Li, Shanda and Welleck, Sean and Yang, Yiming , title =. arXiv preprint arXiv:2408.00724 , year =

  20. [20]

    Proceedings of ICLR , year =

    Huang, Jie and Chen, Xinyun and Mishra, Swaroop and Zheng, Huaixiu Steven and Yu, Adams Wei and Song, Xinying and Zhou, Denny , title =. Proceedings of ICLR , year =

  21. [21]

    Proceedings of the Natural Legal Language Processing Workshop , year =

    Purushothama, Abhishek and Min, Junghyun and Waldon, Brandon and Schneider, Nathan , title =. Proceedings of the Natural Legal Language Processing Workshop , year =

  22. [22]

    arXiv preprint arXiv:2505.12864 , year=

    Fan, Yu and Ni, Jingwei and Merane, Jakob and Tian, Yang and Hermstr. arXiv preprint arXiv:2505.12864 , year=