Pith. sign in

REVIEW 3 major objections 3 minor 18 references

Allocating a four-image visual budget to close-up tiles beats whole pages by 10.6 points in one six-project test, then loses by 4.1 points on 23 broader projects—so no evidence packet dominates.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 14:24 UTC pith:TPEXXINS

load-bearing objection A carefully controlled, unusually honest evidence-allocation experiment whose central 'tradeoff' claim overreaches slightly because the two blocks are not protocol-matched. the 3 major comments →

arxiv 2607.29058 v1 pith:TPEXXINS submitted 2026-07-31 cs.AI

Evidence-Grounded Constraint Checking in Construction Documents

classification cs.AI
keywords constraint checkingevidence allocationresolution-breadth tradeoffmultimodal document reviewvisual retrievalfail-closed decision rulesconstruction drawingscalibration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether, when a review system can attach only four images to its retrieved text evidence, it should spend them on whole-page overviews or on overlapping close-ups of the single most relevant page. To answer it, the authors build a constraint-checking pipeline that turns extracted facts into four deterministic decision states—meets, does not meet, missing information, uncertain—under a fail-closed precedence that never lets missing evidence become a pass. In a repeated test across six projects and four systems, the close-up policy improves expert-referenced decision accuracy by 10.6 percentage points; in a broader block of 23 disjoint projects it is 4.1 points worse, and an equal-image sensitivity also favors page breadth. The paper's central claim is therefore that no single evidence packet dominates: the measured effect is a resolution–breadth tradeoff, and safe deployment needs constraint-aware evidence routing plus expert review rather than a fixed zoom policy.

Core claim

The core discovery is that a bounded visual evidence budget has no universally best allocation. Holding instruction, text retrieval, output contract, and model fixed, reallocating four image attachments from retrieved page overviews (page-RAG) to one overview plus three overlapping half-span crops of the top page's ink bounding box (region-RAG) improves project-family standardized decision accuracy by 10.6 percentage points (95% project-cluster bootstrap CI 4.3–18.0; exact p=.031) in a six-project, four-system panel where every project contribution is positive. On 120 tasks from 23 disjoint projects, the same fixed comparison changes accuracy by −4.1 points (95% CI −10.2–1.9; p=.209), and th

What carries the argument

The central object is a deterministic four-state constraint reducer with a fail-closed precedence: given typed facts with expected and observed values, state, confidence, and source references, g(F) returns DOESNOTMEET if any required fact is violated, MISSINGINFORMATION if something required is missing, UNCERTAIN if evidence is unstable, and MEETS only when every required fact is verified. It carries the argument by guaranteeing that the same facts always yield the same decision and that missing or low-confidence evidence cannot be laundered into a pass. Paired with it is the experimental allocator: page-RAG spends the four-image cap on up to four distinct retrieved page overviews, while re

Load-bearing premise

The load-bearing premise is that the six held-out projects in the repeated test are representative enough that the 10.6-point gain reflects a general effect of local visual resolution rather than a quirk of those projects; the paper itself notes that six clusters cannot create project diversity absent from the data and that the 23-project block, designed after the repeated result, is not an independent confirmation.

What would settle it

Pre-register a 30-project, one-call-per-cell allocation experiment with the same four-image cap and identical text evidence. If region-RAG minus page-RAG is significantly positive across the project clusters, the paper's tradeoff claim (no dominant packet) is falsified in favor of tiling; if it is significantly negative, the repeated-test gain is confirmed project-specific. A secondary check: whether exact finding F1 stays low (below about 10%) even when decision accuracy improves, which would settle the noticing-versus-localizing interpretation.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • No fixed evidence policy should be shipped as a general default for large-format document review; the safe direction is a rule-aware allocator that spends resolution on local geometric constraints and page/file breadth on cross-document comparison or index checks.
  • Four-state fail-closed decisions make 'missing information' and 'uncertain' operationally distinct from MEETS, so review queues can triage them separately, but only if fact-level source tracing is implemented and measured—it is a contract here, not yet a validated safety improvement.
  • Repeat-run agreement cannot replace calibrated confidence: the five-call majority sits at 0.86–0.93 even when wrong, expected calibration error ranges from 0.36 to 0.67, and escalating below 0.8 confidence still leaves 43–80% retained risk, so production escalation needs evidence-completeness or a separately calibrated correctness model.
  • Decision accuracy and report completeness are different endpoints: the 10.6-point triage gain coexists with finding F1 near 9%, so deployment metrics must score exact finding sets, and future benchmarks need region-level proof annotations to measure trace correctness directly.
  • Because the breadth block was designed after seeing the repeated result, the two blocks should be read as one signal of non-dominance, and any claim of a general tiling benefit requires a pre-registered replication on new projects.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same resolution–breadth logic should transfer to other large-format professional documents where a judgment links text to a spatial region—for example, engineering drawings, zoning exhibits, or clinical image reports—so a testable prediction is that the resolution benefit grows with page scale and with tasks whose answer sits in a small local detail, and shrinks for multi-
  • Editorial inference: a content-aware router that decides the packet from the extracted fact state (unresolved entities, missing source links, conflicting observations) rather than from task family alone could outperform both fixed policies; the family-policy ablation in the paper suggests family labels are too coarse to make that routing decision.
  • Editorial inference: the all-positive six-project contribution pattern, though encouraging, is consistent with a cluster-level confound; an independent pre-registered replication on new projects—ideally with fact-level region annotations—would tell whether the 10.6-point gain is a stable property of tiling or an artefact of the six chosen projects.
  • Editorial inference: the false-pass reduction under region-RAG in the repeated test suggests higher-resolution evidence helps the model notice violations, while the persistently low finding F1 suggests noticing is not localizing; the two failure modes likely need separate interventions—better evidence routing versus better structured output or post-verification.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents an evidence-grounded constraint-checking pipeline for construction documents, with deterministic four-state rule reduction, typed facts, source provenance, and escalation. The main empirical contribution is a controlled evidence-allocation experiment: under a four-image cap, region-RAG (one page overview plus three overlapping tiles) is compared with page-RAG (four page overviews) while text retrieval and response logic are held fixed. In a repeated six-project, four-system test, region-RAG improves standardized decision accuracy by 10.6 percentage points (95% CI 4.3–18.0; exact p=.031). In a disjoint 23-project breadth extension using two systems, the estimate is -4.1 points (95% CI -10.2–1.9; exact p=.209). The paper interprets this sign reversal as a resolution–breadth tradeoff rather than a universal advantage for region-focused evidence, and it reports low finding localization, high false-pass rates, and miscalibrated repeated-run confidence.

Significance. If the central claim were fully supported, the paper would make a useful contribution: it isolates a packet-level evidence-allocation variable in multimodal document review, provides a reproducible artifact with auditable hashes, and is admirably explicit about limitations, including the small number of independent projects and the absence of inter-annotator agreement. The architecture separates deterministic rule execution from learned extraction, which is a sensible design for safety-critical constraint checking. However, the headline 'tradeoff' interpretation is currently stronger than the evidence, because the sign reversal is confounded with system and repetition changes across the two blocks, and because the exact-test procedure for the small-cluster inference is not fully transparent. These are fixable within the manuscript's scope.

major comments (3)
  1. [Section 6 / Section 5.3] The central claim that the data 'identif[y] a resolution–breadth tradeoff' rests on the sign reversal between the repeated block (+10.6, p=.031, four systems, five calls per cell) and the breadth block (-4.1, p=.209, two Gemini systems, one call per cell). The two blocks differ not only in projects but also in systems and repetition count, so the reversal is confounded. Moreover, -4.1 is not statistically significant. At most, the breadth block fails to replicate the positive effect; it does not establish that page-RAG is superior in a broader population. To support a tradeoff claim, the authors should provide a protocol-matched sensitivity (e.g., the same two Gemini systems and same call multiplicity across both project sets, or a one-call-per-cell analysis of the repeated block) before interpreting the sign reversal as a tradeoff. The equal-four-image sensitivity in Section 5.3 still h
  2. [Appendix C / Section 5.1] The statistical protocol states that exact two-sided sign-flip tests 'enumerate all sign assignments' and then reports 'This yields 26 assignments in the repeated test and 223 in the breadth extension.' With six nonzero project contributions, there are 64 possible sign assignments; with the reported 5 positive, 11 negative, and 7 zero breadth contributions, there are 2^16 = 65,536 assignments. The numbers 26 and 223 are therefore inconsistent with a full enumeration. Because exact p-values are load-bearing for both the primary positive result and the null breadth result, the procedure needs to be clarified: are zero contributions dropped? Is this the number of distinct absolute test statistics rather than assignments? If it is a Monte Carlo approximation, the label 'exact' is inappropriate. Without this clarification, the reported p-values cannot be independently verified.
  3. [Section 4.2 / Section 7] The primary positive estimate is based on six independent projects. The paper acknowledges in Section 7 that the bootstrap 'cannot create project diversity absent from the data.' This limitation is real and affects generalizability: with six clusters, the exact sign-flip p=.031 is essentially driven by the fact that all six project contributions are positive, but those six projects may not represent the broader distribution. The paper should make more prominent that the positive result is a six-project finding and should avoid language suggesting a robust general rule (e.g., 'improves project-family standardized decision accuracy' in the abstract without immediately qualifying the project count). This is not a fatal flaw, but it should be tied more directly to the tradeoff interpretation.
minor comments (3)
  1. [References] The Robertson & Zaragoza reference has an odd spacing: 'F oundations and Trends' should be 'Foundations and Trends.'
  2. [Section 3.3 / Appendix B] The phrase-match minimum min(2,k) is described as intentionally strict. It would be helpful to state explicitly how often the strict matcher rejects a semantically valid paraphrase, since the secondary finding endpoints depend on this choice; the paper already notes this in Section 7, but a one-sentence quantification would aid interpretation.
  3. [Table 1] The 'Panel mean' row pools four systems, but the table lists only three decimal places for 40.3 and 51.0; consider aligning decimal formatting and reporting the pooled cell count (2,400 trials) nearby for clarity.

Circularity Check

0 steps flagged

No circularity: the primary result is a controlled empirical contrast, not a derivation from fitted inputs.

full rationale

The paper's central claim is an empirical finding: §5.1 reports a 10.6 percentage-point repeated-test gain and §5.3 reports a -4.1 point breadth-extension estimate. Neither of these is derived from a fitted parameter. Equation (3) is only a standardization/weighting formula for the endpoint; it does not encode the treatment effect. Equation (1) is a deterministic decision reducer (DNM/MI/U/MEETS precedence); it is an implementation contract, not an estimator whose input is the outcome. The hand-set threshold τ=0.8 (eq. 2) and the min(2,k) phrase-matching rule affect escalation and secondary finding metrics only; they are not fitted to the primary accuracy contrast and are not renamed as predictions. No parameter is fit on the repeated panel and then 'predicted' on the breadth block; the paper explicitly states the breadth extension 'is not an independent confirmation' (§7) and reports it separately. The acknowledged limitations—six independent projects, post-hoc breadth-block design, differing systems and repetitions across blocks—are statistical-design concerns, not circularity. The external benchmark (AEC-Bench) is from different authors (Mankodiya et al., 2026), and no load-bearing self-citation, uniqueness theorem, or ansatz-via-citation is present. The result is therefore self-contained evidence, and the honest limitation statements strengthen rather than undermine that assessment.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The paper's primary numeric contrast rests on no fitted mathematical parameters; the main free choices are hand-set thresholds for escalation and phrase matching that affect secondary endpoints. The central empirical claim assumes inherited expert references are valid ground truth and that project-level inference is appropriate; these are stated domain assumptions rather than ad hoc constructs. No new physical entities are introduced.

free parameters (2)
  • Escalation confidence threshold τ = 0.8
    Hand-chosen in Equation (2); used for selective-risk and escalation analyses, not for the primary decision-accuracy contrast.
  • Phrase-match minimum k = 2
    Appendix B requires at least min(2,k) distinct reference phrases for a finding match; a hand-set threshold affecting precision and recall, not the four-way decision endpoint.
axioms (4)
  • domain assumption AEC-Bench expert references are valid, complete ground truth for constraint decisions
    Sections 4.1 and 7: references are inherited from the public release and documented as expert-generated, but no independent relabeling or inter-annotator agreement is available; all accuracy and false-pass claims depend on this premise.
  • domain assumption Fail-closed state precedence in Equation (1) is the correct decision semantics
    The reducer maps facts to decisions using precedence violated > missing > uncertain > meets; assuming a violation dominates absent evidence is a design choice, not independently validated against expert decisions.
  • standard math Source projects are independent sampling units for bootstrap and exact tests
    Appendix C resamples project clusters and enumerates sign assignments; this assumes project contributions are independent and exchangeable despite shared families and task variants.
  • domain assumption BM25 text retrieval surfaces the pages relevant to each constraint
    Section 3.4 holds retrieval fixed across conditions; if the top-12 ranked excerpts omit the decisive page, neither image policy can recover it, so the contrast measures allocation conditional on retrieval quality rather than end-to-end evidence sufficiency.

pith-pipeline@v1.3.0-daily-deepseek · 11516 in / 12264 out tokens · 134588 ms · 2026-08-03T14:24:06.114207+00:00 · methodology

0 comments
read the original abstract

Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and document revisions. We present an evidence-grounded pipeline that normalizes extracted facts, executes four-state rules deterministically, retains source spans, and escalates unresolved cases. We evaluate its PDF evidence allocator on 160 reference-based tasks from 29 construction projects using a repeated four-system test and a disjoint two-system breadth extension. In the repeated test, reallocating a four-image budget from retrieved page overviews to one overview and three overlapping tiles improves project-family standardized decision accuracy by 10.6 percentage points (95% project-cluster bootstrap CI: 4.3 to 18.0; exact p = 0.031). This effect does not persist in the broader block: Region-RAG changes accuracy by -4.1 points (95% CI: -10.2 to 1.9; exact p = 0.209), while an equal-image sensitivity favors page breadth. Exact finding-set recovery remains low, false passes remain common, and repeated-run agreement is poorly calibrated. The results identify a resolution-breadth trade-off rather than a universal advantage for region-focused evidence, motivating rule-aware evidence routing and expert review.

Figures

Figures reproduced from arXiv: 2607.29058 by Hugo Berard, Rashid Mushkani, Shin Koseki.

Figure 1
Figure 1. Figure 1: Constraint-checking architecture. Solid paths are instantiated for PDFs in the experiment; the dashed CAD/IFC path is an adapter contract. The rule engine emits both a decision and an auditable trace. The benchmark lacks region-level reference annotations, so the experiment evaluates decisions and finding content, not trace localization accuracy. Shared task instruction + ordered BM25 text excerpts Page-RA… view at source ↗
Figure 2
Figure 2. Figure 2: Controlled evidence-allocation intervention. Text evi￾dence and response logic are fixed; only the four visual attachment slots differ. 3.5. Typed output, trace, and escalation The model returns the public task’s typed records. A field￾strict parser rejects missing, extra, null, or empty fields. The deterministic rule layer then maps records to Equation (1); it does not ask a second model to revise the ans… view at source ↗
Figure 3
Figure 3. Figure 3: Region-RAG minus page-RAG decision accuracy. Inter￾vals are source-project cluster bootstraps. The repeated-test panel is the primary contrast; system rows and the breadth extension are secondary. 5.2. Heterogeneity and paired error audit The descriptive family contrasts are positive in all seven families but vary substantially. Region-minus-page accuracy is 26.1 percentage points for cross-reference resol… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

18 extracted references · 4 linked inside Pith

  1. [1]

    2603.29199 , archivePrefix =

    Mankodiya, Harsh and Gallik, Chase and Galanos, Theodoros and Mulyar, Andriy , year =. 2603.29199 , archivePrefix =

  2. [2]

    2607.15418 , archivePrefix =

    Jung, Yoonhwa and Fu, Junryu and Golparvar-Fard, Mani , year =. 2607.15418 , archivePrefix =

  3. [3]

    Mathew, Minesh and Karatzas, Dimosthenis and Jawahar, C. V. , booktitle =. 2021 , doi =

  4. [4]

    2024 , eprint =

    Ma, Yubo and Zang, Yuhang and Chen, Liangyu and Chen, Meiqi and Jiao, Yizhu and Li, Xinze and Lu, Xinyuan and Liu, Ziyu and Ma, Yan and Dong, Xiaoyi and Zhang, Pan and Pan, Liangming and Jiang, Yu-Gang and Wang, Jiaqi and Cao, Yixin and Sun, Aixin , booktitle =. 2024 , eprint =

  5. [5]

    2411.04952 , archivePrefix =

    Cho, Jaemin and Mahata, Debanjan and Irsoy, Ozan and He, Yujie and Bansal, Mohit , year =. 2411.04952 , archivePrefix =

  6. [6]

    International Conference on Learning Representations , year =

    Faysse, Manuel and Sibille, Hugues and Wu, Tony and Omrani, Bilel and Viaud, Gautier and Hudelot, C. International Conference on Learning Representations , year =. 2407.01449 , archivePrefix =

  7. [7]

    2022 , doi =

    Huang, Yupan and Lv, Tengchao and Cui, Lei and Lu, Yutong and Wei, Furu , booktitle =. 2022 , doi =

  8. [8]

    2022 , doi =

    Kim, Geewook and Hong, Teakgyu and Yim, Moonbin and Nam, JeongYeon and Park, Jinyoung and Yim, Jinyeong and Hwang, Wonseok and Yun, Sangdoo and Han, Dongyoon and Park, Seunghyun , booktitle =. 2022 , doi =

  9. [9]

    Automation in Construction , volume =

    Automatic Rule-Based Checking of Building Designs , author =. Automation in Construction , volume =. 2009 , doi =

  10. [10]

    Classification of Rules for Automated

    Solihin, Wawan and Eastman, Charles , journal =. Classification of Rules for Automated. 2015 , doi =

  11. [11]

    Advances in Neural Information Processing Systems , volume =

    Selective Classification for Deep Neural Networks , author =. Advances in Neural Information Processing Systems , volume =. 2017 , eprint =

  12. [12]

    Proceedings of the 34th International Conference on Machine Learning , series =

    On Calibration of Modern Neural Networks , author =. Proceedings of the 34th International Conference on Machine Learning , series =

  13. [13]

    2023 , eprint =

    Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , booktitle =. 2023 , eprint =

  14. [14]

    and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =

    Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =. 2024 , eprint =

  15. [15]

    Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =

    Measurement and Fairness , author =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , doi =

  16. [16]

    and Paullada, Amandalynne and Denton, Emily and Hanna, Alex , year =

    Raji, Inioluwa Deborah and Bender, Emily M. and Paullada, Amandalynne and Denton, Emily and Hanna, Alex , year =. 2111.15366 , archivePrefix =

  17. [17]

    The Probabilistic Relevance Framework:

    Robertson, Stephen and Zaragoza, Hugo , journal =. The Probabilistic Relevance Framework:. 2009 , doi =

  18. [18]

    Scandinavian Journal of Statistics , volume =

    A Simple Sequentially Rejective Multiple Test Procedure , author =. Scandinavian Journal of Statistics , volume =. 1979 , doi =