REVIEW 3 major objections 3 minor 18 references
Allocating a four-image visual budget to close-up tiles beats whole pages by 10.6 points in one six-project test, then loses by 4.1 points on 23 broader projects—so no evidence packet dominates.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 14:24 UTC pith:TPEXXINS
load-bearing objection A carefully controlled, unusually honest evidence-allocation experiment whose central 'tradeoff' claim overreaches slightly because the two blocks are not protocol-matched. the 3 major comments →
Evidence-Grounded Constraint Checking in Construction Documents
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The core discovery is that a bounded visual evidence budget has no universally best allocation. Holding instruction, text retrieval, output contract, and model fixed, reallocating four image attachments from retrieved page overviews (page-RAG) to one overview plus three overlapping half-span crops of the top page's ink bounding box (region-RAG) improves project-family standardized decision accuracy by 10.6 percentage points (95% project-cluster bootstrap CI 4.3–18.0; exact p=.031) in a six-project, four-system panel where every project contribution is positive. On 120 tasks from 23 disjoint projects, the same fixed comparison changes accuracy by −4.1 points (95% CI −10.2–1.9; p=.209), and th
What carries the argument
The central object is a deterministic four-state constraint reducer with a fail-closed precedence: given typed facts with expected and observed values, state, confidence, and source references, g(F) returns DOESNOTMEET if any required fact is violated, MISSINGINFORMATION if something required is missing, UNCERTAIN if evidence is unstable, and MEETS only when every required fact is verified. It carries the argument by guaranteeing that the same facts always yield the same decision and that missing or low-confidence evidence cannot be laundered into a pass. Paired with it is the experimental allocator: page-RAG spends the four-image cap on up to four distinct retrieved page overviews, while re
Load-bearing premise
The load-bearing premise is that the six held-out projects in the repeated test are representative enough that the 10.6-point gain reflects a general effect of local visual resolution rather than a quirk of those projects; the paper itself notes that six clusters cannot create project diversity absent from the data and that the 23-project block, designed after the repeated result, is not an independent confirmation.
What would settle it
Pre-register a 30-project, one-call-per-cell allocation experiment with the same four-image cap and identical text evidence. If region-RAG minus page-RAG is significantly positive across the project clusters, the paper's tradeoff claim (no dominant packet) is falsified in favor of tiling; if it is significantly negative, the repeated-test gain is confirmed project-specific. A secondary check: whether exact finding F1 stays low (below about 10%) even when decision accuracy improves, which would settle the noticing-versus-localizing interpretation.
If this is right
- No fixed evidence policy should be shipped as a general default for large-format document review; the safe direction is a rule-aware allocator that spends resolution on local geometric constraints and page/file breadth on cross-document comparison or index checks.
- Four-state fail-closed decisions make 'missing information' and 'uncertain' operationally distinct from MEETS, so review queues can triage them separately, but only if fact-level source tracing is implemented and measured—it is a contract here, not yet a validated safety improvement.
- Repeat-run agreement cannot replace calibrated confidence: the five-call majority sits at 0.86–0.93 even when wrong, expected calibration error ranges from 0.36 to 0.67, and escalating below 0.8 confidence still leaves 43–80% retained risk, so production escalation needs evidence-completeness or a separately calibrated correctness model.
- Decision accuracy and report completeness are different endpoints: the 10.6-point triage gain coexists with finding F1 near 9%, so deployment metrics must score exact finding sets, and future benchmarks need region-level proof annotations to measure trace correctness directly.
- Because the breadth block was designed after seeing the repeated result, the two blocks should be read as one signal of non-dominance, and any claim of a general tiling benefit requires a pre-registered replication on new projects.
Where Pith is reading between the lines
- Editorial inference: the same resolution–breadth logic should transfer to other large-format professional documents where a judgment links text to a spatial region—for example, engineering drawings, zoning exhibits, or clinical image reports—so a testable prediction is that the resolution benefit grows with page scale and with tasks whose answer sits in a small local detail, and shrinks for multi-
- Editorial inference: a content-aware router that decides the packet from the extracted fact state (unresolved entities, missing source links, conflicting observations) rather than from task family alone could outperform both fixed policies; the family-policy ablation in the paper suggests family labels are too coarse to make that routing decision.
- Editorial inference: the all-positive six-project contribution pattern, though encouraging, is consistent with a cluster-level confound; an independent pre-registered replication on new projects—ideally with fact-level region annotations—would tell whether the 10.6-point gain is a stable property of tiling or an artefact of the six chosen projects.
- Editorial inference: the false-pass reduction under region-RAG in the repeated test suggests higher-resolution evidence helps the model notice violations, while the persistently low finding F1 suggests noticing is not localizing; the two failure modes likely need separate interventions—better evidence routing versus better structured output or post-verification.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an evidence-grounded constraint-checking pipeline for construction documents, with deterministic four-state rule reduction, typed facts, source provenance, and escalation. The main empirical contribution is a controlled evidence-allocation experiment: under a four-image cap, region-RAG (one page overview plus three overlapping tiles) is compared with page-RAG (four page overviews) while text retrieval and response logic are held fixed. In a repeated six-project, four-system test, region-RAG improves standardized decision accuracy by 10.6 percentage points (95% CI 4.3–18.0; exact p=.031). In a disjoint 23-project breadth extension using two systems, the estimate is -4.1 points (95% CI -10.2–1.9; exact p=.209). The paper interprets this sign reversal as a resolution–breadth tradeoff rather than a universal advantage for region-focused evidence, and it reports low finding localization, high false-pass rates, and miscalibrated repeated-run confidence.
Significance. If the central claim were fully supported, the paper would make a useful contribution: it isolates a packet-level evidence-allocation variable in multimodal document review, provides a reproducible artifact with auditable hashes, and is admirably explicit about limitations, including the small number of independent projects and the absence of inter-annotator agreement. The architecture separates deterministic rule execution from learned extraction, which is a sensible design for safety-critical constraint checking. However, the headline 'tradeoff' interpretation is currently stronger than the evidence, because the sign reversal is confounded with system and repetition changes across the two blocks, and because the exact-test procedure for the small-cluster inference is not fully transparent. These are fixable within the manuscript's scope.
major comments (3)
- [Section 6 / Section 5.3] The central claim that the data 'identif[y] a resolution–breadth tradeoff' rests on the sign reversal between the repeated block (+10.6, p=.031, four systems, five calls per cell) and the breadth block (-4.1, p=.209, two Gemini systems, one call per cell). The two blocks differ not only in projects but also in systems and repetition count, so the reversal is confounded. Moreover, -4.1 is not statistically significant. At most, the breadth block fails to replicate the positive effect; it does not establish that page-RAG is superior in a broader population. To support a tradeoff claim, the authors should provide a protocol-matched sensitivity (e.g., the same two Gemini systems and same call multiplicity across both project sets, or a one-call-per-cell analysis of the repeated block) before interpreting the sign reversal as a tradeoff. The equal-four-image sensitivity in Section 5.3 still h
- [Appendix C / Section 5.1] The statistical protocol states that exact two-sided sign-flip tests 'enumerate all sign assignments' and then reports 'This yields 26 assignments in the repeated test and 223 in the breadth extension.' With six nonzero project contributions, there are 64 possible sign assignments; with the reported 5 positive, 11 negative, and 7 zero breadth contributions, there are 2^16 = 65,536 assignments. The numbers 26 and 223 are therefore inconsistent with a full enumeration. Because exact p-values are load-bearing for both the primary positive result and the null breadth result, the procedure needs to be clarified: are zero contributions dropped? Is this the number of distinct absolute test statistics rather than assignments? If it is a Monte Carlo approximation, the label 'exact' is inappropriate. Without this clarification, the reported p-values cannot be independently verified.
- [Section 4.2 / Section 7] The primary positive estimate is based on six independent projects. The paper acknowledges in Section 7 that the bootstrap 'cannot create project diversity absent from the data.' This limitation is real and affects generalizability: with six clusters, the exact sign-flip p=.031 is essentially driven by the fact that all six project contributions are positive, but those six projects may not represent the broader distribution. The paper should make more prominent that the positive result is a six-project finding and should avoid language suggesting a robust general rule (e.g., 'improves project-family standardized decision accuracy' in the abstract without immediately qualifying the project count). This is not a fatal flaw, but it should be tied more directly to the tradeoff interpretation.
minor comments (3)
- [References] The Robertson & Zaragoza reference has an odd spacing: 'F oundations and Trends' should be 'Foundations and Trends.'
- [Section 3.3 / Appendix B] The phrase-match minimum min(2,k) is described as intentionally strict. It would be helpful to state explicitly how often the strict matcher rejects a semantically valid paraphrase, since the secondary finding endpoints depend on this choice; the paper already notes this in Section 7, but a one-sentence quantification would aid interpretation.
- [Table 1] The 'Panel mean' row pools four systems, but the table lists only three decimal places for 40.3 and 51.0; consider aligning decimal formatting and reporting the pooled cell count (2,400 trials) nearby for clarity.
Circularity Check
No circularity: the primary result is a controlled empirical contrast, not a derivation from fitted inputs.
full rationale
The paper's central claim is an empirical finding: §5.1 reports a 10.6 percentage-point repeated-test gain and §5.3 reports a -4.1 point breadth-extension estimate. Neither of these is derived from a fitted parameter. Equation (3) is only a standardization/weighting formula for the endpoint; it does not encode the treatment effect. Equation (1) is a deterministic decision reducer (DNM/MI/U/MEETS precedence); it is an implementation contract, not an estimator whose input is the outcome. The hand-set threshold τ=0.8 (eq. 2) and the min(2,k) phrase-matching rule affect escalation and secondary finding metrics only; they are not fitted to the primary accuracy contrast and are not renamed as predictions. No parameter is fit on the repeated panel and then 'predicted' on the breadth block; the paper explicitly states the breadth extension 'is not an independent confirmation' (§7) and reports it separately. The acknowledged limitations—six independent projects, post-hoc breadth-block design, differing systems and repetitions across blocks—are statistical-design concerns, not circularity. The external benchmark (AEC-Bench) is from different authors (Mankodiya et al., 2026), and no load-bearing self-citation, uniqueness theorem, or ansatz-via-citation is present. The result is therefore self-contained evidence, and the honest limitation statements strengthen rather than undermine that assessment.
Axiom & Free-Parameter Ledger
free parameters (2)
- Escalation confidence threshold τ =
0.8
- Phrase-match minimum k =
2
axioms (4)
- domain assumption AEC-Bench expert references are valid, complete ground truth for constraint decisions
- domain assumption Fail-closed state precedence in Equation (1) is the correct decision semantics
- standard math Source projects are independent sampling units for bootstrap and exact tests
- domain assumption BM25 text retrieval surfaces the pages relevant to each constraint
read the original abstract
Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and document revisions. We present an evidence-grounded pipeline that normalizes extracted facts, executes four-state rules deterministically, retains source spans, and escalates unresolved cases. We evaluate its PDF evidence allocator on 160 reference-based tasks from 29 construction projects using a repeated four-system test and a disjoint two-system breadth extension. In the repeated test, reallocating a four-image budget from retrieved page overviews to one overview and three overlapping tiles improves project-family standardized decision accuracy by 10.6 percentage points (95% project-cluster bootstrap CI: 4.3 to 18.0; exact p = 0.031). This effect does not persist in the broader block: Region-RAG changes accuracy by -4.1 points (95% CI: -10.2 to 1.9; exact p = 0.209), while an equal-image sensitivity favors page breadth. Exact finding-set recovery remains low, false passes remain common, and repeated-run agreement is poorly calibrated. The results identify a resolution-breadth trade-off rather than a universal advantage for region-focused evidence, motivating rule-aware evidence routing and expert review.
Figures
Reference graph
Works this paper leans on
-
[1]
Mankodiya, Harsh and Gallik, Chase and Galanos, Theodoros and Mulyar, Andriy , year =. 2603.29199 , archivePrefix =
-
[2]
Jung, Yoonhwa and Fu, Junryu and Golparvar-Fard, Mani , year =. 2607.15418 , archivePrefix =
-
[3]
Mathew, Minesh and Karatzas, Dimosthenis and Jawahar, C. V. , booktitle =. 2021 , doi =
2021
-
[4]
2024 , eprint =
Ma, Yubo and Zang, Yuhang and Chen, Liangyu and Chen, Meiqi and Jiao, Yizhu and Li, Xinze and Lu, Xinyuan and Liu, Ziyu and Ma, Yan and Dong, Xiaoyi and Zhang, Pan and Pan, Liangming and Jiang, Yu-Gang and Wang, Jiaqi and Cao, Yixin and Sun, Aixin , booktitle =. 2024 , eprint =
2024
-
[5]
Cho, Jaemin and Mahata, Debanjan and Irsoy, Ozan and He, Yujie and Bansal, Mohit , year =. 2411.04952 , archivePrefix =
-
[6]
International Conference on Learning Representations , year =
Faysse, Manuel and Sibille, Hugues and Wu, Tony and Omrani, Bilel and Viaud, Gautier and Hudelot, C. International Conference on Learning Representations , year =. 2407.01449 , archivePrefix =
-
[7]
2022 , doi =
Huang, Yupan and Lv, Tengchao and Cui, Lei and Lu, Yutong and Wei, Furu , booktitle =. 2022 , doi =
2022
-
[8]
2022 , doi =
Kim, Geewook and Hong, Teakgyu and Yim, Moonbin and Nam, JeongYeon and Park, Jinyoung and Yim, Jinyeong and Hwang, Wonseok and Yun, Sangdoo and Han, Dongyoon and Park, Seunghyun , booktitle =. 2022 , doi =
2022
-
[9]
Automation in Construction , volume =
Automatic Rule-Based Checking of Building Designs , author =. Automation in Construction , volume =. 2009 , doi =
2009
-
[10]
Classification of Rules for Automated
Solihin, Wawan and Eastman, Charles , journal =. Classification of Rules for Automated. 2015 , doi =
2015
-
[11]
Advances in Neural Information Processing Systems , volume =
Selective Classification for Deep Neural Networks , author =. Advances in Neural Information Processing Systems , volume =. 2017 , eprint =
2017
-
[12]
Proceedings of the 34th International Conference on Machine Learning , series =
On Calibration of Modern Neural Networks , author =. Proceedings of the 34th International Conference on Machine Learning , series =
-
[13]
2023 , eprint =
Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , booktitle =. 2023 , eprint =
2023
-
[14]
and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =. 2024 , eprint =
2024
-
[15]
Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =
Measurement and Fairness , author =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , doi =
2021
-
[16]
and Paullada, Amandalynne and Denton, Emily and Hanna, Alex , year =
Raji, Inioluwa Deborah and Bender, Emily M. and Paullada, Amandalynne and Denton, Emily and Hanna, Alex , year =. 2111.15366 , archivePrefix =
-
[17]
The Probabilistic Relevance Framework:
Robertson, Stephen and Zaragoza, Hugo , journal =. The Probabilistic Relevance Framework:. 2009 , doi =
2009
-
[18]
Scandinavian Journal of Statistics , volume =
A Simple Sequentially Rejective Multiple Test Procedure , author =. Scandinavian Journal of Statistics , volume =. 1979 , doi =
1979
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.