Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

MCMit uses a constant-latency branch instruction and CNN discriminators to cut mid-circuit measurement errors in quantum circuits.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

MCMit proposes a constant-latency multi-control branch instruction, transformer and CNN discriminators, plus static MCM elimination and stochastic branching, evaluated on Qubic with QPU traces to cut latency by 70% and logical error rates by up to 9.4x.

T0 review reviewed 2026-07-01 challenge →

load-bearing objection MCMit combines a constant-latency branch instruction with CNN/transformer discriminators and software patches, but the accuracy and error-rate gains rest on offline trace tests rather than closed-loop runs with the new hardware. the 2 major comments →

arxiv 2604.25863 v2 pith:VB2SH2EY submitted 2026-04-28 quant-ph

MCMit: Mid-Circuit Measurement Error Mitigation

classification quant-ph
keywords mid-circuit measurementquantum error correctiondynamic circuitserror mitigationqubit state discriminationclassical feedbackhardware-software co-design
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents MCMit as a hardware-software co-design that targets errors arising from mid-circuit measurements and classical feedback in dynamic quantum circuits. It introduces a multi-control branch instruction to reduce feedback latency, CNN and transformer models for more accurate qubit-state discrimination even at short measurement times, and software methods including static MCM elimination and stochastic branching to handle residual errors. The approach is evaluated on experimentally extracted QPU traces, showing concrete gains in latency, accuracy, logical error rates for QEC, and overall fidelity. A sympathetic reader would care because these improvements directly address bottlenecks that limit circuit depth and reliability in both distributed quantum computing and error-corrected algorithms.

Core claim

MCMit is a hardware-software co-design for mitigating branching and latency-induced errors in mid-circuit measurements that introduces a scalable constant-latency multi-control branch instruction for faster classical feedback, transformer and CNN qubit-state discriminators that maintain high accuracy at short measurement durations, and software techniques of static MCM elimination and stochastic branching to address remaining errors.

What carries the argument

The constant-latency multi-control branch instruction paired with CNN and transformer qubit-state discriminators.

Load-bearing premise

The CNN and transformer discriminators trained on extracted QPU traces will retain their accuracy gains when executed in real time on actual hardware controllers.

What would settle it

A side-by-side run of the same quantum error correction circuit on the same QPU with and without the MCMit branch instruction and discriminators, measuring the resulting logical error rates.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Feedback latency drops by up to 70 percent, supporting circuit depths up to seven times larger than the Qubic baseline.
  • The CNN discriminator delivers up to 62 percent higher accuracy for short measurement windows.
  • Logical error rates in quantum error correction drop by factors between 1.2 and 9.4.
  • Software mitigation raises circuit fidelity by 18 to 30 percent compared with baseline methods.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Faster feedback could make mid-circuit measurements practical in larger-scale error-corrected algorithms that currently exceed coherence limits.
  • The discriminators might be combined with existing calibration routines to further reduce the need for long measurement times.
  • Testing the full stack on a physical controller would reveal whether the reported accuracy holds under real-time scheduling constraints.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper proposes MCMit, a hardware-software co-design for mitigating errors from mid-circuit measurements (MCMs) and classical feedback in dynamic quantum circuits for DQC and QEC. It introduces a constant-latency multi-control branch instruction on the Qubic controller, CNN and transformer qubit-state discriminators trained on experimental traces, and software techniques (static MCM elimination, stochastic branching). Evaluation on experimentally extracted QPU readout traces claims up to 70% feedback latency reduction (7× circuit depth improvement), up to 62% higher CNN accuracy at short durations (yielding 1.2×–9.4× lower QEC logical error rates), and 18–30% fidelity gains from software mitigation over baselines.

Significance. If the integrated performance claims hold, MCMit would meaningfully advance practical dynamic-circuit execution by jointly tackling latency and discrimination bottlenecks that currently limit QEC and DQC. The use of real QPU traces and controller implementation provides concrete grounding; the co-design framing and reported numerical gains on logical error rates are the primary contributions.

major comments (2)
  1. [Abstract / Evaluation] Abstract and Evaluation (results on discriminators): The headline claims that the CNN achieves up to 62% higher accuracy for short measurement durations and produces 1.2×–9.4× lower logical error rates rest on offline training/testing of the discriminator on static experimentally extracted traces. The manuscript provides no closed-loop experiments that combine the new constant-latency branch instruction with the CNN/transformer in real-time feedback; any timing skew or signal-path change introduced by the branch could alter the input statistics seen by the discriminator, directly undermining the reported accuracy and logical-error improvements.
  2. [Implementation / Evaluation] Implementation and Evaluation sections: The branch-instruction latency reduction (70%) and circuit-depth improvement (7×) are presented separately from the discriminator accuracy results. No integrated benchmark is shown that measures end-to-end logical error rates or fidelity when the new branch, controller firmware, and real-time discriminator are all active together, which is required to substantiate the holistic co-design claim.
minor comments (1)
  1. [Abstract] The abstract states results are obtained 'on Qubic' yet the discriminator numbers are explicitly from offline trace processing; a short clarifying sentence on the distinction between controller implementation and evaluation methodology would improve readability.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback emphasizing the importance of integrated evaluation for our hardware-software co-design claims. We agree that the current manuscript evaluates components separately using real QPU traces and controller measurements, without a single closed-loop benchmark. We will revise the abstract, evaluation sections, and add a limitations discussion to clarify the methodology, assumptions, and scope of the reported gains. Our point-by-point responses follow.

read point-by-point responses
  1. Referee: [Abstract / Evaluation] Abstract and Evaluation (results on discriminators): The headline claims that the CNN achieves up to 62% higher accuracy for short measurement durations and produces 1.2×–9.4× lower logical error rates rest on offline training/testing of the discriminator on static experimentally extracted traces. The manuscript provides no closed-loop experiments that combine the new constant-latency branch instruction with the CNN/transformer in real-time feedback; any timing skew or signal-path change introduced by the branch could alter the input statistics seen by the discriminator, directly undermining the reported accuracy and logical-error improvements.

    Authors: We acknowledge that the discriminator accuracy and logical-error results derive from offline training and testing on static traces extracted from the QPU, while the branch-instruction latency is measured independently on the Qubic controller. The manuscript does not include closed-loop real-time experiments integrating both. The traces reflect actual hardware noise and readout characteristics; the branch instruction operates on the classical feedback path after discrimination and does not modify the analog signal path or measurement duration seen by the discriminator. Nevertheless, the referee's point is valid regarding potential unmodeled interactions. In revision we will (i) explicitly qualify the abstract and evaluation claims as component-wise results, (ii) add text stating the assumption that reduced classical latency does not alter discriminator input statistics, and (iii) include a limitations paragraph noting the value of future closed-loop validation. revision: partial

  2. Referee: [Implementation / Evaluation] Implementation and Evaluation sections: The branch-instruction latency reduction (70%) and circuit-depth improvement (7×) are presented separately from the discriminator accuracy results. No integrated benchmark is shown that measures end-to-end logical error rates or fidelity when the new branch, controller firmware, and real-time discriminator are all active together, which is required to substantiate the holistic co-design claim.

    Authors: The evaluations are intentionally modular: latency and depth improvements are quantified via controller implementation and trace-driven circuit simulation, while discriminator performance uses the same experimental traces for training and inference accuracy. No single end-to-end benchmark combining live branch execution, firmware, and real-time discrimination is reported. This separation follows from the distinct experimental setups (hardware controller timing versus offline trace processing). We agree this limits the strength of the integrated co-design narrative. In the revision we will reorganize the evaluation section to present the components as complementary parts of the co-design, add an explicit statement that end-to-end logical-error measurements under simultaneous operation remain future work, and temper the abstract accordingly while retaining the individual quantitative results. revision: partial

Circularity Check

0 steps flagged

No circularity; claims rest on direct experimental evaluation rather than fitted or self-referential derivations.

full rationale

The paper presents its core results—70% latency reduction, 62% accuracy gain for the CNN discriminator, 1.2–9.4× logical-error improvement, and 18–30% fidelity gain—as outcomes of implementing the branch instruction, CNN/transformer discriminators, and software mitigation on Qubic hardware and measuring them against experimentally extracted QPU readout traces. No equations, parameter fits, or self-citations are shown that would reduce any reported gain to an input by construction; the evaluation chain remains external to the claimed improvements.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only input supplies no identifiable free parameters, axioms, or invented entities; full paper would be needed to audit these.

reviewed 2026-07-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of MCMit: Mid-Circuit Measurement Error Mitigation." pith.science (2026). https://pith.science/paper/VB2SH2EY

@misc{pith2026260425863,
  author       = {Pith},
  title        = {Pith review of: MCMit: Mid-Circuit Measurement Error Mitigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VB2SH2EY}},
  note         = {Machine review of arXiv:2604.25863}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Distributed Quantum Computing (DQC) and Quantum Error Correction (QEC) rely on dynamic circuits that include Mid-Circuit Measurements (MCMs) and classical feedback. These operations present a major bottleneck: MCMs suffer from high error rates that lead to real-time branching errors, while MCM and classical feedback latencies amplify decoherence errors. Current hardware controllers, qubit-state discriminators, and software error mitigation techniques fail to address these challenges holistically. We propose MCMit, a hardware-software co-design to mitigate branching and latency-induced errors. MCMit introduces a scalable, constant-latency multi-control branch instruction for faster classical feedback and two qubit-state discriminators, a transformer, and a CNN, with high accuracy even under short measurement durations. On the software side, static MCM elimination and stochastic branching complement the hardware by mitigating residual branching errors that persist despite hardware improvements. We implement MCMit on Qubic and evaluate it using experimentally extracted QPU readout traces. Our branch instruction reduces feedback latency by up to 70%, improving circuit depths by up to $7\times$ over Qubic. Our CNN discriminator achieves up to 62% higher accuracy for short measurement durations than the baselines, leading to 1.2$\times$--9.4$\times$ lower logical error rates in QEC. Last, our software mitigation improves fidelity by 18--30% over baseline methods.

Figures

Figures reproduced from arXiv: 2604.25863 by Aleksandra \'Swierkowska, Benjamin Lienhard, Emmanouil Giortamis, Felix Gust, Innocenzo Fulginiti, Martin Schulz, Pramod Bhatotia, Sandra Stankovic, Xiaorang Guo, Yanbin Chen.

Figure 1
Figure 1. Figure 1: Superconducting readout (§ 2.1). (a) Superconducting qubit readout pipeline on an FPGA controller. (b) The MCM in a tele￾portation circuit can produce a readout trace of 1𝜇𝑠: 500 samples spaced 2ns apart. (c) The readout trace is input to a Feed-forward Neural Network comprising N layers. The network discriminates 0 and 1 using a decision boundary that separates the states. measurement rounds, directly com… view at source ↗
Figure 2
Figure 2. Figure 2: MECH [112] performance analysis. (a) Impact of MCM errors as a ratio to the 2-qubit gate errors. (b) Impact of MCM latency as a ratio to the 2-qubit gate latency. (c) Impact of classical feedback latency as a ratio to the 2-qubit gate latency. There is a linear performance improvement with MCM error and (classical) latency reduction. 2 Background In this section, we detail the superconducting qubit readout… view at source ↗
Figure 3
Figure 3. Figure 3: Logical error rate of the surface code on the view at source ↗
Figure 5
Figure 5. Figure 5: MCMit workflow (§ 4.2). (a) Compile-time workflow and (b) runtime workflow. joint outcomes of multiple qubits. This instruction enables more com￾plex, constant-latency feedback protocols that are not possible with existing branching. This functionality is particularly advantageous for routines such as parity checks and majority voting [12, 106], where MCM results are used to directly address a pre-computed… view at source ↗
Figure 6
Figure 6. Figure 6: MCMit controller (§ 5). Blue boxes show modified components, yellow boxes show ADC/DAC components, and green boxes show example data tables. Steps (1)-(8) are executed for a conditional feedback operation based on an MCM result. Runtime workflow. For each set of MCMs that generates a set of {I, Q} traces, (1) we first discriminate the qubit-states and extract the discrimination confidence. (2) If it is bel… view at source ↗
Figure 7
Figure 7. Figure 7: The MCMit qubit-state discriminators (§ 6). exhibits a large size that stresses FPGA resources ( view at source ↗
Figure 8
Figure 8. Figure 8: Dynamic circuit simplification (§ 7.1). The original dynamic circuit is replaced by a static sub-circuit containing a probabilistic gate, whose outcome is decided at compile time. their associated control logic (§ 7.1), (2) measurement hardening via repetition codes, parity checks, and repeated measurements to detect and correct bitflip errors (§ 7.2), and (3) stochastic branching, which probabilistically … view at source ↗
Figure 9
Figure 9. Figure 9: Software MCM error mitigation (§ 7). (a) GHZ state dynamic circuit. (b) Dynamic circuit simplification removes MCMs when possible and simplifies classical conditional logic. (c) Measurement hardening leverages repetition codes, parity checks, and flag qubits to detect and/or correct errors. (d) Stochastic branching factors MCM errors into branching decisions. 10 50 100 250 500 750 1000 Number of instances … view at source ↗
Figure 10
Figure 10. Figure 10: Classical feedback latency impact (§ 8.2). The x-axis shows the number of instances of the GHZ/CNOT circuit in a higher-level application. MCMit achieves 57.3% and 37.8% higher fidelity than Qubic, on average. 7.3 Stochastic Branching To address branching errors caused by MCM bitflips in dynamic cir￾cuits, we leverage confusion matrix [73] data provided by the MCMit controller to inform stochastic compila… view at source ↗
Figure 11
Figure 11. Figure 11: MCMit software error mitigation impact on fidelity (§ view at source ↗
Figure 12
Figure 12. Figure 12: Readout duration impact on fidelity (§ 8.3). The x￾axis shows the number of teleportation steps. A 250ns readout achieves 6% higher fidelity than that of 750ns, on average. manner, respectively). “Raw” indicates completely unmitigated re￾sults, and the MCMit and Qiskit M3 results do not use other error mitigation techniques, for fairness. Each experiment runs for 10.000 shots on the same calibration cycle… view at source ↗
Figure 13
Figure 13. Figure 13: Impact of varying readout duration and fidelity on logical error rate. view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Oraqle: An Empirical Analysis of Qubit Readout and Discriminators in Quantum Error Correction

    quant-ph 2026-08 conditional novelty 6.0

    Using real 5-qubit traces, this study shows readout windows can be cut to ~600 ns with negligible QEC penalty and small discriminators match large ones.

Reference graph

Works this paper leans on

61 extracted references · 61 canonical work pages · cited by 1 Pith paper · 12 internal anchors

  1. [1]

    System card: Claude opus 4 & claude sonnet 4

    Anthropic. System card: Claude opus 4 & claude sonnet 4. https://www-cdn.anthropic.com/ 6d8a8055020700718b0c49369f60816ba2a7c285.pdf, 2025

  2. [2]

    Detection of Interest Points Using Symmetry

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. InICCV, 2015. doi: 10.1109/ICCV .2015.279

  3. [3]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923, 2025

  4. [4]

    Large language models must be taught to know what they don’t know

    Umang Bhatt, Katherine Collins, Samuel Dooley, Micah Goldblum, Nate Gruver, Sanyam Kapoor, Arka Pal, Manley Roberts, Adrian Weller, and Andrew Wilson. Large language models must be taught to know what they don’t know. InNeurIPS, 2024. doi: 10.52202/079017-2729. URL https://doi.org/10. 52202/079017-2729

  5. [5]

    Germany: Berlin, Greece: Athens, ..., Japan: ____

    Jiefeng Chen, Jinsung Yoon, Sayna Ebrahimi, Sercan Arik, Tomas Pfister, and Somesh Jha. Adaptation with self-evaluation to improve selective prediction in LLMs. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5190– 5213, Singapore, December 2023. Association for Computation...

  6. [6]

    Training Deep Nets with Sublinear Memory Cost

    Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost.arXiv preprint arXiv:1604.06174, 2016. URLhttps://arxiv.org/abs/1604.06174

  7. [7]

    C. K. Chow. On optimum recognition error and reject tradeoff.IEEE Trans. Inf. Theory, 16(1):41–46,

  8. [8]

    doi: 10.1109/TIT.1970.1054406

  9. [9]

    Training Verifiers to Solve Math Word Problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

  10. [10]

    Boosting with abstention

    Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri. Boosting with abstention. In NeurIPS, 2016. URL https://proceedings.neurips.cc/paper_files/paper/2016/file/ 7634ea65a4e6d9041cfd3f7de18e334a-Paper.pdf

  11. [11]

    Improving selective visual question answering by learning from your peers.arXiv preprint arXiv:2306.08751, 2023

    Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, and Marcus Rohrbach. Improving selective visual question answering by learning from your peers.arXiv preprint arXiv:2306.08751, 2023. URL https://arxiv.org/abs/ 2306.08751

  12. [12]

    FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

    Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning.arXiv preprint arXiv:2307.08691, 2023. URLhttps://arxiv.org/abs/2307.08691

  13. [13]

    Be my ai.https://www.bemyeyes.com/be-my-ai, 2026

    Be My Eyes. Be my ai.https://www.bemyeyes.com/be-my-ai, 2026. Accessed on 03/05/2026

  14. [14]

    Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

    Xingyu Fu, Yushi Hu, Ranjay Krishna, Mari Ostendorf, Dan Roth, Weijia Shi, Noah Smith, and Luke Zettlemoyer. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. InNeurIPS, 2024. doi: 10.52202/079017-4423. URL https://doi.org/10.52202/079017-4423. NeurIPS 2024

  15. [15]

    Selective classification for deep neural networks

    Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. InNeurIPS, 2017. 10

  16. [16]

    Selectivenet: A deep neural network with an integrated reject option

    Yonatan Geifman and Ran El-Yaniv. Selectivenet: A deep neural network with an integrated reject option. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors,Proceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 2151–2159. PMLR, 09–15 Jun 2019. URLhttps://proceedings.mlr.press/v...

  17. [17]

    Lookout: Assisted vision

    Google. Lookout: Assisted vision. https://support.google.com/accessibility/android/ answer/9031274, 2026. Accessed on 03/05/2026

  18. [18]

    In2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. InCVPR, 2017. doi: 10.1109/CVPR.2017.670

  19. [19]

    arXiv preprint arXiv:2405.02917 , year =

    Tobias Groot and Matias Valdenegro-Toro. Overconfidence is key: Verbalized uncertainty evaluation in large language and vision-language models, 2024. URLhttps://arxiv.org/abs/2405.02917

  20. [20]

    Div8k: Diverse 8k resolution image dataset.2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pages 3512–3516, 2019

    Shuhang Gu, Andreas Lugmayr, Martin Danelljan, Manuel Fritsche, Julien Lamour, and Radu Timofte. Div8k: Diverse 8k resolution image dataset.2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pages 3512–3516, 2019

  21. [21]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025

  22. [22]

    Vizwiz grand challenge: Answering visual questions from blind people

    Danna Gurari et al. Vizwiz grand challenge: Answering visual questions from blind people. InCVPR,

  23. [23]

    doi: 10.1109/CVPR.2018.00380

  24. [24]

    LoRA: Low-Rank Adaptation of Large Language Models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models.arXiv preprint arXiv:2106.09685, 10, 2021

  25. [25]

    Verbalized confidence triggers self-verification : Emergent behavior without explicit reasoning supervision

    Chaeyun Jang, Moonseok Choi, Yegon Kim, Hyungi Lee, and Juho Lee. Verbalized confidence triggers self-verification : Emergent behavior without explicit reasoning supervision. InICML 2025 Workshop on Reliable and Responsible F oundation Models, 2025. URL https://openreview.net/forum?id= Ub3eXwQ0uA

  26. [26]

    Why Language Models Hallucinate

    Adam Tauman Kalai, Ofir Nachum, Santosh S Vempala, and Edwin Zhang. Why language models hallucinate.arXiv preprint arXiv:2509.04664, 2025

  27. [27]

    Cross-dimension affinity distillation for 3d em neuron segmentation,

    Zaid Khan and Yun Fu. Consistency and uncertainty: Identifying unreliable responses from black-box vision-language models for selective visual question answering. InCVPR, pages 10854–10863, 2024. doi: 10.1109/CVPR52733.2024.01032

  28. [28]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023

  29. [29]

    Let’s verify step by step

    Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. InThe Twelfth International Conference on Learning Representations, 2024

  30. [30]

    Intelligent document processing market analysis 2025-2028: Idp at the crossroads

    Dan Lucarini and Alan Pelz-Sharpe. Intelligent document processing market analysis 2025-2028: Idp at the crossroads. Technical report, Deep Analysis, 2025. URL https://www.deep-analysis.net/ intelligent-document-processing-market-analysis-2025-2028/

  31. [31]

    Spatial recall index for machine learning algorithms

    Patrick Müller, Mattis Brummel, and Alexander Braun. Spatial recall index for machine learning algorithms. In Graham D. Finlayson and Sophie Triantaphillidou, editors,London Imaging Meeting 2021: Imaging for Deep Learning, LIM 2021, online, September 20-22, 2021, pages 58–62. Society for Imaging Science and Technology, 2021. doi: 10.2352/ISSN.2694-118X.20...

  32. [32]

    Harmony: Hidden activation representations and model output- aware uncertainty estimation for vision-language models

    Erum Mushtaq, Zalan Fabian, Yavuz Faruk Bakman, Anil Ramakrishna, Mahdi Soltanolkotabi, and Salman Avestimehr. Harmony: Hidden activation representations and model output- aware uncertainty estimation for vision-language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025. URL https://openacc...

  33. [33]

    Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf, 2025

    OpenAI. Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf, 2025. 11

  34. [34]

    Openai o3 and o4-mini system card

    OpenAI. Openai o3 and o4-mini system card. https://cdn.openai.com/pdf/ 2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf, 2025

  35. [35]

    Thinking with images.https://openai.com/index/thinking-with-images/, 2025

    OpenAI. Thinking with images.https://openai.com/index/thinking-with-images/, 2025

  36. [36]

    Human-adversarial visual question answering, 2021

    Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Alberto Lopez Magana, Wojciech Galuba, Devi Parikh, and Douwe Kiela. Human-adversarial visual question answering, 2021. URL https: //arxiv.org/abs/2106.02280

  37. [37]

    Real open-ended track leaderboard results, vqa challenge 2021.https://visualqa.org/roe.html, 2021

    Ayush Shrivastava, Yash Goyal, Dhruv Batra, Devi Parikh, and Aishwarya Agrawal. Real open-ended track leaderboard results, vqa challenge 2021.https://visualqa.org/roe.html, 2021

  38. [38]

    , Bansal, H

    Nishad Singhi, Hritik Bansal, Arian Hosseini, Aditya Grover, Kai-Wei Chang, Marcus Rohrbach, and Anna Rohrbach. When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning, 2025. URLhttps://arxiv.org/abs/2504.01005

  39. [39]

    selective prediction

    Tejas Srinivasan, Jack Hessel, Tanmay Gupta, Bill Yuchen Lin, Yejin Choi, Jesse Thomason, and Khyathi Chandu. Selective “selective prediction”: Reducing unnecessary abstention in vision-language reasoning. InFindings of the Association for Computational Linguistics: ACL 2024, pages 12935–12948, 2024. doi: 10.18653/v1/2024.findings-acl.767. URLhttps://acla...

  40. [40]

    Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning, May 2025

    Alex Su, Haozhe Wang, Weimin Ren, Fangzhen Lin, and Wenhu Chen. Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning, May 2025

  41. [41]

    Gemma 3 Technical Report

    Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean-bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Bey...

  42. [42]

    Post-abstention: Towards reliably re-attempting the abstained instances in QA

    Neeraj Varshney and Chitta Baral. Post-abstention: Towards reliably re-attempting the abstained instances in QA. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), pages 967–982, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.55. URLhttp...

  43. [43]

    40 intelligent document processing statistics

    Megon Venter. 40 intelligent document processing statistics. https://www.pdfreaderpro.com/blog/ document-processing-statistics, 2025

  44. [44]

    Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.arXiv preprint arXiv:2408.15556, 2024

    Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, and Dacheng Tao. Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models, 2024. URLhttps://arxiv.org/abs/2408.15556

  45. [45]

    Transactions of the Association for Computational Linguistics13, 529–556 (2025).https://doi.org/ 10.1162/tacl_a_00754,https://aclanthology.org/2025.tacl-1.26/

    Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang. Know your limits: A survey of abstention in large language models.Transactions of the Association for Computational Linguistics, 13:529–556, 2025. doi: 10.1162/tacl_a_00754. URL https://direct.mit. edu/tacl/article/doi/10.1162/tacl_a_00754

  46. [46]

    Reliable visual question answering: Abstain rather than answer incorrectly

    Spencer Whitehead, Suzanne Petryk, Vedaad Shakib, Joseph Gonzalez, Trevor Darrell, Anna Rohrbach, and Marcus Rohrbach. Reliable visual question answering: Abstain rather than answer incorrectly. InComputer Vision – ECCV 2022 Workshops, pages 148–166. Springer, Cham, 2022. doi: 10.1007/978-3-031-20059-5_ 9

  47. [47]

    V*: Guided visual search as a core mechanism in multimodal llms

    Penghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. InCVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  48. [48]

    The art of abstention: Selective prediction and error regularization for natural language processing

    Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. The art of abstention: Selective prediction and error regularization for natural language processing. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021. doi: 10.18653/v1/2021.acl-long.84. URL https://aclanthology.org/2021.acl-long.84

  49. [49]

    Qwen3 Technical Report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

  50. [50]

    On verbalized confidence scores for LLMs

    Daniel Yang, Yao-Hung Hubert Tsai, and Makoto Yamada. On verbalized confidence scores for LLMs. In ICLR Workshop: Quantify Uncertainty and Hallucination in F oundation Models: The Next Frontier in Reliable AI, 2025. URLhttps://openreview.net/forum?id=CVRdNQvFPE. 12

  51. [51]

    Selective-LAMA: Selective prediction for confidence-aware evaluation of language models

    Hiyori Yoshikawa and Naoaki Okazaki. Selective-LAMA: Selective prediction for confidence-aware evaluation of language models. In Andreas Vlachos and Isabelle Augenstein, editors,Findings of the Association for Computational Linguistics: EACL 2023, pages 2017–2028, Dubrovnik, Croatia, May

  52. [52]

    doi: 10.18653/v1/2023.findings-eacl.150

    Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-eacl.150. URL https: //aclanthology.org/2023.findings-eacl.150/

  53. [53]

    MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

    Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin, and Tieniu Tan. Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?, 2024. URL https://arxiv.org/abs/2408.13257

  54. [54]

    Thyme: Think Beyond Images

    Yi-Fan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu, Wei Chen, Xiao Hu, Bin Wen, Kaiyu Jiang, Changyi Liu, Tianke Zhang, Haonan Fan, Kaibing Chen, Jiankang Chen, Haojie Ding, Kaiyu Tang, Zhang Zhang, Liang Wang, Fan Yang, Tingting Gao, and Guorui Zhou. Thyme: Think beyond images, 2025. URL https://arxiv.org/abs/2508.11630

  55. [55]

    LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

    Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. Llamafactory: Unified efficient fine-tuning of 100+ language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V olume 3: System Demonstrations), Bangkok, Thailand, 2024. Association for Computational Linguist...

  56. [56]

    Thinking with Images

    Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning, May 2025

  57. [57]

    Pooled” evaluates C@5 on the set of question-answer pairs aggregated over all repetitions. “Avg. per-rep

    Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. Towards complex document understanding by discrete reasoning. InACM MM, 2022. doi: 10.1145/3503161.3548422. URL https://doi.org/10.1145/3503161.3548422. 13 Supplementary Material We briefly describe how the appendix is organized. Section A defines the selective-prediction m...

  58. [58]

    Selected

    generated additional distractor options, until 4 answers per question are obtained. To generate multiple choices for the Thyme hold out set, we draw incorrect answers (hard negatives) from Pixel-Reasoner. If these are not enough to reach 4 choices (including the ground-truth), we generate distractor options with GPT-5. The exact prompt templates used for ...

  59. [59]

    Fine-grained Single-instance Perception

    It contains 800 questions on “Fine-grained Single-instance Perception” (aligns with V* Bench’s attribute recognition), and “Fine-grained Cross-instance Perception” (aligns with V* Bench’s spatial relationship reasoning). We use the 8K resolution version. The domain of the images is broader. Gathered from DIV8K [19], questions pertain not only natural imag...

  60. [60]

    Note you are not provided this final image, and only the crop, which the model should only use to give the final answer

    **Crop Sufficiency**: Is the provided image crop sufficient to support the model’s response? Does it contain all the necessary visual information referenced in the response? If the model explicitly states they use the global view to answer this question, you should consider this as not grounded in the prompt. Note you are not provided this final image, an...

  61. [61]

    red car" -> \boxed{Yes} - If the crop shows a partial view that doesn’t contain enough information to answer -> \boxed{No} - If the crop shows a dog but the model answers

    **Answer Coherence**: Is the model’s response coherent with what is actually visible in the image? Or is the model hallucinating information or obtaining it from elsewhere (not from the image)? Think step by step about both aspects, then provide your final assessment. Output your final decision as \boxed{Yes} if the answer is well-grounded in the image cr...

This paper was first reviewed by grok-4.3 on July 1, 2026.