REVIEW 2 major objections 1 minor 1 cited by
MCMit uses a constant-latency branch instruction and CNN discriminators to cut mid-circuit measurement errors in quantum circuits.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-07-01 08:33 UTC pith:VB2SH2EY
load-bearing objection MCMit combines a constant-latency branch instruction with CNN/transformer discriminators and software patches, but the accuracy and error-rate gains rest on offline trace tests rather than closed-loop runs with the new hardware. the 2 major comments →
MCMit: Mid-Circuit Measurement Error Mitigation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MCMit is a hardware-software co-design for mitigating branching and latency-induced errors in mid-circuit measurements that introduces a scalable constant-latency multi-control branch instruction for faster classical feedback, transformer and CNN qubit-state discriminators that maintain high accuracy at short measurement durations, and software techniques of static MCM elimination and stochastic branching to address remaining errors.
What carries the argument
The constant-latency multi-control branch instruction paired with CNN and transformer qubit-state discriminators.
Load-bearing premise
The CNN and transformer discriminators trained on extracted QPU traces will retain their accuracy gains when executed in real time on actual hardware controllers.
What would settle it
A side-by-side run of the same quantum error correction circuit on the same QPU with and without the MCMit branch instruction and discriminators, measuring the resulting logical error rates.
If this is right
- Feedback latency drops by up to 70 percent, supporting circuit depths up to seven times larger than the Qubic baseline.
- The CNN discriminator delivers up to 62 percent higher accuracy for short measurement windows.
- Logical error rates in quantum error correction drop by factors between 1.2 and 9.4.
- Software mitigation raises circuit fidelity by 18 to 30 percent compared with baseline methods.
Where Pith is reading between the lines
- Faster feedback could make mid-circuit measurements practical in larger-scale error-corrected algorithms that currently exceed coherence limits.
- The discriminators might be combined with existing calibration routines to further reduce the need for long measurement times.
- Testing the full stack on a physical controller would reveal whether the reported accuracy holds under real-time scheduling constraints.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MCMit, a hardware-software co-design for mitigating errors from mid-circuit measurements (MCMs) and classical feedback in dynamic quantum circuits for DQC and QEC. It introduces a constant-latency multi-control branch instruction on the Qubic controller, CNN and transformer qubit-state discriminators trained on experimental traces, and software techniques (static MCM elimination, stochastic branching). Evaluation on experimentally extracted QPU readout traces claims up to 70% feedback latency reduction (7× circuit depth improvement), up to 62% higher CNN accuracy at short durations (yielding 1.2×–9.4× lower QEC logical error rates), and 18–30% fidelity gains from software mitigation over baselines.
Significance. If the integrated performance claims hold, MCMit would meaningfully advance practical dynamic-circuit execution by jointly tackling latency and discrimination bottlenecks that currently limit QEC and DQC. The use of real QPU traces and controller implementation provides concrete grounding; the co-design framing and reported numerical gains on logical error rates are the primary contributions.
major comments (2)
- [Abstract / Evaluation] Abstract and Evaluation (results on discriminators): The headline claims that the CNN achieves up to 62% higher accuracy for short measurement durations and produces 1.2×–9.4× lower logical error rates rest on offline training/testing of the discriminator on static experimentally extracted traces. The manuscript provides no closed-loop experiments that combine the new constant-latency branch instruction with the CNN/transformer in real-time feedback; any timing skew or signal-path change introduced by the branch could alter the input statistics seen by the discriminator, directly undermining the reported accuracy and logical-error improvements.
- [Implementation / Evaluation] Implementation and Evaluation sections: The branch-instruction latency reduction (70%) and circuit-depth improvement (7×) are presented separately from the discriminator accuracy results. No integrated benchmark is shown that measures end-to-end logical error rates or fidelity when the new branch, controller firmware, and real-time discriminator are all active together, which is required to substantiate the holistic co-design claim.
minor comments (1)
- [Abstract] The abstract states results are obtained 'on Qubic' yet the discriminator numbers are explicitly from offline trace processing; a short clarifying sentence on the distinction between controller implementation and evaluation methodology would improve readability.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback emphasizing the importance of integrated evaluation for our hardware-software co-design claims. We agree that the current manuscript evaluates components separately using real QPU traces and controller measurements, without a single closed-loop benchmark. We will revise the abstract, evaluation sections, and add a limitations discussion to clarify the methodology, assumptions, and scope of the reported gains. Our point-by-point responses follow.
read point-by-point responses
-
Referee: [Abstract / Evaluation] Abstract and Evaluation (results on discriminators): The headline claims that the CNN achieves up to 62% higher accuracy for short measurement durations and produces 1.2×–9.4× lower logical error rates rest on offline training/testing of the discriminator on static experimentally extracted traces. The manuscript provides no closed-loop experiments that combine the new constant-latency branch instruction with the CNN/transformer in real-time feedback; any timing skew or signal-path change introduced by the branch could alter the input statistics seen by the discriminator, directly undermining the reported accuracy and logical-error improvements.
Authors: We acknowledge that the discriminator accuracy and logical-error results derive from offline training and testing on static traces extracted from the QPU, while the branch-instruction latency is measured independently on the Qubic controller. The manuscript does not include closed-loop real-time experiments integrating both. The traces reflect actual hardware noise and readout characteristics; the branch instruction operates on the classical feedback path after discrimination and does not modify the analog signal path or measurement duration seen by the discriminator. Nevertheless, the referee's point is valid regarding potential unmodeled interactions. In revision we will (i) explicitly qualify the abstract and evaluation claims as component-wise results, (ii) add text stating the assumption that reduced classical latency does not alter discriminator input statistics, and (iii) include a limitations paragraph noting the value of future closed-loop validation. revision: partial
-
Referee: [Implementation / Evaluation] Implementation and Evaluation sections: The branch-instruction latency reduction (70%) and circuit-depth improvement (7×) are presented separately from the discriminator accuracy results. No integrated benchmark is shown that measures end-to-end logical error rates or fidelity when the new branch, controller firmware, and real-time discriminator are all active together, which is required to substantiate the holistic co-design claim.
Authors: The evaluations are intentionally modular: latency and depth improvements are quantified via controller implementation and trace-driven circuit simulation, while discriminator performance uses the same experimental traces for training and inference accuracy. No single end-to-end benchmark combining live branch execution, firmware, and real-time discrimination is reported. This separation follows from the distinct experimental setups (hardware controller timing versus offline trace processing). We agree this limits the strength of the integrated co-design narrative. In the revision we will reorganize the evaluation section to present the components as complementary parts of the co-design, add an explicit statement that end-to-end logical-error measurements under simultaneous operation remain future work, and temper the abstract accordingly while retaining the individual quantitative results. revision: partial
Circularity Check
No circularity; claims rest on direct experimental evaluation rather than fitted or self-referential derivations.
full rationale
The paper presents its core results—70% latency reduction, 62% accuracy gain for the CNN discriminator, 1.2–9.4× logical-error improvement, and 18–30% fidelity gain—as outcomes of implementing the branch instruction, CNN/transformer discriminators, and software mitigation on Qubic hardware and measuring them against experimentally extracted QPU readout traces. No equations, parameter fits, or self-citations are shown that would reduce any reported gain to an input by construction; the evaluation chain remains external to the claimed improvements.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of MCMit: Mid-Circuit Measurement Error Mitigation." pith.science (2026). https://pith.science/paper/VB2SH2EY
@misc{pith2026260425863,
author = {Pith},
title = {Pith review of: MCMit: Mid-Circuit Measurement Error Mitigation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VB2SH2EY}},
note = {Machine review of arXiv:2604.25863}
}
read the original abstract
Distributed Quantum Computing (DQC) and Quantum Error Correction (QEC) rely on dynamic circuits that include Mid-Circuit Measurements (MCMs) and classical feedback. These operations present a major bottleneck: MCMs suffer from high error rates that lead to real-time branching errors, while MCM and classical feedback latencies amplify decoherence errors. Current hardware controllers, qubit-state discriminators, and software error mitigation techniques fail to address these challenges holistically. We propose MCMit, a hardware-software co-design to mitigate branching and latency-induced errors. MCMit introduces a scalable, constant-latency multi-control branch instruction for faster classical feedback and two qubit-state discriminators, a transformer, and a CNN, with high accuracy even under short measurement durations. On the software side, static MCM elimination and stochastic branching complement the hardware by mitigating residual branching errors that persist despite hardware improvements. We implement MCMit on Qubic and evaluate it using experimentally extracted QPU readout traces. Our branch instruction reduces feedback latency by up to 70%, improving circuit depths by up to $7\times$ over Qubic. Our CNN discriminator achieves up to 62% higher accuracy for short measurement durations than the baselines, leading to 1.2$\times$--9.4$\times$ lower logical error rates in QEC. Last, our software mitigation improves fidelity by 18--30% over baseline methods.
Figures
Forward citations
Cited by 1 Pith paper
-
Oraqle: An Empirical Analysis of Qubit Readout and Discriminators in Quantum Error Correction
Using real 5-qubit traces, this study shows readout windows can be cut to ~600 ns with negligible QEC penalty and small discriminators match large ones.
Reference graph
Works this paper leans on
-
[1]
System card: Claude opus 4 & claude sonnet 4
Anthropic. System card: Claude opus 4 & claude sonnet 4. https://www-cdn.anthropic.com/ 6d8a8055020700718b0c49369f60816ba2a7c285.pdf, 2025
work page 2025
-
[2]
Detection of Interest Points Using Symmetry
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. InICCV, 2015. doi: 10.1109/ICCV .2015.279
-
[3]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[4]
Large language models must be taught to know what they don’t know
Umang Bhatt, Katherine Collins, Samuel Dooley, Micah Goldblum, Nate Gruver, Sanyam Kapoor, Arka Pal, Manley Roberts, Adrian Weller, and Andrew Wilson. Large language models must be taught to know what they don’t know. InNeurIPS, 2024. doi: 10.52202/079017-2729. URL https://doi.org/10. 52202/079017-2729
-
[5]
Germany: Berlin, Greece: Athens, ..., Japan: ____
Jiefeng Chen, Jinsung Yoon, Sayna Ebrahimi, Sercan Arik, Tomas Pfister, and Somesh Jha. Adaptation with self-evaluation to improve selective prediction in LLMs. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5190– 5213, Singapore, December 2023. Association for Computation...
-
[6]
Training Deep Nets with Sublinear Memory Cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost.arXiv preprint arXiv:1604.06174, 2016. URLhttps://arxiv.org/abs/1604.06174
work page internal anchor Pith review Pith/arXiv arXiv 2016
-
[7]
C. K. Chow. On optimum recognition error and reject tradeoff.IEEE Trans. Inf. Theory, 16(1):41–46,
-
[8]
doi: 10.1109/TIT.1970.1054406
-
[9]
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[10]
Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri. Boosting with abstention. In NeurIPS, 2016. URL https://proceedings.neurips.cc/paper_files/paper/2016/file/ 7634ea65a4e6d9041cfd3f7de18e334a-Paper.pdf
work page 2016
-
[11]
Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, and Marcus Rohrbach. Improving selective visual question answering by learning from your peers.arXiv preprint arXiv:2306.08751, 2023. URL https://arxiv.org/abs/ 2306.08751
-
[12]
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning.arXiv preprint arXiv:2307.08691, 2023. URLhttps://arxiv.org/abs/2307.08691
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[13]
Be my ai.https://www.bemyeyes.com/be-my-ai, 2026
Be My Eyes. Be my ai.https://www.bemyeyes.com/be-my-ai, 2026. Accessed on 03/05/2026
work page 2026
-
[14]
Visual sketchpad: Sketching as a visual chain of thought for multimodal language models
Xingyu Fu, Yushi Hu, Ranjay Krishna, Mari Ostendorf, Dan Roth, Weijia Shi, Noah Smith, and Luke Zettlemoyer. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. InNeurIPS, 2024. doi: 10.52202/079017-4423. URL https://doi.org/10.52202/079017-4423. NeurIPS 2024
-
[15]
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. InNeurIPS, 2017. 10
work page 2017
-
[16]
Selectivenet: A deep neural network with an integrated reject option
Yonatan Geifman and Ran El-Yaniv. Selectivenet: A deep neural network with an integrated reject option. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors,Proceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 2151–2159. PMLR, 09–15 Jun 2019. URLhttps://proceedings.mlr.press/v...
work page 2019
-
[17]
Google. Lookout: Assisted vision. https://support.google.com/accessibility/android/ answer/9031274, 2026. Accessed on 03/05/2026
-
[18]
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. InCVPR, 2017. doi: 10.1109/CVPR.2017.670
-
[19]
arXiv preprint arXiv:2405.02917 , year =
Tobias Groot and Matias Valdenegro-Toro. Overconfidence is key: Verbalized uncertainty evaluation in large language and vision-language models, 2024. URLhttps://arxiv.org/abs/2405.02917
-
[20]
Shuhang Gu, Andreas Lugmayr, Martin Danelljan, Manuel Fritsche, Julien Lamour, and Radu Timofte. Div8k: Diverse 8k resolution image dataset.2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pages 3512–3516, 2019
work page 2019
-
[21]
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[22]
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari et al. Vizwiz grand challenge: Answering visual questions from blind people. InCVPR,
-
[23]
doi: 10.1109/CVPR.2018.00380
-
[24]
LoRA: Low-Rank Adaptation of Large Language Models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models.arXiv preprint arXiv:2106.09685, 10, 2021
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[25]
Chaeyun Jang, Moonseok Choi, Yegon Kim, Hyungi Lee, and Juho Lee. Verbalized confidence triggers self-verification : Emergent behavior without explicit reasoning supervision. InICML 2025 Workshop on Reliable and Responsible F oundation Models, 2025. URL https://openreview.net/forum?id= Ub3eXwQ0uA
work page 2025
-
[26]
Why Language Models Hallucinate
Adam Tauman Kalai, Ofir Nachum, Santosh S Vempala, and Edwin Zhang. Why language models hallucinate.arXiv preprint arXiv:2509.04664, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[27]
Cross-dimension affinity distillation for 3d em neuron segmentation,
Zaid Khan and Yun Fu. Consistency and uncertainty: Identifying unreliable responses from black-box vision-language models for selective visual question answering. InCVPR, pages 10854–10863, 2024. doi: 10.1109/CVPR52733.2024.01032
-
[28]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023
work page 2023
-
[29]
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. InThe Twelfth International Conference on Learning Representations, 2024
work page 2024
-
[30]
Intelligent document processing market analysis 2025-2028: Idp at the crossroads
Dan Lucarini and Alan Pelz-Sharpe. Intelligent document processing market analysis 2025-2028: Idp at the crossroads. Technical report, Deep Analysis, 2025. URL https://www.deep-analysis.net/ intelligent-document-processing-market-analysis-2025-2028/
work page 2025
-
[31]
Spatial recall index for machine learning algorithms
Patrick Müller, Mattis Brummel, and Alexander Braun. Spatial recall index for machine learning algorithms. In Graham D. Finlayson and Sophie Triantaphillidou, editors,London Imaging Meeting 2021: Imaging for Deep Learning, LIM 2021, online, September 20-22, 2021, pages 58–62. Society for Imaging Science and Technology, 2021. doi: 10.2352/ISSN.2694-118X.20...
-
[32]
Erum Mushtaq, Zalan Fabian, Yavuz Faruk Bakman, Anil Ramakrishna, Mahdi Soltanolkotabi, and Salman Avestimehr. Harmony: Hidden activation representations and model output- aware uncertainty estimation for vision-language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025. URL https://openacc...
work page 2025
-
[33]
Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf, 2025
OpenAI. Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf, 2025. 11
work page 2025
-
[34]
Openai o3 and o4-mini system card
OpenAI. Openai o3 and o4-mini system card. https://cdn.openai.com/pdf/ 2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf, 2025
work page 2025
-
[35]
Thinking with images.https://openai.com/index/thinking-with-images/, 2025
OpenAI. Thinking with images.https://openai.com/index/thinking-with-images/, 2025
work page 2025
-
[36]
Human-adversarial visual question answering, 2021
Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Alberto Lopez Magana, Wojciech Galuba, Devi Parikh, and Douwe Kiela. Human-adversarial visual question answering, 2021. URL https: //arxiv.org/abs/2106.02280
-
[37]
Real open-ended track leaderboard results, vqa challenge 2021.https://visualqa.org/roe.html, 2021
Ayush Shrivastava, Yash Goyal, Dhruv Batra, Devi Parikh, and Aishwarya Agrawal. Real open-ended track leaderboard results, vqa challenge 2021.https://visualqa.org/roe.html, 2021
work page 2021
-
[38]
Nishad Singhi, Hritik Bansal, Arian Hosseini, Aditya Grover, Kai-Wei Chang, Marcus Rohrbach, and Anna Rohrbach. When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning, 2025. URLhttps://arxiv.org/abs/2504.01005
-
[39]
Tejas Srinivasan, Jack Hessel, Tanmay Gupta, Bill Yuchen Lin, Yejin Choi, Jesse Thomason, and Khyathi Chandu. Selective “selective prediction”: Reducing unnecessary abstention in vision-language reasoning. InFindings of the Association for Computational Linguistics: ACL 2024, pages 12935–12948, 2024. doi: 10.18653/v1/2024.findings-acl.767. URLhttps://acla...
-
[40]
Alex Su, Haozhe Wang, Weimin Ren, Fangzhen Lin, and Wenhu Chen. Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning, May 2025
work page 2025
-
[41]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean-bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Bey...
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[42]
Post-abstention: Towards reliably re-attempting the abstained instances in QA
Neeraj Varshney and Chitta Baral. Post-abstention: Towards reliably re-attempting the abstained instances in QA. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), pages 967–982, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.55. URLhttp...
-
[43]
40 intelligent document processing statistics
Megon Venter. 40 intelligent document processing statistics. https://www.pdfreaderpro.com/blog/ document-processing-statistics, 2025
work page 2025
-
[44]
Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, and Dacheng Tao. Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models, 2024. URLhttps://arxiv.org/abs/2408.15556
-
[45]
Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang. Know your limits: A survey of abstention in large language models.Transactions of the Association for Computational Linguistics, 13:529–556, 2025. doi: 10.1162/tacl_a_00754. URL https://direct.mit. edu/tacl/article/doi/10.1162/tacl_a_00754
-
[46]
Reliable visual question answering: Abstain rather than answer incorrectly
Spencer Whitehead, Suzanne Petryk, Vedaad Shakib, Joseph Gonzalez, Trevor Darrell, Anna Rohrbach, and Marcus Rohrbach. Reliable visual question answering: Abstain rather than answer incorrectly. InComputer Vision – ECCV 2022 Workshops, pages 148–166. Springer, Cham, 2022. doi: 10.1007/978-3-031-20059-5_ 9
-
[47]
V*: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. InCVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
work page 2024
-
[48]
The art of abstention: Selective prediction and error regularization for natural language processing
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. The art of abstention: Selective prediction and error regularization for natural language processing. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021. doi: 10.18653/v1/2021.acl-long.84. URL https://aclanthology.org/2021.acl-long.84
-
[49]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[50]
On verbalized confidence scores for LLMs
Daniel Yang, Yao-Hung Hubert Tsai, and Makoto Yamada. On verbalized confidence scores for LLMs. In ICLR Workshop: Quantify Uncertainty and Hallucination in F oundation Models: The Next Frontier in Reliable AI, 2025. URLhttps://openreview.net/forum?id=CVRdNQvFPE. 12
work page 2025
-
[51]
Selective-LAMA: Selective prediction for confidence-aware evaluation of language models
Hiyori Yoshikawa and Naoaki Okazaki. Selective-LAMA: Selective prediction for confidence-aware evaluation of language models. In Andreas Vlachos and Isabelle Augenstein, editors,Findings of the Association for Computational Linguistics: EACL 2023, pages 2017–2028, Dubrovnik, Croatia, May
work page 2023
-
[52]
doi: 10.18653/v1/2023.findings-eacl.150
Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-eacl.150. URL https: //aclanthology.org/2023.findings-eacl.150/
-
[53]
Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin, and Tieniu Tan. Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?, 2024. URL https://arxiv.org/abs/2408.13257
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[54]
Yi-Fan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu, Wei Chen, Xiao Hu, Bin Wen, Kaiyu Jiang, Changyi Liu, Tianke Zhang, Haonan Fan, Kaibing Chen, Jiankang Chen, Haojie Ding, Kaiyu Tang, Zhang Zhang, Liang Wang, Fan Yang, Tingting Gao, and Guorui Zhou. Thyme: Think beyond images, 2025. URL https://arxiv.org/abs/2508.11630
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[55]
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. Llamafactory: Unified efficient fine-tuning of 100+ language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V olume 3: System Demonstrations), Bangkok, Thailand, 2024. Association for Computational Linguist...
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[56]
Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning, May 2025
work page 2025
-
[57]
Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. Towards complex document understanding by discrete reasoning. InACM MM, 2022. doi: 10.1145/3503161.3548422. URL https://doi.org/10.1145/3503161.3548422. 13 Supplementary Material We briefly describe how the appendix is organized. Section A defines the selective-prediction m...
-
[58]
generated additional distractor options, until 4 answers per question are obtained. To generate multiple choices for the Thyme hold out set, we draw incorrect answers (hard negatives) from Pixel-Reasoner. If these are not enough to reach 4 choices (including the ground-truth), we generate distractor options with GPT-5. The exact prompt templates used for ...
-
[59]
Fine-grained Single-instance Perception
It contains 800 questions on “Fine-grained Single-instance Perception” (aligns with V* Bench’s attribute recognition), and “Fine-grained Cross-instance Perception” (aligns with V* Bench’s spatial relationship reasoning). We use the 8K resolution version. The domain of the images is broader. Gathered from DIV8K [19], questions pertain not only natural imag...
-
[60]
**Crop Sufficiency**: Is the provided image crop sufficient to support the model’s response? Does it contain all the necessary visual information referenced in the response? If the model explicitly states they use the global view to answer this question, you should consider this as not grounded in the prompt. Note you are not provided this final image, an...
-
[61]
**Answer Coherence**: Is the model’s response coherent with what is actually visible in the image? Or is the model hallucinating information or obtaining it from elsewhere (not from the image)? Think step by step about both aspects, then provide your final assessment. Output your final decision as \boxed{Yes} if the answer is well-grounded in the image cr...
This paper was first reviewed by grok-4.3 on July 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.