REVIEW 1 major objections 42 references
ZK-Tracer: A High-Performance Heterogeneous Accelerator for Zero-Knowledge VM Trace Generation
T0 review · 1 major / 0 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read A heterogeneous ASIC called ZK-Tracer speeds zkVM trace generation by up to 1829 times over a high-performance CPU.
desk verdict The abstract claims 1829x trace generation speedup and 963x end-to-end but supplies zero details on baselines, workloads, or measurement method. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Heterogeneous accelerator built from one Main Trace Unit plus parallel Permutation Trace Units, controlled through a lightweight instruction-set extension for host offloading.
What would settle it
A benchmark run on current zkVM workloads that shows trace generation occupying less than 10 percent of total ZKP runtime, or an ASIC measurement that fails to exceed 100x speedup in trace generation.
Extended reading notes
Core claim
ZK-Tracer is the first accelerator architecture built specifically for the zkVM frontend. Its heterogeneous layout consists of a Main Trace Unit handling core execution traces and parallel Permutation Trace Units handling the permutation traces required by the proof system. The unit exposes a fine-grained interface via a lightweight instruction set extension so that software can offload trace tasks efficiently. Fabricated ASIC results establish the 1829x trace-generation speedup over multi-core CPU and the 963x end-to-end system gain when the accelerator is combined with existing proving hardware.
Load-bearing premise
The frontend execution and trace generation phase has become the dominant performance bottleneck in zkVM systems.
Editorial extensions
If this is right
- Trace generation ceases to limit zkVM throughput once the accelerator is present.
- End-to-end ZKP pipelines achieve roughly three orders of magnitude higher performance when the frontend accelerator is added to existing backend hardware.
- zkVMs become practical for workloads whose size was previously ruled out by frontend cost.
- Software can be restructured around the new offload interface to keep the accelerator fed.
Reading between the lines
- Similar trace-generation hardware could be adapted to other proof systems that rely on execution traces.
- The reported speedups assume the host CPU remains the source of program execution; a fully integrated design might change the bottleneck again.
- The 963x end-to-end figure depends on the relative sizes of frontend and backend workloads in a given application.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes ZK-Tracer, the first hardware accelerator architecture for the zkVM frontend trace generation phase. It introduces a heterogeneous design with a Main Trace Unit and parallel Permutation Trace Units, exposed to host software via a lightweight instruction set extension for task offloading. ASIC implementation results are claimed to deliver up to 1829× speedup in trace generation versus a high-performance multi-core CPU and 963× end-to-end improvement for the full ZKP system when paired with existing backend accelerators.
Significance. If the reported speedups can be substantiated with complete benchmark descriptions, workloads, baselines, and measurement details, the work would be significant for identifying and addressing the frontend as an emerging bottleneck in zkVMs. The heterogeneous architecture and integration approach, if validated, could enable practical system-level gains in zero-knowledge proof generation.
major comments (1)
- [Abstract] Abstract: the central performance claims (1829× trace-generation speedup and 963× end-to-end improvement) are stated without any description of benchmarks, workloads, measurement methodology, error bars, comparison baselines, ASIC process node, or host offload/data-movement costs. This absence renders the primary empirical results unevaluable and is load-bearing for the paper's contribution.
Simulated Author's Rebuttal
We thank the referee for the constructive comment on the abstract. We agree that additional context is needed to make the central claims evaluable and will revise the manuscript to address this.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central performance claims (1829× trace-generation speedup and 963× end-to-end improvement) are stated without any description of benchmarks, workloads, measurement methodology, error bars, comparison baselines, ASIC process node, or host offload/data-movement costs. This absence renders the primary empirical results unevaluable and is load-bearing for the paper's contribution.
Authors: We agree the abstract is too terse on evaluation details. In the revised manuscript we will expand the abstract to briefly specify: the workloads (representative zkVM programs including Fibonacci, sorting, and SHA256 circuits), the CPU baseline (high-performance 64-core Xeon-class processor), the ASIC process node, and that reported end-to-end figures include host offload and data-movement overheads. Full methodology, run counts, and any variance will continue to appear in the evaluation section. This change directly addresses the evaluability concern while preserving abstract length. revision: yes
Circularity Check
No circularity: empirical hardware benchmarks only
full rationale
The paper reports measured ASIC performance results (1829x trace-generation speedup and 963x end-to-end) against external CPU and backend baselines. No equations, derivations, fitted parameters, predictions, or first-principles claims appear in the abstract or described architecture. The load-bearing steps are direct empirical measurements, not reductions to self-defined quantities or self-citations, so the result is self-contained against external benchmarks.
Assumptions & free parameters
Cite this review
Pith. "Pith review of ZK-Tracer: A High-Performance Heterogeneous Accelerator for Zero-Knowledge VM Trace Generation." pith.science (2026). https://pith.science/paper/4Q4VPTYB
@misc{pith2026260525493,
author = {Pith},
title = {Pith review of: ZK-Tracer: A High-Performance Heterogeneous Accelerator for Zero-Knowledge VM Trace Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/4Q4VPTYB}},
note = {Machine review of arXiv:2605.25493}
}
read the original abstract
Zero-knowledge virtual machines (zkVMs) are a key technology for driving the large-scale adoption of zero-knowledge proofs (ZKP), but their performance bottlenecks severely limit their practicality. While current hardware acceleration research has exclusively focused on backend proving, we identify that the frontend execution and trace generation phase is rapidly emerging as the new system bottleneck. To address this challenge, we propose ZK-Tracer, the first hardware accelerator architecture specifically designed for the zkVM frontend. ZK-Tracer features a novel heterogeneous design comprising a Main Trace Unit and parallel Permutation Trace Units. It exposes a fine-grained interface to the host software through a lightweight instruction set extension, enabling efficient task offloading. Our ASIC implementation results demonstrate that ZK-Tracer achieves up to 1829x speedup in trace generation over a high-performance multi-core CPU. When integrated with existing backend proving accelerators, it delivers a remarkable 963x end-to-end performance improvement for the entire ZKP system.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Anees Ahmed, Nojan Sheybani, Davi Moreno, Nges Brian Njungle, Tengkai Gong, Michel Kinsy, and Farinaz Koushanfar. 2024. Amaze: Accelerated mimc hardware architecture for zero-knowledge applications on the edge. InProceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design. 1–9
2024
-
[2]
Bing-Jyue Chen, Suppakit Waiwitlikhit, Ion Stoica, and Daniel Kang. 2024. Zkml: An optimizing system for ml inference in zero-knowledge proofs. InProceedings of the Nineteenth European Conference on Computer Systems. 560–574
2024
-
[3]
Alhad Daftardar, Jianqiao Mo, Joey Ah-kiow, Benedikt Bünz, Ramesh Karri, Siddharth Garg, and Brandon Reagen. 2025. Need for zkSpeed: Accelerating HyperPlonk for Zero-Knowledge Proofs. InProceedings of the 52nd Annual Inter- national Symposium on Computer Architecture. 1986–2001
2025
-
[4]
Alhad Daftardar, Brandon Reagen, and Siddharth Garg. 2024. Szkp: A scalable accelerator architecture for zero-knowledge proofs. InProceedings of the 2024 International Conference on Parallel Architectures and Compilation Techniques. 271–283
2024
-
[5]
George Danezis, Cedric Fournet, Markulf Kohlweiss, and Bryan Parno. 2013. Pinocchio coin: building zerocoin from a succinct pairing-based proof system. In Proceedings of the First ACM workshop on Language support for privacy-enhancing technologies. 27–30
2013
-
[6]
Antoine Delignat-Lavaud, Cédric Fournet, Markulf Kohlweiss, and Bryan Parno
-
[7]
509 certificates into elegant anonymous credentials with the magic of verifiable computation
Cinderella: Turning shabby X. 509 certificates into elegant anonymous credentials with the magic of verifiable computation. In2016 IEEE Symposium on Security and Privacy (SP). IEEE, 235–254
-
[9]
Boyuan Feng, Zheng Wang, Yuke Wang, Shu Yang, and Yufei Ding. 2024. Zeno: A type-based optimization framework for zero knowledge neural network infer- ence. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1. 450–464
2024
Show all 42 references
-
[10]
gem5. 2024. gem5. https://github.com/gem5/gem5. Version 24.1 Accessed: 2025-10-10
2024
-
[11]
Shafi Goldwasser, Silvio Micali, and Chales Rackoff. 2019. The knowledge com- plexity of interactive proof-systems. InProviding sound foundations for cryptog- raphy: On the work of shafi goldwasser and silvio micali. 203–225
2019
-
[12]
Florian Hirner, Ahmet Can Mert, and Sujoy Sinha Roy. 2024. Proteus: A pipelined ntt architecture generator.IEEE Transactions on Very Large Scale Integration (VLSI) Systems32, 7 (2024), 1228–1238
2024
-
[13]
Zhuoran Ji, Zhiyuan Zhang, Jiming Xu, and Lei Ju. 2024. Accelerating multi- scalar multiplication for efficient zero knowledge proofs with multi-gpu systems. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating...
2024
-
[14]
Zhuoran Ji, Jianyu Zhao, Peimin Gao, Xiangkai Yin, and Lei Ju. 2025. Accelerat- ing Number Theoretic Transform with Multi-GPU Systems for Efficient Zero Knowledge Proof. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages a...
2025
-
[15]
Ahmed Kosba, Andrew Miller, Elaine Shi, Zikai Wen, and Charalampos Papa- manthou. 2016. Hawk: The blockchain model of cryptography and privacy- preserving smart contracts. In2016 IEEE symposium on security and privacy (SP). IEEE, 839–858
2016
-
[16]
Shang Li, Zhiyuan Yang, Dhiraj Reddy, Ankur Srivastava, and Bruce Jacob. 2020. DRAMsim3: A cycle-accurate, thermal-capable DRAM simulator.IEEE Computer Architecture Letters19, 2 (2020), 106–109
2020
-
[17]
Changxu Liu, Hao Zhou, Lan Yang, Yifei Feng, Zheng Wu, Zhuoyuan Yang, Yinlong Li, Shiyong Wu, and Fan Yang. 2025. AcclMT: A Highly Resource- Efficient and Flexible Poseidon Hash-Based Merkle Tree Architecture. In2025 62nd ACM/IEEE Design Automation Conference (DAC). IEEE, 1–7
2025
-
[18]
Changxu Liu, Hao Zhou, Lan Yang, Zheng Wu, Patrick Dai, Yinlong Li, Shiyong Wu, and Fan Yang. 2024. Myosotis: An Efficiently Pipelined and Parameterized Multi-Scalar Multiplication Architecture via Data Sharing.IEEE Transactions on Computer-Aided Design of Integrated Circuits ...
2024
-
[19]
Changxu Liu, Hao Zhou, Lan Yang, Jiamin Xu, Patrick Dai, and Fan Yang. 2024. Gypsophila: A scalable and bandwidth-optimized multi-scalar multiplication architecture. InProceedings of the 61st ACM/IEEE Design Automation Conference. 1–6
2024
-
[20]
Tao Lu, Yuxun Chen, Zonghui Wang, Xiaohang Wang, Wenzhi Chen, and Ji- aheng Zhang. 2025. BatchZK: A Fully Pipelined GPU-Accelerated System for Batch Generation of Zero-Knowledge Proofs. InProceedings of the 30th ACM International Conference on Architectural Support for Program...
2025
-
[21]
Tao Lu, Chengkun Wei, Ruijing Yu, Chaochao Chen, Wenjing Fang, Lei Wang, Zeke Wang, and Wenzhi Chen. 2023. Cuzk: Accelerating zero-knowledge proof with a faster parallel multi-scalar multiplication algorithm on gpus.IACR Transac- tions on Cryptographic Hardware and Embedded Sy...
2023
-
[22]
Weiliang Ma, Qian Xiong, Xuanhua Shi, Xiaosong Ma, Hai Jin, Haozhao Kuang, Mingyu Gao, Ye Zhang, Haichen Shen, and Weifang Hu. 2023. Gzkp: A gpu accelerated zero-knowledge proof system. InProceedings of the 28th ACM Inter- national Conference on Architectural Support for Progr...
2023
-
[23]
Bjorn Oude Roelink, Mohammed El-Hajj, and Dipti Sarmah. 2024. Systematic review: Comparing zk-SNARK, zk-STARK, and bulletproof protocols for privacy- preserving authentication.Security and Privacy7, 5 (2024), e401
2024
-
[24]
Pengcheng Qiu, Guiming Wu, Tingqiang Chu, Changzheng Wei, Runzhou Luo, Ying Yan, Wei Wang, and Hui Zhang. 2024. MSMAC: Accelerating Multi-Scalar Multiplication for Zero-Knowledge Proof. InProceedings of the 61st ACM/IEEE Design Automation Conference. 1–6
2024
-
[25]
Andy Ray, Benjamin Devlin, Fu Yong Quah, and Rahul Yesantharao. 2024. Hard- caml msm: A high-performance split cpu-fpga multi-scalar multiplication engine. InProceedings of the 2024 ACM/SIGDA International Symposium on Field Pro- grammable Gate Arrays. 33–39
2024
-
[26]
RISC Zero. 2021. risc0. https://github.com/risc0/risc0. Accessed: 2025-11-16
2021
-
[27]
Nikola Samardzic, Simon Langowski, Srinivas Devadas, and Daniel Sanchez. 2024. Accelerating zero-knowledge proofs through hardware-algorithm co-design. In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 366–379
2024
-
[28]
Eli Ben Sasson, Alessandro Chiesa, Christina Garman, Matthew Green, Ian Miers, Eran Tromer, and Madars Virza. 2014. Zerocash: Decentralized anonymous payments from bitcoin. In2014 IEEE symposium on security and privacy. IEEE, 459–474
2014
-
[29]
succinctlabs. 2024. sp1. https://github.com/succinctlabs/sp1/tree/v1.0.1. Ac- cessed: 2025-11-10
2024
-
[30]
2024.SP1 Technical Whitepaper
succinctlabs. 2024.SP1 Technical Whitepaper. Technical Whitepa- per. succinctlabs. https://drive.google.com/file/d/1aTCELr2b2Kc1NS- wZ0YYLKdw1Y2HcLTr/view Accessed: 2025-11-10
2024
-
[31]
syntacore. 2017. scr1. https://github.com/syntacore/scr1. Accessed: 2025-11-10
2017
-
[32]
Justin Thaler. 2025. The path to secure and efficient zkVMs: How to track progress. https://a16zcrypto.com/posts/article/secure-efficient-zkvms-progress/
2025
-
[33]
Cheng Wang and Mingyu Gao. 2023. Sam: A scalable accelerator for number theoretic transform using multi-dimensional decomposition. In2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD). IEEE, 1–9
2023
-
[34]
Cheng Wang and Mingyu Gao. 2025. UniZK: Accelerating Zero-Knowledge Proof with Unified Hardware and Flexible Kernel Mapping. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1. 1101–1117
2025
-
[35]
Xi Wang, John D Leidel, Brody Williams, Alan Ehret, Miguel Mark, Michel A Kinsy, and Yong Chen. 2021. xbgas: A global address space extension on risc-v for high performance computing. In2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 454–463
2021
-
[36]
Xi Wang, Brody Williams, John D Leidel, Alan Ehret, Michel Kinsy, and Yong Chen. 2020. Remote atomic extension (rae) for scalable high performance com- puting. In2020 57th ACM/IEEE Design Automation Conference (DAC). IEEE, 1–6
2020
-
[37]
Tiancheng Xie, Jiaheng Zhang, Zerui Cheng, Fan Zhang, Yupeng Zhang, Yongzheng Jia, Dan Boneh, and Dawn Song. 2022. zkbridge: Trustless cross-chain bridges made practical. InProceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 3003–3017
2022
-
[38]
Yongkui Yang, Zhenyan Lu, Jingwei Zeng, Xingguo Liu, Xuehai Qian, and Zhibin Yu. 2024. Falic: An FPGA-Based Multi-Scalar Multiplication Accelerator for Zero-Knowledge Proof.IEEE Trans. Comput.(2024)
2024
-
[39]
Zhengbang Yang, Lutan Zhao, Peinan Li, Han Liu, Kai Li, Boyan Zhao, Dan Meng, and Rui Hou. 2025. LegoZK: A Dynamically Reconfigurable Accelerator for Zero- Knowledge Proof. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 113–126
2025
-
[40]
Jiaheng Zhang, Zhiyong Fang, Yupeng Zhang, and Dawn Song. 2020. Zero knowledge proofs for decision tree predictions and accuracy. InProceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security. 2039–2053
2020
-
[41]
Ye Zhang, Shuo Wang, Xian Zhang, Jiangbin Dong, Xingzhong Mao, Fan Long, Cong Wang, Dong Zhou, Mingyu Gao, and Guangyu Sun. 2021. Pipezk: Accel- erating zero-knowledge proof with a pipelined architecture. In2021 ACM/IEEE 48th Annual International Symposium on Computer Architec...
2021
-
[42]
Hao Zhou, Changxu Liu, Lan Yang, Li Shang, and Fan Yang. 2024. A Fully Pipelined Reconfigurable Montgomery Modular Multiplier Supporting Variable Bit-Widths.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems43, 12 (2024), 4653–4665
2024
-
[43]
Hao Zhou, Changxu Liu, Lan Yang, Li Shang, and Fan Yang. 2024. ReZK: A Highly Reconfigurable Accelerator for Zero-Knowledge Proof.IEEE Transactions on Circuits and Systems I: Regular Papers(2024)
2024
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.