REVIEW 4 major objections 5 minor 97 references
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read LLMulator claims a pre-trained LLM can predict dataflow-accelerator performance at 12.2% average error across unseen applications, inputs, and hardware.
desk verdict Static modeling and data synthesis are solid contributions, but the dynamic-calibration headline numbers are inflated by DPO updates on the test workloads' own ground-truth profiles, so the generalization claim needs a clean held-out evaluation before it can be believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Numeric modeling is the mechanism that carries the argument: instead of regressing a single number, the model predicts each digit of the performance value as a separate classification decision, ordered from most to least significant, with beam search to correct high-order errors and logits that read as confidence. Supporting it are the progressive tokenizer, which isolates numeric tokens so their length scales with digit count, and the dynamic calibration loop, which uses direct-preference optimization on preference pairs (predicted vs. profiled performance) to adapt to input-dependent control flow. The dataset synthesizer supplies the breadth of (program, hardware, performance) triples that
What would settle it
Evaluate LLMulator on the same workloads in direct mode, where the program and parameters are the only inputs and no synthesis-extracted module or multiplexer counts are supplied; if average MAPE rises to the baselines' level, the end-to-end advantage for users who skip synthesis would disappear.
Extended reading notes
Core claim
LLMulator treats dataflow performance prediction as a language-model task: input is the dataflow graph program, operator implementations, hardware parameters, and runtime input scalars; output is a vector of power, area, flip-flops, and cycles. Its core discovery is that three design choices together make this generalization work. First, progressive numeric modeling: numbers inside the program are isolated and tokenized digit-wise, and the predicted performance value is decoded digit-by-digit from most-significant to least-significant digit using categorical classification with beam search, yielding confidence estimates at each position and reducing extreme-value errors. Second, dynamic pred
Load-bearing premise
The accuracy claims assume the intermediate hardware features the model reasons over are available at prediction time; if a user needs a fast prediction without running synthesis, the model must generate those features itself, and the paper does not test whether that preserves accuracy.
Editorial extensions
If this is right
- Design-space exploration for dataflow accelerators could use LLMulator as a fast filter before expensive synthesis, since a prediction costs about a second versus minutes to hours for HLS and physical synthesis.
- Cost models that output confidence at each digit position would let designers know when an early estimate is unreliable and needs a synthesis check.
- The dynamic calibration recipe means a deployed cost model can keep improving as new input profiles arrive, without retraining from scratch.
- The progressive data synthesizer's approach could be reused as a data-generation recipe for other learned hardware models, since adding those synthesized examples also lowers error for the baselines tested.
- If the reasoning-data variant is used, the model exposes intermediate RTL-level features as a chain of thought, making predictions more interpretable than a black-box regression.
Reading between the lines
- The paper's accuracy numbers are measured with intermediate RTL-level features supplied by a synthesis tool. A user who wants the speed advantage without running synthesis would need the model to generate those features itself; that path is not evaluated, so the end-to-end speed claim is an open question.
- The digit-wise categorical decoding is a general trick: any learned predictor with a bounded target metric could adopt it, though the paper only demonstrates it for power, area, flip-flops, and cycles.
- Attention masking in dynamic prediction hints at a compositional extension: predict per-operator costs independently and cache them across design-space iterations, making latency scale with the number of changed operators rather than the whole graph.
- A testable extension is to run LLMulator on a held-out accelerator whose control flow depends on input in an unseen way, such as variable-length sequence models, and measure whether the same few DPO iterations still converge.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. LLMulator proposes an LLM-based cost model for dataflow accelerators that (1) tokenizes numeric values in programs and predicts performance as digit-wise classification with beam search, (2) calibrates cycle predictions online via DPO using profiler feedback, and (3) augments training data with progressively synthesized programs and RTL-level "thinking" features. The paper reports a 12.2% average MAPE across static and dynamic metrics, a 9.7% improvement from dynamic calibration, and convergence of cycle-prediction error to 11.2% after several iterations.
Significance. The problem is important and the three proposed components are individually plausible. Treating numeric outputs as categorical digit sequences is a sensible way to avoid regression saturation, and using DPO with execution feedback for input-adaptive calibration is a novel direction. The progressive data synthesizer also addresses a real limitation of existing datasets. However, the current evaluation does not support the central generalization claims: the dynamic-calibration numbers are computed on the same workloads whose ground truth drives the DPO updates, and the static predictions appear to require RTL features extracted by the same SiliconCompiler flow the model is meant to replace. If these issues are fixed with a held-out calibration protocol and a clearly specified inference-time feature pipeline, the contribution could be significant; as it stands, the headline accuracy is not an out-of-sample prediction result.
major comments (4)
- [§5.1, Fig. 4, Eq. (2); §7.2, Table 3; §7.4, Table 11] The dynamic-calibration evaluation leaks test labels. The DPO procedure in Eq. (2) uses the profiler's ground-truth y_w for a state {x, data} to construct preference pairs, and Section 7.4 states that LLMulator is "dynamically calibrated using input profiles collected during TPU runs" for the same workloads. Section 7.2 then reports the Dynamic-Cycles MAPE (16.4% after 5 iterations) on those same workloads. This is in-sample fitting, not prediction, and it is not a like-for-like comparison with GNNHLS, TLP, or Tenset-MLP, which are not updated with the corresponding ground-truth profiles. The abstract's "9.7% improvement" and "11.2% convergence" are therefore not supported as generalization results. The authors need a held-out calibration set: DPO updates on a subset of inputs, evaluation on disjoint inputs, and reporting of both before- and after-calibration MAPE on that held-out set.
- [§6.2, Figs. 8–10; §7.2, Table 3; §7.1, Table 4] The inference-time availability of the "thinking" features is unspecified. The reasoning format inserts SiliconCompiler-extracted RTL-level features (module counts, MUX counts, estimated areas) into the model input, and Table 3's static power/area/FF results are obtained with this format. If these features require running HLS/synthesis at inference time, then the static results are not end-to-end cost-model predictions, and the runtime advantage over synthesis claimed in Table 4 is not demonstrated because the feature-extraction cost is omitted. If the model is expected to predict or bypass these features, that variant must be evaluated explicitly. Please state the inference protocol and report end-to-end latency including any required EDA steps.
- [§7.3, Table 6] The confidence-calibration claim is based on a Pearson correlation of -0.44 between the final-layer logit and MSE on 12 randomly sampled workloads. This sample size is too small to establish a calibration property, and a correlation with MSE is not a calibration curve. A proper evaluation would report expected calibration error (ECE) or a reliability diagram across confidence bins, with confidence intervals, on a larger held-out sample.
- [§7.2, Table 3] The reported aggregate 12.2% MAPE mixes static metrics (trained and evaluated on the synthesized dataset) and dynamic-cycle metrics (obtained after DPO calibration on the test workloads). The headline comparison against TLP and GNNHLS therefore conflates two different evaluation protocols. The paper should separate the static and dynamic claims, and for the dynamic claim it should compare every method under the same calibration/update protocol, or explicitly state that the baselines are not dynamically calibrated.
minor comments (5)
- [§2] The second bullet under "Numerical range compression distortion" contains an explicit placeholder: "(specific experimental numbers or benchmark results to be added)". A submitted manuscript should not contain such placeholders, and the surrounding claims about threefold edge errors and >40% relative error are not yet supported by the text.
- [Table 3] The table formatting is corrupted: rows for Polybench kernels appear merged with values, e.g., "correlation1.0% 1.6% ..." and "covariance28.7% ...". This makes the per-benchmark results difficult or impossible to verify.
- [Table 7] The column label "No-A" is never defined. Also, the cycles column shows several workloads where the full pipeline is worse than the No-A configuration (e.g., Tab. 2-6: 51.4% → 67.3%; Tab. 2-8: 12.9% → 31.6%). This contradicts the blanket statement that the dataset synthesizer delivers consistent improvement and needs explanation.
- [Figure 11] The Timeloop comparison figure is largely unreadable: the text is tiny, axis labels are missing, and the legend is ambiguous. Please redraw it with legible fonts and explicit per-workload labels.
- [§7.1] The evaluation reports pass@5 sampling but gives no standard deviation or confidence intervals across sampling runs. Given the small benchmark set and the stochasticity of LLM decoding, this makes it hard to assess whether the reported differences are meaningful.
Circularity Check
Headline accuracy is inflated by DPO calibration on the evaluated workloads' ground-truth profiles and by RTL feature leakage into static area predictions.
-
fitted input called prediction
[Section 5.1 (Eq. 2) and Section 7.2 (Table 3)]
""(3) Profiler feedback: The environment (e.g., Siliconcompiler) returns the “ground-truth” performance y_w for the same {x, data} ... Construct a preference triplet ({x, data}, y_w, y_l) ... R(θ)=E[log σ(β(log π_θ(y_w|{x,data})/π_ref(y_w|{x,data}) − log π_θ(y_l|{x,data})/π_ref(y_l|{x,data})))] ... our dynamic calibration framework reduces MAPE from initial 28.9% to 16.4% after 5 DPO iterations.""
The DPO update in Eq. 2 directly optimizes the model to raise the log-probability of the ground-truth label y_w for the same input {x, data} that is later scored. Section 7.2 reports dynamic-cycle MAPE after 5 such DPO iterations on the Table 2 workloads, and Section 7.4 says the model is "dynamically calibrated using input profiles collected during TPU runs" for the same real-world workloads. No held-out calibration set is described. Thus the reported cycle-error improvements (28.9%→16.4%, and the 11.2% convergence claim) are fits to the evaluation labels, not independent predictions, making the headline generalization numbers test-label-influenced by construction.
-
self definitional
[Section 6.2, Reasoning Data Formatting and Figures 8-9]
""Intermediate feature extraction. We utilize SiliconCompiler [63] to extract critical RTL-level features (e.g., module counts, multiplexer numbers)... [Dataflow Graph][Operator Program] <think>[RTL Level Semantic Analysis Result]</think><Power>[Power Profiled Value]</Power><Area>[Area Profiled Value]</Area> ... Number of modules instantiated: 81 ... Estimated resources area: 1399 ... Estimated area of MUX21: 584.5""
The target output is <Area>[Area Profiled Value]</Area>, but the input <think> block already contains SiliconCompiler-extracted area estimates such as "Estimated resources area: 1399" and "Estimated area of MUX21: 584.5." The model can satisfy the area-prediction task by copying an input feature that is itself an area estimate produced by the same compilation flow. The reported static area MAPE reduction (e.g., from 27.1% to 14.2%) is therefore partially a restatement of compiler-provided estimates rather than a first-principles prediction. The paper also never verifies that these RTL features are available at inference or that the model can generate them accurately, so the end-to-end prediction claim is unsupported.
full rationale
The paper's central generalization claims are partially circular. First, the dynamic-cycle results are test-label-influenced: Eq. 2 performs DPO against the ground-truth y_w for the same {x, data} that is later evaluated, and Sections 7.2/7.4 report cycle MAPE after 5 such DPO iterations on those same workloads without describing a held-out calibration set. This is the strongest circular step, directly undermining the 11.2% convergence and the headline 12.2% MAPE insofar as it aggregates dynamic cycles. Second, the static area/FF predictions are compromised by feeding SiliconCompiler's estimated-area and MUX-count features into the input while asking for area as output; the model can copy those estimates, so the static metric improvements are not clean first-principles predictions. The remaining components—numeric tokenization, progressive data synthesis, and attention-mask acceleration—have independent content and are not themselves circular. I also note the explicit missing-support flag in Section 2: "Experimental results (specific experimental numbers or benchmark results to be added)" is a placeholder for the claimed edge-error evidence; this is not circularity but weakens the motivation. Overall, because the headline quantitative claims reduce in part to fitting test labels and to echoing compiler estimates, a score of 7 is appropriate.
Assumptions & free parameters
free parameters (4)
- Digit base D =
10 (decimal)
- Dataset composition ratio =
30% AST, 50% dataflow, 20% LLM-generated
- Memory delay values in synthesizer =
2, 5, 10 cycles
- DPO temperature beta
assumptions (4)
- domain assumption Pre-trained LLM weights encode program-semantic knowledge that transfers to numerical hardware-cost prediction
- domain assumption The open-source EDA profiling flow (Bambu, OpenROAD, Verilator, SkyWater130nm) produces ground-truth labels that are accurate and representative
- ad hoc to paper RTL-level 'thinking' features (module counts, MUX counts) are available at inference time without running the synthesis that the model is meant to replace
- domain assumption The three-stage synthetic data generator covers the distribution of real dataflow applications and hardware configurations
Cite this review
Pith. "Pith review of LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow." pith.science (2026). https://pith.science/paper/SJIPSXOR
@misc{pith2026250817826,
author = {Pith},
title = {Pith review of: LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow},
year = {2026},
howpublished = {\url{https://pith.science/paper/SJIPSXOR}},
note = {Machine review of arXiv:2508.17826}
}
read the original abstract
Accurate and fast performance prediction for dataflow-based accelerators is vital for efficient hardware design and design space exploration, yet existing methods struggle to generalize across architectures, applications, and input-dependent control flows. We present LLMulator, a progressive numeric modeling framework leveraging the program semantic knowledge of pre-trained large language models (LLMs) for robust, hardware- and application-aware prediction. Our numeric model treats performance values as categorical token sequences, enabling range-agnostic estimates and confidence-aware predictions for unseen applications. To handle input-dependent control flows, we introduce a reinforcement learning-based dynamic calibration method, reducing cycle prediction error by 9.7% over static models and converging to 11.2% error after a few iterations. For cross-hardware generalization, we develop a progressive data augmentation strategy that generates diverse datasets covering multi-level dataflow structures, memory parameters, and loop mapping primitives, significantly boosting prediction accuracy across architectures and configurations.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B
Tutu Ajayi, Vidya A. Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B. Kahng, Minsoo Kim, Jeongsup Lee, Uday Mallappa, Marina Neseem, Geraldo Pradipta, Sherief Reda, Mehdi Saligane, Sachin S. Sapatnekar, Carl Sechen, Mohamed Shalan, William Swartz, Lutong Wang, Zhehong Wang, Mingyu Woo, and Bangqi Xu. 2019. Toward an open-source digita...
2019
-
[2]
Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham LEGHETTAS, Kamel Abdous, Taha Arbaoui, Karima BENATCHBA, and Saman amarasinghe. 2021. A deep learning based cost model for automatic code optimization. Proceedings of Machine Learning and Systems 3 (2021), 181–193
2021
-
[3]
Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Ab- durrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil, and Saman Amarasinghe. 2019. Tiramisu: A polyhedral compiler for expressing fast and portable code. In 2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 193–205
2019
-
[4]
Gergö Barany. 2017. Liveness-driven random program generation. InInternational Symposium on Logic-Based Program Synthesis and Transformation . Springer, 112– 127
2017
-
[5]
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long- document transformer. arXiv preprint arXiv:2004.05150 (2020)
arXiv 2020
-
[6]
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. (2009), 41–48
2009
-
[7]
Jun Bi, Qi Guo, Xiaqing Li, Yongwei Zhao, Yuanbo Wen, Yuxuan Guo, Enshuai Zhou, Xing Hu, Zidong Du, Ling Li, Huaping Chen, and Tianshi Chen. 2023. Heron: Automatically constrained high-performance library generation for deep learning accelerators. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages an...
2023
-
[8]
Alexander Brauckmann, Elizabeth Polgreen, Tobias Grosser, and Michael F. P. O’Boyle. 2023. mlirSynth: Automatic, Retargetable Program Raising in Multi- Level IR Using Program Synthesis. In 2023 32nd International Conference on Parallel Architectures and Compilation Techniques (PACT). 39–50. https://doi.org/ 10.1109/PACT58117.2023.00012
Show all 97 references
-
[9]
Kaiyan Chang, Kun Wang, Nan Yang, Ying Wang, Dantong Jin, Wenlong Zhu, Zhirong Chen, Cangyuan Li, Hao Yan, Yunhao Zhou, Zhuoliang Zhao, Yuan Cheng, Yudong Pan, Yiqi Liu, Mengdi Wang, Shengwen Liang, Yinhe Han, Huawei Li, and Xiaowei Li. 2024. Data is all you need: Finetuning L...
2024
-
[10]
Bodhisatwa Chatterjee, Neeraj Jadhav, Sharjeel Khan, and Santosh Pande. 2024. Phaedrus: Exploring Dynamic Application Behavior with Lightweight Generative Models and Large-Language Models. arXiv preprint arXiv:2412.06994 (2024)
2024
-
[11]
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th US...
2018
-
[12]
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to optimize tensor programs. Advances in Neural Information Processing Systems 31 (2018)
2018
-
[13]
Yu-Hsin Chen, Joel Emer, and Vivienne Sze. 2016. Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks. ACM SIGARCH computer architecture news 44, 3 (2016), 367–379
2016
-
[14]
François Chollet. 2017. Xception: Deep learning with depthwise separable con- volutions. In Proceedings of the IEEE conference on computer vision and pattern recognition. 1251–1258
2017
-
[15]
Chris Cummins, Zacharias V Fisches, Tal Ben-Nun, Torsten Hoefler, Michael FP O’Boyle, and Hugh Leather. 2021. Programl: A graph-based program representa- tion for data flow analysis and compiler optimizations. InInternational Conference on Machine Learning. PMLR, 2244–2253
2021
-
[16]
Gautier Dagan, Gabriel Synnaeve, and Baptiste Roziere. 2024. Getting the most out of your tokenizer for pre-training and domain adaptation. In International Conference on Machine Learning . PMLR, 9784–9805
2024
-
[17]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang 13 MICRO ’25, October 18–22, 2025, Seoul, Republic of Korea Kaiyan Chang, Wenlong Zhu, Shengwen Liang, Huawei Li, and Ying Wang Zhang, X...
2025 arXiv
-
[18]
Stanislas Dehaene. 2011. The number sense: How the mind creates mathematics . OUP USA
2011
-
[19]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human...
2019
-
[20]
Zidong Du, Robert Fasthuber, Tianshi Chen, Paolo Ienne, Ling Li, Tao Luo, Xiaobing Feng, Yunji Chen, and Olivier Temam. 2015. ShiDianNao: Shifting vision processing closer to the sensor. In Proceedings of the 42nd annual international symposium on computer architecture . 92–104
2015
-
[21]
Lieven Eeckhout, Robert H Bell Jr, Bastiaan Stougie, Koen De Bosschere, and Lizy K John. 2004. Control flow modeling in statistical simulation for accurate and efficient processor design studies. ACM SIGARCH Computer Architecture News 32, 2 (2004), 350
2004
-
[22]
Fabrizio Ferrandi, Vito Giovanni Castellana, Serena Curzel, Pietro Fezzardi, Michele Fiorito, Marco Lattuada, Marco Minutoli, Christian Pilato, and Antonino Tumeo. 2021. Invited: Bambu: an Open-Source Research Framework for the High-Level Synthesis of Complex Applications. In ...
2021
-
[23]
Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with PyTorch Geometric. arXiv preprint arXiv:1903.02428 (2019)
2019 arXiv
-
[24]
Ross Girshick. 2015. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision. 1440–1448
2015
-
[25]
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in neural information processing systems 27 (2014)
2014
-
[26]
Google and SkyWater Technology Foundry. 2020. SkyWater Open Source PDK: Open Source Process Design Kit for Usage with SkyWater Technology Foundry’s 130nm Node. GitHub Repository. https://github.com/google/skywater-pdk Repository includes documentation, libraries, and tooling files
2020
-
[27]
Scott Grauer-Gray, Lifan Xu, Robert Searles, Sudhee Ayalasomayajula, and John Cavazos. 2012. Auto-tuning a high-level language targeted to GPU codes. In 2012 Innovative Parallel Computing (InPar). 1–10. https://doi.org/10.1109/InPar.2012. 6339595
2012 doi
-
[28]
Yuntao Gui, Yidi Wu, Han Yang, Tatiana Jin, Boyang Li, Qihui Zhou, James Cheng, and Fan Yu. 2022. HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and Optimization. In SC22: International Conference for High Performance Computing, Networking, Storage a...
2022 arXiv
-
[29]
Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. Advances in neural information processing systems 30 (2017)
2017
-
[30]
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision . 2961–2969
2017
-
[31]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence 37, 9 (2015), 1904–1916
2015
-
[32]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[33]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations . https: //openreview.net/forum?id=nZeVKeeFYf9
2022
-
[34]
Wenhao Hu, Jinhao Duan, Chunchen Wei, Li Zhang, Yue Zhang, and Kaidi Xu
-
[35]
Qijing Huang, Minwoo Kang, Grace Dinh, Thomas Norell, Aravind Kalaiah, James Demmel, John Wawrzynek, and Yakun Sophia Shao. 2021. Cosa: Scheduling by constrained optimization for spatial accelerators. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architectur...
2021
-
[36]
Forrest Iandola, Matt Moskewicz, Sergey Karayev, Ross Girshick, Trevor Darrell, and Kurt Keutzer. 2014. Densenet: Implementing efficient convnet descriptor pyramids. arXiv preprint arXiv:1404.1869 (2014)
2014 arXiv
-
[37]
Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning. pmlr, 448–456
2015
-
[38]
Lana Josipović, Shabnam Sheikhha, Andrea Guerrieri, Paolo Ienne, and Jordi Cortadella. 2021. Buffer placement and sizing for high-performance dataflow circuits. ACM Transactions on Reconfigurable Technology and Systems (TRETS) 15, 1 (2021), 1–32
2021
-
[39]
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazi...
2017
-
[40]
Wookeun Jung, Thanh Tuan Dao, and Jaejin Lee. 2021. DeepCuts: a deep learning optimization framework for versatile GPU workloads. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation. 190–205
2021
-
[41]
Sheng-Chun Kao, Suvinay Subramanian, Gaurav Agrawal, Amir Yazdanbakhsh, and Tushar Krishna. 2023. Flat: An optimized dataflow for mitigating attention bottlenecks. In Proceedings of the 28th ACM International Conference on Archi- tectural Support for Programming Languages and ...
2023
-
[42]
Jeyhun Karimov, Tilmann Rabl, and Volker Markl. 2019. Polybench: The first benchmark for polystores. In Performance Evaluation and Benchmarking for the Era of Artificial Intelligence: 10th TPC Technology Conference, TPCTC 2018, Rio de Janeiro, Brazil, August 27–31, 2018, Revis...
2019
-
[43]
Florent Kirchner, Nikolai Kosmatov, Virgile Prevosto, Julien Signoles, and Boris Yakobowski. 2015. Frama-C: A software analysis perspective. Formal aspects of computing 27, 3 (2015), 573–609
2015
-
[44]
Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017. Overcoming catastrophic forget...
2017
-
[45]
Klaus Krogmann, Michael Kuperberg, and Ralf Reussner. 2010. Using Genetic Search for Reverse Engineering of Parametric Behavior Models for Performance Prediction. IEEE Transactions on Software Engineering 36, 6 (2010), 865–877. https://doi.org/10.1109/TSE.2010.69
2010 doi
-
[46]
Hyoukjun Kwon, Prasanth Chatarasi, Michael Pellauer, Angshuman Parashar, Vivek Sarkar, and Tushar Krishna. 2019. Understanding Reuse, Performance, and Hardware Cost of DNN Dataflow: A Data-Centric Approach. InProceedings of the 52nd Annual IEEE/ACM International Symposium on M...
2019
-
[47]
Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. Maestro: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings. IEEE Micro 40, 3 (2020), 20–29
2020
-
[48]
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942 (2019)
2019 arXiv
-
[49]
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient- based learning applied to document recognition. Proc. IEEE 86, 11 (1998), 2278– 2324
1998
-
[50]
Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunning- ham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, and Wenzhe Shi. 2017. Photo-realistic single image super-resolution us- ing a generative adversarial network. In Procee...
2017
-
[51]
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. 2023. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arX...
2023 arXiv
-
[52]
Fuping Li, Ying Wang, Cheng Liu, Huawei Li, and Xiaowei Li. 2022. Noception: a fast ppa prediction framework for network-on-chips using graph neural network. In 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 1035–1040
2022
-
[53]
Ming Li and Paul Vitányi. 2008. An introduction to Kolmogorov complexity and its applications. Vol. 3. Springer
2008
-
[54]
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. 2024. Bevformer: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[55]
Zhaoying Li, Dan Wu, Dhananjaya Wijerathne, and Tulika Mitra. 2022. Lisa: Graph neural network based portable mapping on spatial accelerators. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 444–459
2022
-
[56]
Zhe Lin, Jieru Zhao, Sharad Sinha, and Wei Zhang. 2020. HL-Pow: A learning- based power modeling framework for high-level synthesis. In 2020 25th Asia and South Pacific Design Automation Conference (ASP-DAC). IEEE, 574–580
2020
-
[57]
Shang Liu, Wenji Fang, Yao Lu, Jing Wang, Qijun Zhang, Hongce Zhang, and Zhiyao Xie. 2025. RTLCoder: Fully Open-Source and Efficient LLM-Assisted RTL Code Generation Technique. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 44, 4 (2025), 1448–146...
2025
-
[58]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)
2019 arXiv
-
[59]
Liqiang Lu, Naiqing Guan, Yuyue Wang, Liancheng Jia, Zizhang Luo, Jieming Yin, Jason Cong, and Yun Liang. 2021. TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric Notation. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (IS...
2021
-
[60]
Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Meiqi Guo, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, and Abhinav Rastogi. 2024. Improve Mathematical Reasoning in Language Models by Automated Process Supervision. (2024). arXiv:2406.06592 [cs.CL] htt...
2024 arXiv
-
[61]
Hosein Mohammadi Makrani, Hossein Sayadi, Tinoosh Mohsenin, Setareh Rafati- rad, Avesta Sasan, and Houman Homayoun. 2019. Xppe: cross-platform perfor- mance estimation of hardware accelerators using machine learning. InProceedings of the 24th Asia and South Pacific Design Auto...
2019
-
[62]
Gordon Euhyun Moon, Hyoukjun Kwon, Geonhwa Jeong, Prasanth Chatarasi, Sivasankaran Rajamanickam, and Tushar Krishna. 2021. Evaluating spatial accel- erator architectures with tiled matrix-matrix multiplication. IEEE Transactions on Parallel and Distributed Systems 33, 4 (2021)...
2021
-
[63]
Andreas Olofsson, William Ransohoff, and Noah Moroze. 2022. A Distributed Approach to Silicon Compilation: Invited. In Proceedings of the 59th ACM/IEEE Design Automation Conference (San Francisco, California). 1343–1346
2022
-
[64]
Ying, Anurag Mukkara, Rangharajan Venkatesan, Brucek Khailany, Stephen W
Angshuman Parashar, Priyanka Raina, Yakun Sophia Shao, Yu-Hsin Chen, Victor A. Ying, Anurag Mukkara, Rangharajan Venkatesan, Brucek Khailany, Stephen W. Keckler, and Joel Emer. 2019. Timeloop: A Systematic Approach to DNN Accelerator Evaluation. In Proceedings of the 2019 IEEE...
2019
-
[65]
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Im- proving language understanding by generative pre-training. OpenAI (2018)
2018
-
[66]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2023), 53728–53741
2023
-
[67]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67
2020
-
[68]
James Requeima, John Bronskill, Dami Choi, Richard Turner, and David K Duve- naud. 2024. Llm processes: Numerical predictive distributions conditioned on natural language. Advances in Neural Information Processing Systems 37 (2024), 109609–109671
2024
-
[69]
Abellán, Ajay Joshi, John Kim, and David Kaeli
Kaustubh Shivdikar, Nicolas Bohm Agostini, Malith Jayaweera, Gilbert Jonatan, José L. Abellán, Ajay Joshi, John Kim, and David Kaeli. 2024. NeuraChip: Accel- erating GNN Computations with a Hash-based Decoupled Spatial Accelerator. In 2024 ACM/IEEE 51st Annual International Sy...
2024
-
[70]
Xinyu Sun, Yu Zhang, Shuo Liu, and Yi Zhai. 2024. Crop: An Analytical Cost Model for Cross-Platform Performance Prediction of Tensor Programs. In Pro- ceedings of the 61st ACM/IEEE Design Automation Conference (San Francisco, CA, USA) (DAC ’24). Association for Computing Machi...
2024
-
[71]
Qwen Team. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671 (2024)
2024 arXiv
-
[72]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucu- rull, David Esiobu, Jude Fernandes, Jeremy...
2023 arXiv
-
[73]
Robert Vacareanu, Vlad Andrei Negru, Vasile Suciu, and Mihai Surdeanu. 2024. From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples. In First Conference on Language Modeling. https://openreview.net/forum?id=LzpaUxcNFK
2024
-
[74]
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:180...
2018 arXiv
-
[75]
verilator and Contributors. 2024. Verilator: Open-source SystemVerilog simulator and lint system. https://github.com/verilator/verilator GitHub repository
2024
-
[76]
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tris- tan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gal- louédec. 2020. TRL: Transformer Reinforcement Learning. https://github.com/ huggingface/trl
2020
-
[77]
Fuyu Wang, Minghua Shen, Yufei Ding, and Nong Xiao. 2024. Soter: Analytical Tensor-Architecture Modeling and Automatic Tensor Program Tuning for Spatial Accelerators. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 991–1004
2024
-
[78]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837
2022
-
[79]
Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV) . 3–19
2018
-
[80]
Nan Wu, Hang Yang, Yuan Xie, Pan Li, and Cong Hao. 2022. High-level synthesis performance prediction using gnns: Benchmarking, modeling, and advancing. In Proceedings of the 59th ACM/IEEE Design Automation Conference . 49–54
2022
-
[81]
Zhenglong Wu, Qi Qi, Zirui Zhuang, Haifeng Sun, and Jingyu Wang. 2024. Pre- tokenization of numbers for large language models. In The Second Tiny Papers Track at ICLR 2024
2024
-
[82]
Chunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang, Yuan Wen, Teng Yu, Xiaobing Feng, and Huimin Cui. 2024. Optimizing deep learning inference via global analysis and tensor expressions. In Proceedings of the 29th ACM Inter- national Conference on Architectural Support for Pro...
2024
-
[83]
Zhiyao Xie, Xiaoqing Xu, Matt Walker, Joshua Knebel, Kumaraguru Palaniswamy, Nicolas Hebert, Jiang Hu, Huanrui Yang, Yiran Chen, and Shidhartha Das. 2021. APOLLO: An Automated Power Modeling Framework for Runtime Power In- trospection in High-Volume Commercial Microprocessors....
2021
-
[84]
Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, and Yibo Zhu. 2022. Bolt: Bridging the gap between auto-tuners and hardware-native performance. Proceedings of Machine Learning and Systems 4 (2022), 204–216
2022
-
[85]
Han Xu, Jiayi Ma, Junjun Jiang, Xiaojie Guo, and Haibin Ling. 2020. U2Fusion: A unified unsupervised image fusion network. IEEE transactions on pattern analysis and machine intelligence 44, 1 (2020), 502–518
2020
-
[86]
Xuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter, Jing Pu, Ankita Nayak, Steven Bell, Kaidi Cao, Heonjae Ha, Priyanka Raina, Christos Kozyrakis, and Mark Horowitz
-
[87]
Fisher Yu and Vladlen Koltun. 2015. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122 (2015)
2015 arXiv
-
[88]
Yuan Yu, Martín Abadi, Paul Barham, Eugene Brevdo, Mike Burrows, Andy Davis, Jeff Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Murray, and Xiaoqiang Zheng. 2018. Dynamic control flow in large-scale machine learning. In P...
2018
-
[89]
Yi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu, Jie Peng, Jianmin Ji, and Yanyong Zhang. 2023. TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating S...
2023
-
[90]
Jie Zhao, Bojie Li, Wang Nie, Zhen Geng, Renwei Zhang, Xiong Gao, Bin Cheng, Chen Wu, Yun Cheng, Zheng Li, Peng Di, Kun Zhang, and Xuefeng Jin. 2021. AKG: automatic kernel generation for neural processing units using polyhedral trans- formations. In Proceedings of the 42nd ACM...
2021
-
[91]
Gonzalez, and Ion Stoica
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica. 2020. Ansor: Generating High-Performance Tensor Programs for Deep Learning. In 14th USENIX Symposium on Operating S...
2020
-
[92]
Gonzalez, Ion Stoica, and Ameer Haj Ali
Lianmin Zheng, Ruochen Liu, Junru Shao, Tianqi Chen, Joseph E. Gonzalez, Ion Stoica, and Ameer Haj Ali. 2021. Tenset: A Large-Scale Program Performance Dataset for Learned Tensor Compilers. In Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS)
2021
-
[93]
Size Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia, Guangyu Sun, Runsheng Wang, and Yun Liang. 2023. Tileflow: A Framework for Modeling Fusion Dataflow via Tree-Based Analysis. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . 1271–1288
2023
-
[94]
Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng. 2020. Flex- tensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous Systems. In Proceedings of the 25th In- ternational Conference on Architectural Support for ...
2020
-
[95]
Yuan Zhou, Haoxing Ren, Yanqing Zhang, Ben Keller, Brucek Khailany, and Zhiru Zhang. 2019. PRIMAL: Power inference using machine learning. In Proceedings of the 56th Annual Design Automation Conference 2019 . 1–6. 16
2019
-
[2020]
In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems
Interstellar: Using halide’s scheduling language to analyze dnn accelerators. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems . 369–383
-
[2025]
arXiv:2503.10452 [cs.CL] https: //arxiv.org/abs/2503.10452
DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation. arXiv:2503.10452 [cs.CL] https: //arxiv.org/abs/2503.10452
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.