REVIEW 3 major objections 5 minor 2 cited by
Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A transferred surrogate predicts RTL-level energy-delay with 99% rank accuracy using 61% fewer simulations, and keeping RTL simulation in the optimization loop yields designs 2.7x better in energy-delay product.
desk verdict Useful integration of transfer learning and online RTL-in-the-loop BO for DLA design, but the headline accuracy and sample-efficiency claims are unverifiable as written due to a 212-vs-1600 dataset inconsistency and a missing delay-only ablation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Starlight is built from deep kernel learning (DKL): a variational autoencoder's encoder network (trained with a predictor head that imposes a smooth EDP gradient on the latent space) is transferred by hard weight sharing from Starlight-Low, then attached to a Gaussian process with a Matérn kernel that supplies the uncertainty estimate needed for Bayesian optimization. Polaris wraps this surrogate in an outer hardware loop (enumerating discrete array/scratchpad/accumulator choices) and an inner per-layer software loop (sampling 10,000 Sobol candidates per iteration), selects candidates with an Upper Confidence Bound acquisition function, and evaluates them on the RTL simulator, feeding each result back into Starlight.
What would settle it
Repeat the Starlight training and the Polaris search using delay-only labels, comparing FireSim delay against Timeloop delay with no shared energy term; if the rank correlation or the 61% sample savings drop materially, the central transfer-learning claim is inflated by the shared energy model, and if they hold, the claim is robust.
Extended reading notes
Core claim
The paper's central claim is that the encoder of a variational autoencoder trained on cheap analytical-model evaluations of a DLA (Starlight-Low) can be transplanted into a deep-kernel-learning model and fine-tuned on a small set of RTL-simulation evaluations to produce Starlight, a surrogate that predicts RTL-measured energy-delay product with Spearman rank correlation 0.99 while using 61% fewer high-fidelity samples than the DOSA baseline. On top of that, the paper claims Polaris—a Bayesian optimizer that keeps an RTL simulator inside the optimization loop—consistently finds designs with lower EDP than offline optimizers, beating DOSA's best designs by an average of 2.7x in under 3.3 hours. The authors argue this is the first demonstration that RTL simulation in the loop, rather than only as a final check, materially improves the quality of the produced hardware/software co-designs.
Load-bearing premise
The load-bearing assumption is that the knowledge transferred from the analytical model to the RTL simulator comes from genuinely shared performance structure, but because the hybrid energy-delay label uses the same analytical-model energy term in both source and target, the transfer gain could be partly an artifact of that shared measurement.
Editorial extensions
If this is right
- Training a high-fidelity DLA performance model can be done with 61% fewer RTL simulations, reducing the one-time data-collection bottleneck.
- Design space exploration can use RTL simulation as the evaluator without giving up search breadth, because Starlight evaluates ~6,500 configurations per second while Polaris spends RTL time only on chosen candidates.
- Online evaluation beats offline evaluation: keeping the high-fidelity simulator in the optimization loop yields designs that are faithful when translated to real hardware, not just optimal under the proxy.
- Polaris reaches parity with a 6-hour DOSA run in under 35 minutes and surpasses it by 2.7x in EDP within 3.3 hours.
- Because Starlight's initial accuracy is already high, transfer learning itself provides the head start that makes sample-efficient Bayesian optimization possible.
Reading between the lines
- The same transfer-learning recipe could be reused whenever a cheap low-fidelity estimator and a slow high-fidelity validator share structure, so Starlight's architecture is a template for other accelerator families beyond the one evaluated here.
- The online-versus-offline result implies that even a 0.99-rank-accurate surrogate leaves a systematic fidelity gap; designers should therefore choose in-loop high-fidelity evaluation whenever the RTL time per candidate is affordable.
- If the delay-only experiments promised in the footnote were published, they would either confirm that the method's gains are robust (no energy-term artifact) or bound how much of the 61% saving is due to the shared Timeloop energy model.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Starlight, a transfer-learned deep kernel learning (DKL) model that predicts the energy-delay product (EDP) of the Gemmini DLA, and Polaris, a Bayesian optimization tool that uses Starlight as a surrogate with RTL simulation in the optimization loop. Starlight is first trained as a variational autoencoder with a predictor on Timeloop analytical-model data, and then its encoder is transferred and fine-tuned with FireSim RTL data. The paper claims that Starlight reaches Spearman rank correlation 0.99 against FireSim, trains with 61% fewer high-fidelity samples than the DOSA baseline, and supports Polaris in finding designs that reduce EDP by 2.7x over DOSA within 3.3 hours. The evaluation compares Polaris against Offline Random, DOSA, and Spotlight on four workloads, and includes wall-clock time comparisons. The central methodology is promising, but the current manuscript leaves key quantitative claims ambiguous because of an unresolved dataset-size inconsistency and because the delay-only validation that rules out label contamination appears only in a footnote without supporting results.
Significance. If the claims are correct, this is a significant contribution to accelerator design-space exploration: it is the first work to transfer a performance model trained on a low-fidelity analytical model to predict high-fidelity RTL outcomes, and it provides evidence that online RTL-in-the-loop optimization can outperform offline proxy optimization. The paper also gives creditworthy concrete predictions, including 6,500 predictions per second, Spearman rho values, and wall-clock times in Table II, and it compares against two strong baselines (DOSA and Spotlight). The principal weakness is verifiability: the reported FireSim dataset size (212 samples) is inconsistent with the training-set-size axis in Figure 10, and the claim that delay-only experiments behave identically is asserted but never shown. These issues affect the headline accuracy and sample-efficiency claims, so the paper cannot be accepted in its current form.
major comments (3)
- [Section IV-B, Section VII-A2, Figure 10] Section IV-B states that the authors collected 216 Timeloop samples and 212 FireSim samples, but Figure 10 plots training-set sizes up to 1,600 and Section VII-A2 refers to 'the full training set.' With an 80/20 split of 212 samples, only about 170 training samples would be available, making the x-axis in Figure 10 impossible. If the dataset actually contains at least 1,600 samples, then Section IV-B undercounts it by roughly 7.5x, which would materially change the 61% sample-savings claim and the transfer-learning comparison. Please reconcile the dataset size, report the exact number of training samples used for each curve in Figure 10, and recompute the headline sample-efficiency numbers accordingly.
- [Section IV-B footnote 2, Section IV-C] The transferability justification in Section IV-C is based on KL divergence between Timeloop-EDP and the hybrid EDP label (Timeloop energy x FireSim delay). Because both distributions share the Timeloop energy term, a low KL divergence does not demonstrate that the delay signal transfers. Footnote 2 claims that all experiments were reproduced using delay-only measurements and that the behavior is identical, but no delay-only results are presented. Please provide the delay-only versions of the accuracy and training-set-size experiments, report the delay-only KL divergences, and clarify in the abstract and Section I that the EDP label is a hybrid measure rather than one measured entirely by RTL simulation.
- [Abstract, Section I, Section VIII] The 61% sample-savings claim is repeated in the Abstract, Section I, and Section VIII, but the manuscript never defines the comparison precisely: it does not state DOSA's training-set size, the number of FireSim samples used to train Starlight in the final configuration, or the formula used to compute 61%. Given the dataset-size ambiguity in Figure 10, this central claim is not evaluable as written. Please add a table or explicit sentence that states the exact sample counts for Starlight and the DOSA baseline and shows how 61% is derived.
minor comments (5)
- [Section IV-B] The sentence 'The datasets are collected by performing Sobol sampling [59] cut for space: —a sampling method...' contains a broken phrase 'cut for space:' and should be reworded.
- [Section VI-B, footnote 3] Footnote 3 contains the typo 'coorelation coefficient' and should read 'correlation coefficient.'
- [Section VII-A1] The text says 'Starlight achieves rho >= 0.98 after just 100 trials,' but the context is about epochs of training; this should be '100 epochs' to avoid confusion with the independent trials used for variance reporting.
- [Section V-B, Table I] Section V-B says the hardware design space has '8x32x32 designs,' but Table I lists four spatial-array choices, 32 accumulator sizes, and 32 scratchpad sizes, which is 4x32x32 = 4,096 designs, not 8x32x32. Please correct the count.
- [Table II] Table II uses dashes for Spotlight in the software-DSE rows for ResNet-50 and BERT without an explanation; please state why these entries are missing.
Circularity Check
Transfer-learning justification leans on a hybrid EDP label that shares the Timeloop energy term with the source model; the delay-only control is asserted but not presented, so the transfer evidence is partly self-referential while the core empirical claims retain independent content.
-
self definitional
[Section IV-B (footnote 2) and Section IV-C]
"A limitation of our training data, and consequently of Starlight, is that FireSim does not measure energy consumption, so like prior work [22], we measure energy consumption using Timeloop. For the remainder of this paper, EDP refers to the product of energy consumption as measured by Timeloop and delay as measured by FireSim. 2 To ensure the shared energy consumption measurement from Timeloop is not contaminating Starlight, we reproduce all experiments using only delay measurements, which are independently measured by Timeloop and FireSim."
Starlight-Low is trained to predict Timeloop EDP, while Starlight's target label is defined as Timeloop energy multiplied by FireSim delay. These two labels share the entire energy term by construction. Section IV-C then uses the small KL divergence (0.04) between the two EDP distributions as evidence that transfer learning is applicable, but that distributional similarity is partly manufactured by the common energy factor. The only ablation that would remove the shared term, the delay-only reproduction promised in footnote 2, is asserted but no results are shown. Thus the transfer-learning justification does not independently establish that the delay signal transfers; it leans on a label that is partially self-same with the source model's output.
full rationale
The central empirical claims—Starlight's rank correlation of 0.99 on held-out FireSim hybrid EDP labels and Polaris's 2.7x EDP improvement over DOSA—are genuine measurements against held-out configurations and against external/prior baselines, not quantities recovered from the model's own training targets by construction. The Polaris online Bayesian optimization contribution is independent of the transfer-learning framing: it integrates RTL simulation in the loop, and its advantage is demonstrated against Offline Random, DOSA, and Spotlight. Self-citations to DOSA [22] and Spotlight [52] are used as baselines or as sources of a standard VAE-with-predictor technique, not as load-bearing uniqueness arguments. The dataset-size inconsistency between Section IV-B (212 FireSim samples) and Figure 10 (training-set sizes up to 1,600) is a serious verifiability and reproducibility concern, but it is not circularity. The hybrid EDP definition does, however, partially contaminate the transferability evidence: because both the source and target labels share the Timeloop energy term, the KL-divergence justification in Section IV-C is weaker than it appears, and the delay-only control is mentioned but not delivered. This makes the transfer-learning rationale partially self-referential, but it does not reduce the whole derivation to its inputs, so a moderate score of 4 is appropriate.
Assumptions & free parameters
free parameters (5)
- Latent space dimensionality =
2
- Number of BO iterations (n,m) =
n=8 hardware, m=6 software
- Sobol samples per software iteration =
10,000
- Training/validation split =
80/20
- Number of Polaris trials =
3
assumptions (5)
- domain assumption Timeloop analytical model is a faithful low-fidelity proxy for DLA performance.
- domain assumption FireSim cycle-exact simulation accurately measures DLA delay.
- ad hoc to paper The hybrid EDP (Timeloop energy x FireSim delay) is a valid objective; shared energy does not dominate the transfer signal.
- ad hoc to paper The three software constraints do not exclude optimal mappings.
- ad hoc to paper KL divergence of 0.04 (overall) and 0.12 (lowest 10%) indicates sufficient transferability.
Cite this review
Pith. "Pith review of Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators." pith.science (2026). https://pith.science/paper/PDYGGH2K
@misc{pith2026241215548,
author = {Pith},
title = {Pith review of: Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators},
year = {2026},
howpublished = {\url{https://pith.science/paper/PDYGGH2K}},
note = {Machine review of arXiv:2412.15548}
}
abstract
This paper presents a tool for automatically exploring the design space of deep learning accelerators (DLAs). Our main advancement is Starlight, a data-driven performance model that uses transfer learning to bridge the gap between fast, low-fidelity evaluation methods (such as analytical models) and slow, high-fidelity evaluation methods (such as RTL simulation). Starlight is fast: It can provide 6,500 predictions per second, allowing the evaluation of millions of configurations per hour. Starlight is accurate: It predicts the energy-delay product measured by RTL simulation with 99\% accuracy. And Starlight can be trained efficiently: It can be trained with 61\% fewer samples than DOSA's state-of-the-art data-driven performance predictor. Our second contribution is Polaris, a design-space exploration tool that uses Starlight to efficiently search the large, complex hardware/software co-design space of DLAs. In under 35 minutes, Polaris produces DLA designs that match the performance of designs that take six hours to produce with DOSA. And in under 3.3 hours, Polaris produces DLA designs that reduce energy-delay product by 2.7$\times$ over the best designs found by DOSA.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 2 Pith papers
-
Fovea: Physical-Implication-Aware Wafer-Scale DSE with Decision-Domain-Guided Cross-Fidelity Refinement
A wafer-scale design-space exploration method constructs physically feasible design spaces and prunes candidates by a sampled evaluator-disagreement bound, recovering the exhaustive reference optimum in all 70 tested ...
-
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
DiffAxE uses conditional diffusion models to generate hardware accelerator designs directly from target performance, achieving orders-of-magnitude faster design space exploration with lower error than existing optimiz...
Reference graph
Works this paper leans on
-
[1]
D. Abts, I. Ahmed, A. Bitar, M. Boyd, J. Kim, G. Kimmell, and A. Ling, “Challenges/Opportunities to Enable Dependable Scale-out System with Groq Deterministic Tensor-Streaming Processors,” in Dependable Sys- tems and Networks (DSN-S) , Jun. 2022
work page 2022
-
[2]
BOOM- Explorer: RISC-V BOOM Microarchitecture Design Space Exploration Framework,
C. Bai, Q. Sun, J. Zhai, Y . Ma, B. Yu, and M. D. Wong, “BOOM- Explorer: RISC-V BOOM Microarchitecture Design Space Exploration Framework,” in International Conference On Computer-Aided Design (ICCAD), Nov. 2021
work page 2021
-
[3]
Transfer Learning for Bayesian Optimization: A Survey,
T. Bai, Y . Li, Y . Shen, X. Zhang, W. Zhang, and B. Cui, “Transfer Learning for Bayesian Optimization: A Survey,” arXiv, Feb. 2023
work page 2023
-
[4]
D. Bank, N. Koenigstein, and R. Giryes, “Autoencoders,” in Data Mining and Knowledge Discovery Handbook , L. Rokach, O. Maimon, and E. Shmueli, Eds., 2023
work page 2023
-
[5]
Hyperparameter Optimization: Foundations, Algorithms, Best Practices, and Open Challenges,
B. Bischl, M. Binder, M. Lang, T. Pielok, J. Richter, S. Coors, J. Thomas, T. Ullmann, M. Becker, A.-L. Boulesteix, D. Deng, and M. Lin- dauer, “Hyperparameter Optimization: Foundations, Algorithms, Best Practices, and Open Challenges,” WIREs Data Mining and Knowledge Discovery, no. 2, 2023
work page 2023
-
[6]
Eyeriss: An Energy- Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,
Y .-H. Chen, T. Krishna, J. S. Emer, and V . Sze, “Eyeriss: An Energy- Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” Solid-State Circuits, no. 1, Jan. 2017
work page 2017
-
[7]
Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices,
Y .-H. Chen, T.-J. Yang, J. Emer, and V . Sze, “Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices,” Emerging and Selected Topics in Circuits and Systems, no. 2, Jun. 2019
work page 2019
-
[8]
dMazeRun- ner: Executing Perfectly Nested Loops on Dataflow Accelerators,
S. Dave, Y . Kim, S. Avancha, K. Lee, and A. Shrivastava, “dMazeRun- ner: Executing Perfectly Nested Loops on Dataflow Accelerators,” Transactions on Embedded Computing Systems , no. 5s, Oct. 2019
work page 2019
Show all 71 references
-
[9]
BERT: Pre- training of Deep Bidirectional Transformers for Language Understand- ing,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of Deep Bidirectional Transformers for Language Understand- ing,” arXiv, May 2019
2019
-
[10]
Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey,
P. Dhilleswararao, S. Boppu, M. S. Manikandan, and L. R. Cenkera- maddi, “Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey,” IEEE Access, 2022
2022
-
[11]
A Survey on Deep Learning and Its Applications,
S. Dong, P. Wang, and K. Abbas, “A Survey on Deep Learning and Its Applications,” Computer Science Review , 2021
2021
-
[12]
Acceler- ating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,
M. Emani, V . Vishwanath, C. Adams, M. E. Papka, R. Stevens, L. Florescu, S. Jairath, W. Liu, T. Nama, and A. Sujeeth, “Acceler- ating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,” Computing in Science & Engineering , no. 2, Mar. 2021
2021
-
[13]
An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators,
H. Esmaeilzadeh, S. Ghodrati, A. B. Kahng, J. K. Kim, S. Kinzer, S. Kundu, R. Mahapatra, S. D. Manasi, S. Sapatnekar, Z. Wang, and Z. Zeng, “An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators,” arXiv, Aug. 2023
2023
-
[14]
Physically Accurate Learning-Based Performance Prediction of Hardware-Accelerated ML Algorithms,
H. Esmaeilzadeh, S. Ghodrati, A. B. Kahng, J. K. Kim, S. Kinzer, S. Kundu, R. Mahapatra, S. D. Manasi, S. S. Sapatnekar, Z. Wang, and Z. Zeng, “Physically Accurate Learning-Based Performance Prediction of Hardware-Accelerated ML Algorithms,” in Workshop on Machine Learning for...
2022
-
[15]
Improving Performance Estimation for Design Space Exploration for Convolutional Neural Network Accelerators,
M. Ferianc, H. Fan, D. Manocha, H. Zhou, S. Liu, X. Niu, and W. Luk, “Improving Performance Estimation for Design Space Exploration for Convolutional Neural Network Accelerators,” Electronics, no. 4, Jan. 2021
2021
-
[16]
Practical Transfer Learning for Bayesian Optimization,
M. Feurer, B. Letham, F. Hutter, and E. Bakshy, “Practical Transfer Learning for Bayesian Optimization,” arXiv, Oct. 2022
2022
-
[17]
Tests for Rank Correlation Coefficients, I,
E. C. Fieller, H. O. Hartley, and E. S. Pearson, “Tests for Rank Correlation Coefficients, I,” Biometrika, no. 3-4, Dec. 1957
1957
-
[18]
Multi-fidelity Optimiza- tion via Surrogate Modelling,
A. I. Forrester, A. S ´obester, and A. J. Keane, “Multi-fidelity Optimiza- tion via Surrogate Modelling,” Royal Society A: Mathematical, Physical and Engineering Sciences , no. 2088, Oct. 2007
2007
-
[19]
Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack Integration,
H. Genc, S. Kim, A. Amid, A. Haj-Ali, V . Iyer, P. Prakash, J. Zhao, D. Grubb, H. Liew, H. Mao, A. Ou, C. Schmidt, S. Steffl, J. Wright, I. Stoica, J. Ragan-Kelley, K. Asanovic, B. Nikolic, and Y . S. Shao, “Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation vi...
2021
-
[20]
Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules,
R. G ´omez-Bombarelli, J. N. Wei, D. Duvenaud, J. M. Hern ´andez- Lobato, B. S ´anchez-Lengeling, D. Sheberla, J. Aguilera-Iparraguirre, T. D. Hirzel, R. P. Adams, and A. Aspuru-Guzik, “Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules,” ACS ...
2018
-
[21]
Deep Residual Learning for Image Recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Computer Vision and Pattern Recognition (CVPR) , 2016
2016
-
[22]
DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators,
C. Hong, Q. Huang, G. Dinh, M. Subedar, and Y . S. Shao, “DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators,” in Microarchitecture (MICRO), Dec. 2023
2023
-
[23]
Learning A Continuous and Reconstructible Latent Space for Hard- ware Accelerator Design,
Q. Huang, C. Hong, J. Wawrzynek, M. Subedar, and Y . S. Shao, “Learning A Continuous and Reconstructible Latent Space for Hard- ware Accelerator Design,” in International Symposium on Performance Analysis of Systems and Software (ISPASS) , May 2022
2022
-
[24]
Ten Lessons From Three Generations Shaped Google’s TPUv4i : Industrial Product,
N. P. Jouppi, D. Hyun Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian, J. Laudon, S. Li, P. Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Young, Z. Zhou, and D. Patterson, “Ten Lessons From Three Generations Shaped Google’s TPUv4i : Industrial Product,” in Internationa...
2021
-
[25]
In-Datacenter Performance Analysis of a Tensor Processing Unit,
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V . Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. ...
2017
-
[26]
ConfuciuX: Autonomous Hard- ware Resource Assignment for DNN Accelerators using Reinforcement Learning,
S.-C. Kao, G. Jeong, and T. Krishna, “ConfuciuX: Autonomous Hard- ware Resource Assignment for DNN Accelerators using Reinforcement Learning,” in Microarchitecture (MICRO), Oct. 2020
2020
-
[27]
Firesim: FPGA- Accelerated Cycle-Exact Scale-Out System Simulation in the Public Cloud,
S. Karandikar, H. Mao, D. Kim, D. Biancolin, A. Amid, D. Lee, N. Pemberton, E. Amaro, C. Schmidt, A. Chopra, Q. Huang, K. Kovacs, B. Nikolic, R. Katz, J. Bachrach, and K. Asanovi ´c, “Firesim: FPGA- Accelerated Cycle-Exact Scale-Out System Simulation in the Public Cloud,” in I...
2018
-
[28]
A Learned Performance Model for Tensor Processing Units,
S. Kaufman, P. Phothilimthana, Y . Zhou, C. Mendis, S. Roy, A. Sabne, and M. Burrows, “A Learned Performance Model for Tensor Processing Units,” in Machine Learning and Systems , A. Smola, A. Dimakis, and I. Stoica, Eds., 2021
2021
-
[29]
Full Stack Optimization of Transformer Inference: A Survey,
S. Kim, C. Hooper, T. Wattanawong, M. Kang, R. Yan, H. Genc, G. Dinh, Q. Huang, K. Keutzer, M. W. Mahoney, Y . S. Shao, and A. Gholami, “Full Stack Optimization of Transformer Inference: A Survey,” arXiv, Feb. 2023
2023
-
[30]
Auto-Encoding Variational Bayes,
D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” arXiv, Dec. 2022
2022
-
[31]
Spatial: A language and compiler for application accelerators,
D. Koeplinger, M. Feldman, R. Prabhakar, Y . Zhang, S. Hadjis, R. Fiszel, T. Zhao, L. Nardi, A. Pedram, C. Kozyrakis, and K. Olukotun, “Spatial: A language and compiler for application accelerators,” in Programming Language Design and Implementation (PLDI) , Jun. 2018
2018
-
[32]
ArchGym: An Open-Source Gymnasium for Machine Learning Assisted Architecture Design,
S. Krishnan, A. Yazdanbakhsh, S. Prakash, J. Jabbour, I. Uchendu, S. Ghosh, B. Boroujerdian, D. Richins, D. Tripathy, A. Faust, and V . Janapa Reddi, “ArchGym: An Open-Source Gymnasium for Machine Learning Assisted Architecture Design,” in International Symposium on Computer A...
2023
-
[33]
On Information and Sufficiency,
S. Kullback and R. A. Leibler, “On Information and Sufficiency,” The Annals of Mathematical Statistics , no. 1, 1951
1951
-
[34]
Data-Driven Offline Optimization for Architecting Hardware Accelera- tors,
A. Kumar, A. Yazdanbakhsh, M. Hashemi, K. Swersky, and S. Levine, “Data-Driven Offline Optimization for Architecting Hardware Accelera- tors,” in International Conference on Learning Representations (ICLR) , Oct. 2021
2021
-
[35]
MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,
H. Kwon, P. Chatarasi, V . Sarkar, T. Krishna, M. Pellauer, and A. Parashar, “MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,” IEEE Micro, no. 3, May 2020
2020
-
[36]
MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,
C. Lattner, M. Amini, U. Bondhugula, A. Cohen, A. Davis, J. Pienaar, R. Riddle, T. Shpeisman, N. Vasilache, and O. Zinenko, “MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,” in Code Generation and Optimization (CGO) , Feb. 2021
2021
-
[37]
Powering Extreme-Scale HPC with Cerebras Wafer- Scale Accelerators,
A. Lavely, “Powering Extreme-Scale HPC with Cerebras Wafer- Scale Accelerators,” Cerebras Systems, Inc, Tech. Rep., 2022
2022
-
[38]
A Study of Bayesian Neural Network Surrogates for Bayesian Optimization,
Y . L. Li, T. G. J. Rudner, and A. G. Wilson, “A Study of Bayesian Neural Network Surrogates for Bayesian Optimization,” arXiv, May 2023
2023
-
[39]
Focal Loss for Dense Object Detection,
T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal Loss for Dense Object Detection,” in International Conference on Computer Vision (ICCV), 2017
2017
-
[40]
NAAS: Neural Accelerator Architecture Search,
Y . Lin, M. Yang, and S. Han, “NAAS: Neural Accelerator Architecture Search,” in Design Automation Conference (DAC) , Dec. 2021
2021
-
[41]
The AI Index 2023 Annual Report,
N. Maslej, L. Fattorini, E. Brynjolfsson, J. Etchemendy, K. Ligett, T. Lyons, J. Manyika, H. Ngo, V . Parli, Y . Shoham, R. Wald, J. Clark, and R. Perrault, “The AI Index 2023 Annual Report,” Institute for Human-Centered AI, Tech. Rep., Apr. 2023
2023
-
[42]
Mat ´ern, Spatial Variation, D
B. Mat ´ern, Spatial Variation, D. Brillinger, S. Fienberg, J. Gani, J. Har- tigan, and K. Krickeberg, Eds., 1986
1986
-
[43]
ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,
L. Mei, P. Houshmand, V . Jain, S. Giraldo, and M. Verhelst, “ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,” Transactions on Computers, no. 8, Aug. 2021
2021
-
[44]
STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators,
F. Mu ˜noz-Mart´ınez, J. L. Abell ´an, M. E. Acacio, and T. Krishna, “STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators,” in International Symposium on Workload Characterization (IISWC), Nov. 2021
2021
-
[45]
Practical Design Space Exploration,
L. Nardi, D. Koeplinger, and K. Olukotun, “Practical Design Space Exploration,” in Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS), Oct. 2019
2019
-
[46]
Timeloop: A Systematic Approach to DNN Accelerator Evaluation,
A. Parashar, P. Raina, Y . S. Shao, Y .-H. Chen, V . A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A Systematic Approach to DNN Accelerator Evaluation,” in International Symposium on Performance Analysis of Systems and Software (ISPASS...
2019
-
[47]
Hardware/Software Co-design for Convolutional Neural Networks Acceleration: A Sur- vey and Open Issues,
C. Pham-Quoc, X.-Q. Nguyen, and T. N. Thinh, “Hardware/Software Co-design for Convolutional Neural Networks Acceleration: A Sur- vey and Open Issues,” in Context-Aware Systems and Applications , P. Cong Vinh and A. Rakib, Eds., 2021
2021
-
[48]
A Case for Efficient Accelerator Design Space Exploration via Bayesian Optimization,
B. Reagen, J. M. Hern ´andez-Lobato, R. Adolf, M. Gelbart, P. What- mough, G.-Y . Wei, and D. Brooks, “A Case for Efficient Accelerator Design Space Exploration via Bayesian Optimization,” in International Symposium on Low Power Electronics and Design (ISLPED) , Jul. 2017
2017
-
[49]
Stochastic Backprop- agation and Approximate Inference in Deep Generative Models,
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic Backprop- agation and Approximate Inference in Deep Generative Models,” in International Conference on Machine Learning (ICML) , Jun. 2014
2014
-
[50]
U-Net: Convolutional Net- works for Biomedical Image Segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Net- works for Biomedical Image Segmentation,” in Medical Image Comput- ing and Computer-Assisted Intervention (MICCAI), N. Navab, J. Horneg- ger, W. M. Wells, and A. F. Frangi, Eds., 2015
2015
-
[51]
Learning In- ternal Representations by Error Propagation,
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning In- ternal Representations by Error Propagation,” in Explorations in the Microstructure of Cognition, Jan. 1986
1986
-
[52]
Leveraging Domain Information for the Efficient Automated Design of Deep Learning Accelerators,
C. Sakhuja, Z. Shi, and C. Lin, “Leveraging Domain Information for the Efficient Automated Design of Deep Learning Accelerators,” in High- Performance Computer Architecture (HPCA) , Feb. 2023
2023
-
[53]
AIrchitect: Automating Hardware Architecture and Mapping Optimization,
A. Samajdar, J. M. Joseph, and T. Krishna, “AIrchitect: Automating Hardware Architecture and Mapping Optimization,” in Design, Automa- tion & Test in Europe Conference & Exhibition (DATE) , Apr. 2023
2023
-
[54]
A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim,
A. Samajdar, J. M. Joseph, Y . Zhu, P. Whatmough, M. Mattina, and T. Krishna, “A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim,” in International Symposium on Performance Analysis of Systems and Software (ISPASS), Aug. 2020
2020
-
[55]
Neural Architecture Search and Hardware Accelerator Co- Search: A Survey,
L. Sekanina, “Neural Architecture Search and Hardware Accelerator Co- Search: A Survey,” IEEE Access, 2021
2021
-
[56]
An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks,
K. Seshadri, B. Akin, J. Laudon, R. Narayanaswami, and A. Yazdan- bakhsh, “An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks,” in International Symposium on Workload Character- ization (IISWC), Nov. 2022
2022
-
[57]
Simba: Scaling Deep-Learning Inference with Multi-Chip-Module- Based Architecture,
Y . S. Shao, J. Clemons, R. Venkatesan, B. Zimmer, M. Fojtik, N. Jiang, B. Keller, A. Klinefelter, N. Pinckney, P. Raina, S. G. Tell, Y . Zhang, W. J. Dally, J. Emer, C. T. Gray, B. Khailany, and S. W. Keckler, “Simba: Scaling Deep-Learning Inference with Multi-Chip-Module- Ba...
2019
-
[58]
Artificial Intelli- gence in the IoT Era: A Review of Edge AI Hardware and Software,
T. Sipola, J. Alatalo, T. Kokkonen, and M. Rantonen, “Artificial Intelli- gence in the IoT Era: A Review of Edge AI Hardware and Software,” in Conference of Open Innovations Association (FRUCT) , Apr. 2022
2022
-
[59]
On the Distribution of Points in a Cube and the Ap- proximate Evaluation of Integrals,
I. M. Sobol’, “On the Distribution of Points in a Cube and the Ap- proximate Evaluation of Integrals,” Zhurnal Vychislitel’noi Matematiki i Matematicheskoi Fiziki , no. 4, 1967
1967
-
[60]
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental De- sign,
N. Srinivas, A. Krause, S. Kakade, and M. Seeger, “Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental De- sign,” in International Conference on Machine Learning (ICML) , Jun. 2010
2010
-
[61]
Automated Design of Deep Neural Networks: A Survey and Unified Taxonomy,
E.-G. Talbi, “Automated Design of Deep Neural Networks: A Survey and Unified Taxonomy,” Computing Surveys, no. 2, Mar. 2021
2021
-
[62]
Compute Substrate for Software 2.0,
J. Vasiljevic, L. Bajic, D. Capalija, S. Sokorac, D. Ignjatovic, L. Bajic, M. Trajkovic, I. Hamer, I. Matosevic, A. Cejkov, U. Aydonat, T. Zhou, S. Z. Gilani, A. Paiva, J. Chu, D. Maksimovic, S. A. Chin, Z. Moudallal, A. Rakhmati, S. Nijjar, A. Bhullar, B. Drazic, C. Lee, J. S...
2021
-
[63]
MAGNet: A Modular Accelerator Generator for Neural Networks,
R. Venkatesan, Y . S. Shao, M. Wang, J. Clemons, S. Dai, M. Fojtik, B. Keller, A. Klinefelter, N. Pinckney, P. Raina, Y . Zhang, B. Zimmer, W. J. Dally, J. Emer, S. W. Keckler, and B. Khailany, “MAGNet: A Modular Accelerator Generator for Neural Networks,” in International Con...
2019
-
[64]
Deep Kernel Learning,
A. G. Wilson, Z. Hu, R. Salakhutdinov, and E. P. Xing, “Deep Kernel Learning,” in Artificial Intelligence and Statistics (AISTATS), May 2016
2016
-
[65]
Few-Shot Bayesian Optimization with Deep Kernel Surrogates,
M. Wistuba and J. Grabocka, “Few-Shot Bayesian Optimization with Deep Kernel Surrogates,” in International Conference on Learning Representations (ICLR), Oct. 2020
2020
-
[66]
SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads,
S. L. Xi, Y . Yao, K. Bhardwaj, P. Whatmough, G.-Y . Wei, and D. Brooks, “SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads,” Transactions on Architecture and Code Optimiza- tion (TACO), no. 4, Nov. 2020
2020
-
[67]
HASCO: Towards Agile HArdware and Software CO-design for Tensor Compu- tation,
Q. Xiao, S. Zheng, B. Wu, P. Xu, X. Qian, and Y . Liang, “HASCO: Towards Agile HArdware and Software CO-design for Tensor Compu- tation,” in International Symposium on Computer Architecture (ISCA) , Jun. 2021
2021
-
[68]
Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,
X. Yang, M. Gao, Q. Liu, J. Setter, J. Pu, A. Nayak, S. Bell, K. Cao, H. Ha, P. Raina, C. Kozyrakis, and M. Horowitz, “Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,” in Ar- chitectural Support for Programming Languages and Operating Systems (ASP...
2020
-
[69]
Apollo: Transferable Architecture Exploration,
A. Yazdanbakhsh, C. Angermueller, B. Akin, Y . Zhou, A. Jones, M. Hashemi, K. Swersky, S. Chatterjee, R. Narayanaswami, and J. Laudon, “Apollo: Transferable Architecture Exploration,” Workshop on ML for Systems , 2020
2020
-
[70]
A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators,
D. Zhang, S. Huda, E. Songhori, K. Prabhu, Q. Le, A. Goldie, and A. Mirhoseini, “A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators,” in Architectural Support for Programming Languages and Operating Systems (ASPLOS) , Feb. 2022
2022
-
[71]
A Comprehensive Survey on Transfer Learning,
F. Zhuang, Z. Qi, K. Duan, D. Xi, Y . Zhu, H. Zhu, H. Xiong, and Q. He, “A Comprehensive Survey on Transfer Learning,” IEEE, no. 1, Jan. 2021
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.