Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A transferred surrogate predicts RTL-level energy-delay with 99% rank accuracy using 61% fewer simulations, and keeping RTL simulation in the optimization loop yields designs 2.7x better in energy-delay product.

desk verdict Useful integration of transfer learning and online RTL-in-the-loop BO for DLA design, but the headline accuracy and sample-efficiency claims are unverifiable as written due to a 212-vs-1600 dataset inconsistency and a missing delay-only ablation. read the letter →

arxiv 2412.15548 v1 pith:PDYGGH2K submitted 2024-12-20 cs.AR

classification cs.AR
keywords deeplearningacceleratorsdesignspaceexplorationtransferkernelBayesianoptimizationRTLsimulationperformancemodelinghardware/softwareco-design
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that transfer learning can break the accuracy-versus-speed tradeoff in evaluating deep learning accelerator designs. It introduces Starlight, a performance model trained on cheap analytical estimates and then fine-tuned on a small set of slow RTL simulations, reaching 99% rank correlation with RTL-measured energy-delay product using 61% fewer high-fidelity samples than the prior state of the art. The paper also builds Polaris, a Bayesian optimizer that keeps RTL simulation inside the optimization loop; in under 3.3 hours it finds accelerator designs and software mappings that beat the best designs of the offline baseline DOSA by 2.7x in energy-delay product. If true, this means accurate design evaluation no longer requires huge one-time simulation budgets, and online evaluation produces hardware-faithful designs.

What carries the argument

Starlight is built from deep kernel learning (DKL): a variational autoencoder's encoder network (trained with a predictor head that imposes a smooth EDP gradient on the latent space) is transferred by hard weight sharing from Starlight-Low, then attached to a Gaussian process with a Matérn kernel that supplies the uncertainty estimate needed for Bayesian optimization. Polaris wraps this surrogate in an outer hardware loop (enumerating discrete array/scratchpad/accumulator choices) and an inner per-layer software loop (sampling 10,000 Sobol candidates per iteration), selects candidates with an Upper Confidence Bound acquisition function, and evaluates them on the RTL simulator, feeding each result back into Starlight.

What would settle it

Repeat the Starlight training and the Polaris search using delay-only labels, comparing FireSim delay against Timeloop delay with no shared energy term; if the rank correlation or the 61% sample savings drop materially, the central transfer-learning claim is inflated by the shared energy model, and if they hold, the claim is robust.

Watch

Extended reading notes

Core claim

The paper's central claim is that the encoder of a variational autoencoder trained on cheap analytical-model evaluations of a DLA (Starlight-Low) can be transplanted into a deep-kernel-learning model and fine-tuned on a small set of RTL-simulation evaluations to produce Starlight, a surrogate that predicts RTL-measured energy-delay product with Spearman rank correlation 0.99 while using 61% fewer high-fidelity samples than the DOSA baseline. On top of that, the paper claims Polaris—a Bayesian optimizer that keeps an RTL simulator inside the optimization loop—consistently finds designs with lower EDP than offline optimizers, beating DOSA's best designs by an average of 2.7x in under 3.3 hours. The authors argue this is the first demonstration that RTL simulation in the loop, rather than only as a final check, materially improves the quality of the produced hardware/software co-designs.

Load-bearing premise

The load-bearing assumption is that the knowledge transferred from the analytical model to the RTL simulator comes from genuinely shared performance structure, but because the hybrid energy-delay label uses the same analytical-model energy term in both source and target, the transfer gain could be partly an artifact of that shared measurement.

Editorial extensions

If this is right

  • Training a high-fidelity DLA performance model can be done with 61% fewer RTL simulations, reducing the one-time data-collection bottleneck.
  • Design space exploration can use RTL simulation as the evaluator without giving up search breadth, because Starlight evaluates ~6,500 configurations per second while Polaris spends RTL time only on chosen candidates.
  • Online evaluation beats offline evaluation: keeping the high-fidelity simulator in the optimization loop yields designs that are faithful when translated to real hardware, not just optimal under the proxy.
  • Polaris reaches parity with a 6-hour DOSA run in under 35 minutes and surpasses it by 2.7x in EDP within 3.3 hours.
  • Because Starlight's initial accuracy is already high, transfer learning itself provides the head start that makes sample-efficient Bayesian optimization possible.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same transfer-learning recipe could be reused whenever a cheap low-fidelity estimator and a slow high-fidelity validator share structure, so Starlight's architecture is a template for other accelerator families beyond the one evaluated here.
  • The online-versus-offline result implies that even a 0.99-rank-accurate surrogate leaves a systematic fidelity gap; designers should therefore choose in-loop high-fidelity evaluation whenever the RTL time per candidate is affordable.
  • If the delay-only experiments promised in the footnote were published, they would either confirm that the method's gains are robust (no energy-term artifact) or bound how much of the 61% saving is due to the shared Timeloop energy model.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents Starlight, a transfer-learned deep kernel learning (DKL) model that predicts the energy-delay product (EDP) of the Gemmini DLA, and Polaris, a Bayesian optimization tool that uses Starlight as a surrogate with RTL simulation in the optimization loop. Starlight is first trained as a variational autoencoder with a predictor on Timeloop analytical-model data, and then its encoder is transferred and fine-tuned with FireSim RTL data. The paper claims that Starlight reaches Spearman rank correlation 0.99 against FireSim, trains with 61% fewer high-fidelity samples than the DOSA baseline, and supports Polaris in finding designs that reduce EDP by 2.7x over DOSA within 3.3 hours. The evaluation compares Polaris against Offline Random, DOSA, and Spotlight on four workloads, and includes wall-clock time comparisons. The central methodology is promising, but the current manuscript leaves key quantitative claims ambiguous because of an unresolved dataset-size inconsistency and because the delay-only validation that rules out label contamination appears only in a footnote without supporting results.

Significance. If the claims are correct, this is a significant contribution to accelerator design-space exploration: it is the first work to transfer a performance model trained on a low-fidelity analytical model to predict high-fidelity RTL outcomes, and it provides evidence that online RTL-in-the-loop optimization can outperform offline proxy optimization. The paper also gives creditworthy concrete predictions, including 6,500 predictions per second, Spearman rho values, and wall-clock times in Table II, and it compares against two strong baselines (DOSA and Spotlight). The principal weakness is verifiability: the reported FireSim dataset size (212 samples) is inconsistent with the training-set-size axis in Figure 10, and the claim that delay-only experiments behave identically is asserted but never shown. These issues affect the headline accuracy and sample-efficiency claims, so the paper cannot be accepted in its current form.

major comments (3)
  1. [Section IV-B, Section VII-A2, Figure 10] Section IV-B states that the authors collected 216 Timeloop samples and 212 FireSim samples, but Figure 10 plots training-set sizes up to 1,600 and Section VII-A2 refers to 'the full training set.' With an 80/20 split of 212 samples, only about 170 training samples would be available, making the x-axis in Figure 10 impossible. If the dataset actually contains at least 1,600 samples, then Section IV-B undercounts it by roughly 7.5x, which would materially change the 61% sample-savings claim and the transfer-learning comparison. Please reconcile the dataset size, report the exact number of training samples used for each curve in Figure 10, and recompute the headline sample-efficiency numbers accordingly.
  2. [Section IV-B footnote 2, Section IV-C] The transferability justification in Section IV-C is based on KL divergence between Timeloop-EDP and the hybrid EDP label (Timeloop energy x FireSim delay). Because both distributions share the Timeloop energy term, a low KL divergence does not demonstrate that the delay signal transfers. Footnote 2 claims that all experiments were reproduced using delay-only measurements and that the behavior is identical, but no delay-only results are presented. Please provide the delay-only versions of the accuracy and training-set-size experiments, report the delay-only KL divergences, and clarify in the abstract and Section I that the EDP label is a hybrid measure rather than one measured entirely by RTL simulation.
  3. [Abstract, Section I, Section VIII] The 61% sample-savings claim is repeated in the Abstract, Section I, and Section VIII, but the manuscript never defines the comparison precisely: it does not state DOSA's training-set size, the number of FireSim samples used to train Starlight in the final configuration, or the formula used to compute 61%. Given the dataset-size ambiguity in Figure 10, this central claim is not evaluable as written. Please add a table or explicit sentence that states the exact sample counts for Starlight and the DOSA baseline and shows how 61% is derived.
minor comments (5)
  1. [Section IV-B] The sentence 'The datasets are collected by performing Sobol sampling [59] cut for space: —a sampling method...' contains a broken phrase 'cut for space:' and should be reworded.
  2. [Section VI-B, footnote 3] Footnote 3 contains the typo 'coorelation coefficient' and should read 'correlation coefficient.'
  3. [Section VII-A1] The text says 'Starlight achieves rho >= 0.98 after just 100 trials,' but the context is about epochs of training; this should be '100 epochs' to avoid confusion with the independent trials used for variance reporting.
  4. [Section V-B, Table I] Section V-B says the hardware design space has '8x32x32 designs,' but Table I lists four spatial-array choices, 32 accumulator sizes, and 32 scratchpad sizes, which is 4x32x32 = 4,096 designs, not 8x32x32. Please correct the count.
  5. [Table II] Table II uses dashes for Spotlight in the software-DSE rows for ResNet-50 and BERT without an explanation; please state why these entries are missing.

Circularity Check

1 steps flagged · score 4.0 of 10

Transfer-learning justification leans on a hybrid EDP label that shares the Timeloop energy term with the source model; the delay-only control is asserted but not presented, so the transfer evidence is partly self-referential while the core empirical claims retain independent content.

  1. self definitional [Section IV-B (footnote 2) and Section IV-C]
    "A limitation of our training data, and consequently of Starlight, is that FireSim does not measure energy consumption, so like prior work [22], we measure energy consumption using Timeloop. For the remainder of this paper, EDP refers to the product of energy consumption as measured by Timeloop and delay as measured by FireSim. 2 To ensure the shared energy consumption measurement from Timeloop is not contaminating Starlight, we reproduce all experiments using only delay measurements, which are independently measured by Timeloop and FireSim."

    Starlight-Low is trained to predict Timeloop EDP, while Starlight's target label is defined as Timeloop energy multiplied by FireSim delay. These two labels share the entire energy term by construction. Section IV-C then uses the small KL divergence (0.04) between the two EDP distributions as evidence that transfer learning is applicable, but that distributional similarity is partly manufactured by the common energy factor. The only ablation that would remove the shared term, the delay-only reproduction promised in footnote 2, is asserted but no results are shown. Thus the transfer-learning justification does not independently establish that the delay signal transfers; it leans on a label that is partially self-same with the source model's output.

full rationale

The central empirical claims—Starlight's rank correlation of 0.99 on held-out FireSim hybrid EDP labels and Polaris's 2.7x EDP improvement over DOSA—are genuine measurements against held-out configurations and against external/prior baselines, not quantities recovered from the model's own training targets by construction. The Polaris online Bayesian optimization contribution is independent of the transfer-learning framing: it integrates RTL simulation in the loop, and its advantage is demonstrated against Offline Random, DOSA, and Spotlight. Self-citations to DOSA [22] and Spotlight [52] are used as baselines or as sources of a standard VAE-with-predictor technique, not as load-bearing uniqueness arguments. The dataset-size inconsistency between Section IV-B (212 FireSim samples) and Figure 10 (training-set sizes up to 1,600) is a serious verifiability and reproducibility concern, but it is not circularity. The hybrid EDP definition does, however, partially contaminate the transferability evidence: because both the source and target labels share the Timeloop energy term, the KL-divergence justification in Section IV-C is weaker than it appears, and the delay-only control is mentioned but not delivered. This makes the transfer-learning rationale partially self-referential, but it does not reduce the whole derivation to its inputs, so a moderate score of 4 is appropriate.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claims rest primarily on the fidelity of Timeloop and FireSim as evaluation methods, the validity of the hybrid EDP metric, and the constraints imposed on the software design space. These are domain assumptions and ad hoc choices rather than mathematical axioms. No new physical entities are introduced.

free parameters (5)
  • Latent space dimensionality = 2
    The VAE encoder in Starlight-Low projects inputs to a 2-D latent space (Section IV-D); chosen by hand for visualization and GP compatibility, this limits model capacity and affects transfer quality.
  • Number of BO iterations (n,m) = n=8 hardware, m=6 software
    Polaris runs for 8 hardware and 6 software iterations per layer in HW/SW co-design (Section VI-C); the 2.7x EDP improvement claim depends on this budget.
  • Sobol samples per software iteration = 10,000
    Polaris randomly draws 10,000 Sobol samples from the software space each iteration and picks the best acquisition value (Section V-C); this heuristic replaces exact acquisition maximization.
  • Training/validation split = 80/20
    Unless otherwise specified, models are trained on 80% of the dataset (Section VI-B); the reported rho=0.99 is sensitive to this split's randomness, mitigated by 10 trials.
  • Number of Polaris trials = 3
    Polaris and Spotlight are run for three independent trials; medians are reported without confidence intervals (Section VI-C).
assumptions (5)
  • domain assumption Timeloop analytical model is a faithful low-fidelity proxy for DLA performance.
    Starlight-Low is trained on Timeloop EDPs (Section IV-B) and the transfer premise relies on Timeloop/FireSim similarity (Section IV-C).
  • domain assumption FireSim cycle-exact simulation accurately measures DLA delay.
    FireSim is the high-fidelity ground truth for Starlight's target and Polaris's online evaluations (Sections IV-B, V-C).
  • ad hoc to paper The hybrid EDP (Timeloop energy x FireSim delay) is a valid objective; shared energy does not dominate the transfer signal.
    Section IV-B defines EDP this way and footnote 2 claims delay-only results are identical, but no delay-only data is shown.
  • ad hoc to paper The three software constraints do not exclude optimal mappings.
    Section V-C limits the software space to implementable, utilization-maximizing, evenly-dividing tilings; if these exclude the global optimum, Polaris's advantage may be an artifact.
  • ad hoc to paper KL divergence of 0.04 (overall) and 0.12 (lowest 10%) indicates sufficient transferability.
    Section IV-C uses these values to justify transfer learning, but no null model or significance test is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators." pith.science (2026). https://pith.science/paper/PDYGGH2K

@misc{pith2026241215548,
  author       = {Pith},
  title        = {Pith review of: Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PDYGGH2K}},
  note         = {Machine review of arXiv:2412.15548}
}
abstract

This paper presents a tool for automatically exploring the design space of deep learning accelerators (DLAs). Our main advancement is Starlight, a data-driven performance model that uses transfer learning to bridge the gap between fast, low-fidelity evaluation methods (such as analytical models) and slow, high-fidelity evaluation methods (such as RTL simulation). Starlight is fast: It can provide 6,500 predictions per second, allowing the evaluation of millions of configurations per hour. Starlight is accurate: It predicts the energy-delay product measured by RTL simulation with 99\% accuracy. And Starlight can be trained efficiently: It can be trained with 61\% fewer samples than DOSA's state-of-the-art data-driven performance predictor. Our second contribution is Polaris, a design-space exploration tool that uses Starlight to efficiently search the large, complex hardware/software co-design space of DLAs. In under 35 minutes, Polaris produces DLA designs that match the performance of designs that take six hours to produce with DOSA. And in under 3.3 hours, Polaris produces DLA designs that reduce energy-delay product by 2.7$\times$ over the best designs found by DOSA.

Figures

Figures reproduced from arXiv: 2412.15548 by the authors.

Figure 1
Figure 1. Analytical models can be queried thousands of times [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 3
Figure 3. A Gaussian process that models a ground truth function [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. A Starlight-Low is a neural network that predicts the energy-delay product (EDP) of a DLA as measured by a low￾fidelity method, namely, an analytical model. The encoder network (in blue dotted pattern) from Starlight-Low is transferred to B Starlight, which is a machine learning model based on deep kernel learning that predicts the EDP as measured by a high-fidelity method, namely, an RTL simulator. The decoder netw… view at source ↗
Figures from the paper (8 more)
Figure 6
Figure 6. Figure 6: The distribution of EDPs for the same 2 12 HW/SW configurations as measured by an analytical model and by an RTL simulator. The similarity of the distributions indicates that knowledge can be transferred between models. A limitation of our training data, and consequent…
Figure 7
Figure 7. Figure 7: Polaris is a DSE tool that takes as input layer shapes that define the workload to be optimized and outputs an optimized [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Starlight predicts EDP measured by RTL simulation, [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Accuracy and Spearman rank correlation ( [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: ρ versus the FireSim training set size. We evaluate four model architectures: (1) Starlight, (2) a neural network that leverages transfer learning, (3) a simple fine-tuning of Starlight-Low, and (4) a DKL trained from scratch. The solid line indicates the mean of ten …
Figure 11
Figure 11. Figure 11: We compare the best designs produced by Offline Random, DOSA, Spotlight, and Polaris when performing HW/SW [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]
Figure 12
Figure 12. Figure 12: We compare the best software mappings produced by Offline, DOSA, Spotlight, and Polaris when performing software [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 13
Figure 13. Figure 13: The behavior Polaris and Spotlight when performing HW/SW co-design. Each segment demarcated by a gray dashed [PITH_FULL_IMAGE:figures/full_fig_p011_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fovea: Physical-Implication-Aware Wafer-Scale DSE with Decision-Domain-Guided Cross-Fidelity Refinement

    cs.AR 2026-08 conditional novelty 6.0 of 10

    A wafer-scale design-space exploration method constructs physically feasible design spaces and prunes candidates by a sampled evaluator-disagreement bound, recovering the exhaustive reference optimum in all 70 tested ...

  2. DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration

    cs.AR 2025-08 conditional novelty 6.0 of 10

    DiffAxE uses conditional diffusion models to generate hardware accelerator designs directly from target performance, achieving orders-of-magnitude faster design space exploration with lower error than existing optimiz...

Reference graph

Works this paper leans on

71 extracted references · 71 canonical work pages · cited by 2 Pith papers

  1. [1]

    Challenges/Opportunities to Enable Dependable Scale-out System with Groq Deterministic Tensor-Streaming Processors,

    D. Abts, I. Ahmed, A. Bitar, M. Boyd, J. Kim, G. Kimmell, and A. Ling, “Challenges/Opportunities to Enable Dependable Scale-out System with Groq Deterministic Tensor-Streaming Processors,” in Dependable Sys- tems and Networks (DSN-S) , Jun. 2022

  2. [2]

    BOOM- Explorer: RISC-V BOOM Microarchitecture Design Space Exploration Framework,

    C. Bai, Q. Sun, J. Zhai, Y . Ma, B. Yu, and M. D. Wong, “BOOM- Explorer: RISC-V BOOM Microarchitecture Design Space Exploration Framework,” in International Conference On Computer-Aided Design (ICCAD), Nov. 2021

  3. [3]

    Transfer Learning for Bayesian Optimization: A Survey,

    T. Bai, Y . Li, Y . Shen, X. Zhang, W. Zhang, and B. Cui, “Transfer Learning for Bayesian Optimization: A Survey,” arXiv, Feb. 2023

  4. [4]

    Autoencoders,

    D. Bank, N. Koenigstein, and R. Giryes, “Autoencoders,” in Data Mining and Knowledge Discovery Handbook , L. Rokach, O. Maimon, and E. Shmueli, Eds., 2023

  5. [5]

    Hyperparameter Optimization: Foundations, Algorithms, Best Practices, and Open Challenges,

    B. Bischl, M. Binder, M. Lang, T. Pielok, J. Richter, S. Coors, J. Thomas, T. Ullmann, M. Becker, A.-L. Boulesteix, D. Deng, and M. Lin- dauer, “Hyperparameter Optimization: Foundations, Algorithms, Best Practices, and Open Challenges,” WIREs Data Mining and Knowledge Discovery, no. 2, 2023

  6. [6]

    Eyeriss: An Energy- Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,

    Y .-H. Chen, T. Krishna, J. S. Emer, and V . Sze, “Eyeriss: An Energy- Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” Solid-State Circuits, no. 1, Jan. 2017

  7. [7]

    Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices,

    Y .-H. Chen, T.-J. Yang, J. Emer, and V . Sze, “Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices,” Emerging and Selected Topics in Circuits and Systems, no. 2, Jun. 2019

  8. [8]

    dMazeRun- ner: Executing Perfectly Nested Loops on Dataflow Accelerators,

    S. Dave, Y . Kim, S. Avancha, K. Lee, and A. Shrivastava, “dMazeRun- ner: Executing Perfectly Nested Loops on Dataflow Accelerators,” Transactions on Embedded Computing Systems , no. 5s, Oct. 2019

Show all 71 references
  1. [9]

    BERT: Pre- training of Deep Bidirectional Transformers for Language Understand- ing,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of Deep Bidirectional Transformers for Language Understand- ing,” arXiv, May 2019

  2. [10]

    Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey,

    P. Dhilleswararao, S. Boppu, M. S. Manikandan, and L. R. Cenkera- maddi, “Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey,” IEEE Access, 2022

  3. [11]

    A Survey on Deep Learning and Its Applications,

    S. Dong, P. Wang, and K. Abbas, “A Survey on Deep Learning and Its Applications,” Computer Science Review , 2021

  4. [12]

    Acceler- ating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,

    M. Emani, V . Vishwanath, C. Adams, M. E. Papka, R. Stevens, L. Florescu, S. Jairath, W. Liu, T. Nama, and A. Sujeeth, “Acceler- ating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,” Computing in Science & Engineering , no. 2, Mar. 2021

  5. [13]

    An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators,

    H. Esmaeilzadeh, S. Ghodrati, A. B. Kahng, J. K. Kim, S. Kinzer, S. Kundu, R. Mahapatra, S. D. Manasi, S. Sapatnekar, Z. Wang, and Z. Zeng, “An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators,” arXiv, Aug. 2023

  6. [14]

    Physically Accurate Learning-Based Performance Prediction of Hardware-Accelerated ML Algorithms,

    H. Esmaeilzadeh, S. Ghodrati, A. B. Kahng, J. K. Kim, S. Kinzer, S. Kundu, R. Mahapatra, S. D. Manasi, S. S. Sapatnekar, Z. Wang, and Z. Zeng, “Physically Accurate Learning-Based Performance Prediction of Hardware-Accelerated ML Algorithms,” in Workshop on Machine Learning for...

  7. [15]

    Improving Performance Estimation for Design Space Exploration for Convolutional Neural Network Accelerators,

    M. Ferianc, H. Fan, D. Manocha, H. Zhou, S. Liu, X. Niu, and W. Luk, “Improving Performance Estimation for Design Space Exploration for Convolutional Neural Network Accelerators,” Electronics, no. 4, Jan. 2021

  8. [16]

    Practical Transfer Learning for Bayesian Optimization,

    M. Feurer, B. Letham, F. Hutter, and E. Bakshy, “Practical Transfer Learning for Bayesian Optimization,” arXiv, Oct. 2022

  9. [17]

    Tests for Rank Correlation Coefficients, I,

    E. C. Fieller, H. O. Hartley, and E. S. Pearson, “Tests for Rank Correlation Coefficients, I,” Biometrika, no. 3-4, Dec. 1957

  10. [18]

    Multi-fidelity Optimiza- tion via Surrogate Modelling,

    A. I. Forrester, A. S ´obester, and A. J. Keane, “Multi-fidelity Optimiza- tion via Surrogate Modelling,” Royal Society A: Mathematical, Physical and Engineering Sciences , no. 2088, Oct. 2007

  11. [19]

    Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack Integration,

    H. Genc, S. Kim, A. Amid, A. Haj-Ali, V . Iyer, P. Prakash, J. Zhao, D. Grubb, H. Liew, H. Mao, A. Ou, C. Schmidt, S. Steffl, J. Wright, I. Stoica, J. Ragan-Kelley, K. Asanovic, B. Nikolic, and Y . S. Shao, “Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation vi...

  12. [20]

    Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules,

    R. G ´omez-Bombarelli, J. N. Wei, D. Duvenaud, J. M. Hern ´andez- Lobato, B. S ´anchez-Lengeling, D. Sheberla, J. Aguilera-Iparraguirre, T. D. Hirzel, R. P. Adams, and A. Aspuru-Guzik, “Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules,” ACS ...

  13. [21]

    Deep Residual Learning for Image Recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Computer Vision and Pattern Recognition (CVPR) , 2016

  14. [22]

    DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators,

    C. Hong, Q. Huang, G. Dinh, M. Subedar, and Y . S. Shao, “DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators,” in Microarchitecture (MICRO), Dec. 2023

  15. [23]

    Learning A Continuous and Reconstructible Latent Space for Hard- ware Accelerator Design,

    Q. Huang, C. Hong, J. Wawrzynek, M. Subedar, and Y . S. Shao, “Learning A Continuous and Reconstructible Latent Space for Hard- ware Accelerator Design,” in International Symposium on Performance Analysis of Systems and Software (ISPASS) , May 2022

  16. [24]

    Ten Lessons From Three Generations Shaped Google’s TPUv4i : Industrial Product,

    N. P. Jouppi, D. Hyun Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian, J. Laudon, S. Li, P. Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Young, Z. Zhou, and D. Patterson, “Ten Lessons From Three Generations Shaped Google’s TPUv4i : Industrial Product,” in Internationa...

  17. [25]

    In-Datacenter Performance Analysis of a Tensor Processing Unit,

    N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V . Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. ...

  18. [26]

    ConfuciuX: Autonomous Hard- ware Resource Assignment for DNN Accelerators using Reinforcement Learning,

    S.-C. Kao, G. Jeong, and T. Krishna, “ConfuciuX: Autonomous Hard- ware Resource Assignment for DNN Accelerators using Reinforcement Learning,” in Microarchitecture (MICRO), Oct. 2020

  19. [27]

    Firesim: FPGA- Accelerated Cycle-Exact Scale-Out System Simulation in the Public Cloud,

    S. Karandikar, H. Mao, D. Kim, D. Biancolin, A. Amid, D. Lee, N. Pemberton, E. Amaro, C. Schmidt, A. Chopra, Q. Huang, K. Kovacs, B. Nikolic, R. Katz, J. Bachrach, and K. Asanovi ´c, “Firesim: FPGA- Accelerated Cycle-Exact Scale-Out System Simulation in the Public Cloud,” in I...

  20. [28]

    A Learned Performance Model for Tensor Processing Units,

    S. Kaufman, P. Phothilimthana, Y . Zhou, C. Mendis, S. Roy, A. Sabne, and M. Burrows, “A Learned Performance Model for Tensor Processing Units,” in Machine Learning and Systems , A. Smola, A. Dimakis, and I. Stoica, Eds., 2021

  21. [29]

    Full Stack Optimization of Transformer Inference: A Survey,

    S. Kim, C. Hooper, T. Wattanawong, M. Kang, R. Yan, H. Genc, G. Dinh, Q. Huang, K. Keutzer, M. W. Mahoney, Y . S. Shao, and A. Gholami, “Full Stack Optimization of Transformer Inference: A Survey,” arXiv, Feb. 2023

  22. [30]

    Auto-Encoding Variational Bayes,

    D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” arXiv, Dec. 2022

  23. [31]

    Spatial: A language and compiler for application accelerators,

    D. Koeplinger, M. Feldman, R. Prabhakar, Y . Zhang, S. Hadjis, R. Fiszel, T. Zhao, L. Nardi, A. Pedram, C. Kozyrakis, and K. Olukotun, “Spatial: A language and compiler for application accelerators,” in Programming Language Design and Implementation (PLDI) , Jun. 2018

  24. [32]

    ArchGym: An Open-Source Gymnasium for Machine Learning Assisted Architecture Design,

    S. Krishnan, A. Yazdanbakhsh, S. Prakash, J. Jabbour, I. Uchendu, S. Ghosh, B. Boroujerdian, D. Richins, D. Tripathy, A. Faust, and V . Janapa Reddi, “ArchGym: An Open-Source Gymnasium for Machine Learning Assisted Architecture Design,” in International Symposium on Computer A...

  25. [33]

    On Information and Sufficiency,

    S. Kullback and R. A. Leibler, “On Information and Sufficiency,” The Annals of Mathematical Statistics , no. 1, 1951

  26. [34]

    Data-Driven Offline Optimization for Architecting Hardware Accelera- tors,

    A. Kumar, A. Yazdanbakhsh, M. Hashemi, K. Swersky, and S. Levine, “Data-Driven Offline Optimization for Architecting Hardware Accelera- tors,” in International Conference on Learning Representations (ICLR) , Oct. 2021

  27. [35]

    MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,

    H. Kwon, P. Chatarasi, V . Sarkar, T. Krishna, M. Pellauer, and A. Parashar, “MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,” IEEE Micro, no. 3, May 2020

  28. [36]

    MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,

    C. Lattner, M. Amini, U. Bondhugula, A. Cohen, A. Davis, J. Pienaar, R. Riddle, T. Shpeisman, N. Vasilache, and O. Zinenko, “MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,” in Code Generation and Optimization (CGO) , Feb. 2021

  29. [37]

    Powering Extreme-Scale HPC with Cerebras Wafer- Scale Accelerators,

    A. Lavely, “Powering Extreme-Scale HPC with Cerebras Wafer- Scale Accelerators,” Cerebras Systems, Inc, Tech. Rep., 2022

  30. [38]

    A Study of Bayesian Neural Network Surrogates for Bayesian Optimization,

    Y . L. Li, T. G. J. Rudner, and A. G. Wilson, “A Study of Bayesian Neural Network Surrogates for Bayesian Optimization,” arXiv, May 2023

  31. [39]

    Focal Loss for Dense Object Detection,

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal Loss for Dense Object Detection,” in International Conference on Computer Vision (ICCV), 2017

  32. [40]

    NAAS: Neural Accelerator Architecture Search,

    Y . Lin, M. Yang, and S. Han, “NAAS: Neural Accelerator Architecture Search,” in Design Automation Conference (DAC) , Dec. 2021

  33. [41]

    The AI Index 2023 Annual Report,

    N. Maslej, L. Fattorini, E. Brynjolfsson, J. Etchemendy, K. Ligett, T. Lyons, J. Manyika, H. Ngo, V . Parli, Y . Shoham, R. Wald, J. Clark, and R. Perrault, “The AI Index 2023 Annual Report,” Institute for Human-Centered AI, Tech. Rep., Apr. 2023

  34. [42]

    Mat ´ern, Spatial Variation, D

    B. Mat ´ern, Spatial Variation, D. Brillinger, S. Fienberg, J. Gani, J. Har- tigan, and K. Krickeberg, Eds., 1986

  35. [43]

    ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,

    L. Mei, P. Houshmand, V . Jain, S. Giraldo, and M. Verhelst, “ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,” Transactions on Computers, no. 8, Aug. 2021

  36. [44]

    STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators,

    F. Mu ˜noz-Mart´ınez, J. L. Abell ´an, M. E. Acacio, and T. Krishna, “STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators,” in International Symposium on Workload Characterization (IISWC), Nov. 2021

  37. [45]

    Practical Design Space Exploration,

    L. Nardi, D. Koeplinger, and K. Olukotun, “Practical Design Space Exploration,” in Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS), Oct. 2019

  38. [46]

    Timeloop: A Systematic Approach to DNN Accelerator Evaluation,

    A. Parashar, P. Raina, Y . S. Shao, Y .-H. Chen, V . A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A Systematic Approach to DNN Accelerator Evaluation,” in International Symposium on Performance Analysis of Systems and Software (ISPASS...

  39. [47]

    Hardware/Software Co-design for Convolutional Neural Networks Acceleration: A Sur- vey and Open Issues,

    C. Pham-Quoc, X.-Q. Nguyen, and T. N. Thinh, “Hardware/Software Co-design for Convolutional Neural Networks Acceleration: A Sur- vey and Open Issues,” in Context-Aware Systems and Applications , P. Cong Vinh and A. Rakib, Eds., 2021

  40. [48]

    A Case for Efficient Accelerator Design Space Exploration via Bayesian Optimization,

    B. Reagen, J. M. Hern ´andez-Lobato, R. Adolf, M. Gelbart, P. What- mough, G.-Y . Wei, and D. Brooks, “A Case for Efficient Accelerator Design Space Exploration via Bayesian Optimization,” in International Symposium on Low Power Electronics and Design (ISLPED) , Jul. 2017

  41. [49]

    Stochastic Backprop- agation and Approximate Inference in Deep Generative Models,

    D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic Backprop- agation and Approximate Inference in Deep Generative Models,” in International Conference on Machine Learning (ICML) , Jun. 2014

  42. [50]

    U-Net: Convolutional Net- works for Biomedical Image Segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Net- works for Biomedical Image Segmentation,” in Medical Image Comput- ing and Computer-Assisted Intervention (MICCAI), N. Navab, J. Horneg- ger, W. M. Wells, and A. F. Frangi, Eds., 2015

  43. [51]

    Learning In- ternal Representations by Error Propagation,

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning In- ternal Representations by Error Propagation,” in Explorations in the Microstructure of Cognition, Jan. 1986

  44. [52]

    Leveraging Domain Information for the Efficient Automated Design of Deep Learning Accelerators,

    C. Sakhuja, Z. Shi, and C. Lin, “Leveraging Domain Information for the Efficient Automated Design of Deep Learning Accelerators,” in High- Performance Computer Architecture (HPCA) , Feb. 2023

  45. [53]

    AIrchitect: Automating Hardware Architecture and Mapping Optimization,

    A. Samajdar, J. M. Joseph, and T. Krishna, “AIrchitect: Automating Hardware Architecture and Mapping Optimization,” in Design, Automa- tion & Test in Europe Conference & Exhibition (DATE) , Apr. 2023

  46. [54]

    A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim,

    A. Samajdar, J. M. Joseph, Y . Zhu, P. Whatmough, M. Mattina, and T. Krishna, “A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim,” in International Symposium on Performance Analysis of Systems and Software (ISPASS), Aug. 2020

  47. [55]

    Neural Architecture Search and Hardware Accelerator Co- Search: A Survey,

    L. Sekanina, “Neural Architecture Search and Hardware Accelerator Co- Search: A Survey,” IEEE Access, 2021

  48. [56]

    An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks,

    K. Seshadri, B. Akin, J. Laudon, R. Narayanaswami, and A. Yazdan- bakhsh, “An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks,” in International Symposium on Workload Character- ization (IISWC), Nov. 2022

  49. [57]

    Simba: Scaling Deep-Learning Inference with Multi-Chip-Module- Based Architecture,

    Y . S. Shao, J. Clemons, R. Venkatesan, B. Zimmer, M. Fojtik, N. Jiang, B. Keller, A. Klinefelter, N. Pinckney, P. Raina, S. G. Tell, Y . Zhang, W. J. Dally, J. Emer, C. T. Gray, B. Khailany, and S. W. Keckler, “Simba: Scaling Deep-Learning Inference with Multi-Chip-Module- Ba...

  50. [58]

    Artificial Intelli- gence in the IoT Era: A Review of Edge AI Hardware and Software,

    T. Sipola, J. Alatalo, T. Kokkonen, and M. Rantonen, “Artificial Intelli- gence in the IoT Era: A Review of Edge AI Hardware and Software,” in Conference of Open Innovations Association (FRUCT) , Apr. 2022

  51. [59]

    On the Distribution of Points in a Cube and the Ap- proximate Evaluation of Integrals,

    I. M. Sobol’, “On the Distribution of Points in a Cube and the Ap- proximate Evaluation of Integrals,” Zhurnal Vychislitel’noi Matematiki i Matematicheskoi Fiziki , no. 4, 1967

  52. [60]

    Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental De- sign,

    N. Srinivas, A. Krause, S. Kakade, and M. Seeger, “Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental De- sign,” in International Conference on Machine Learning (ICML) , Jun. 2010

  53. [61]

    Automated Design of Deep Neural Networks: A Survey and Unified Taxonomy,

    E.-G. Talbi, “Automated Design of Deep Neural Networks: A Survey and Unified Taxonomy,” Computing Surveys, no. 2, Mar. 2021

  54. [62]

    Compute Substrate for Software 2.0,

    J. Vasiljevic, L. Bajic, D. Capalija, S. Sokorac, D. Ignjatovic, L. Bajic, M. Trajkovic, I. Hamer, I. Matosevic, A. Cejkov, U. Aydonat, T. Zhou, S. Z. Gilani, A. Paiva, J. Chu, D. Maksimovic, S. A. Chin, Z. Moudallal, A. Rakhmati, S. Nijjar, A. Bhullar, B. Drazic, C. Lee, J. S...

  55. [63]

    MAGNet: A Modular Accelerator Generator for Neural Networks,

    R. Venkatesan, Y . S. Shao, M. Wang, J. Clemons, S. Dai, M. Fojtik, B. Keller, A. Klinefelter, N. Pinckney, P. Raina, Y . Zhang, B. Zimmer, W. J. Dally, J. Emer, S. W. Keckler, and B. Khailany, “MAGNet: A Modular Accelerator Generator for Neural Networks,” in International Con...

  56. [64]

    Deep Kernel Learning,

    A. G. Wilson, Z. Hu, R. Salakhutdinov, and E. P. Xing, “Deep Kernel Learning,” in Artificial Intelligence and Statistics (AISTATS), May 2016

  57. [65]

    Few-Shot Bayesian Optimization with Deep Kernel Surrogates,

    M. Wistuba and J. Grabocka, “Few-Shot Bayesian Optimization with Deep Kernel Surrogates,” in International Conference on Learning Representations (ICLR), Oct. 2020

  58. [66]

    SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads,

    S. L. Xi, Y . Yao, K. Bhardwaj, P. Whatmough, G.-Y . Wei, and D. Brooks, “SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads,” Transactions on Architecture and Code Optimiza- tion (TACO), no. 4, Nov. 2020

  59. [67]

    HASCO: Towards Agile HArdware and Software CO-design for Tensor Compu- tation,

    Q. Xiao, S. Zheng, B. Wu, P. Xu, X. Qian, and Y . Liang, “HASCO: Towards Agile HArdware and Software CO-design for Tensor Compu- tation,” in International Symposium on Computer Architecture (ISCA) , Jun. 2021

  60. [68]

    Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,

    X. Yang, M. Gao, Q. Liu, J. Setter, J. Pu, A. Nayak, S. Bell, K. Cao, H. Ha, P. Raina, C. Kozyrakis, and M. Horowitz, “Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,” in Ar- chitectural Support for Programming Languages and Operating Systems (ASP...

  61. [69]

    Apollo: Transferable Architecture Exploration,

    A. Yazdanbakhsh, C. Angermueller, B. Akin, Y . Zhou, A. Jones, M. Hashemi, K. Swersky, S. Chatterjee, R. Narayanaswami, and J. Laudon, “Apollo: Transferable Architecture Exploration,” Workshop on ML for Systems , 2020

  62. [70]

    A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators,

    D. Zhang, S. Huda, E. Songhori, K. Prabhu, Q. Le, A. Goldie, and A. Mirhoseini, “A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators,” in Architectural Support for Programming Languages and Operating Systems (ASPLOS) , Feb. 2022

  63. [71]

    A Comprehensive Survey on Transfer Learning,

    F. Zhuang, Z. Qi, K. Duan, D. Xi, Y . Zhu, H. Zhu, H. Xiong, and Q. He, “A Comprehensive Survey on Transfer Learning,” IEEE, no. 1, Jan. 2021

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.