REVIEW 4 major objections 6 minor 64 references
Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
T0 review · 4 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read CHASE resolves the topology-mapping deadlock by searching heterogeneous system designs through their target workloads, yielding 6.20× and 2.12× speedups over fixed baselines.
desk verdict Well-built DSE framework, but the headline speedups are simulator predictions for unbuilt machines, and the scale-out extrapolation is unvalidated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Hierarchical typed hardware graphs (package, node, rack, cluster layers) with a physical-constraint verifier that rejects undeployable candidates; and the decoupled two-level loop: an inner mapper (HEFT/PEFT-style list scheduling with cached cost estimates) producing topology-aware event traces; a calibrated event-driven simulator that models compute, point-to-point/collective communication, queueing, and remote-memory contention on a unified timeline; and an outer GNN-based RL optimizer using bottleneck-telemetry priors, workload-normalized rewards, and legality masks. The organizing identity is the fiber-bundle view — every hardware point h carries its own mapping space S_h, so the search
What would settle it
Build a system of 32–64 GPUs matching one of the discovered topologies' structural template (scale-up islands plus a multi-plane or fat-tree scale-out fabric), run the same HPCG and LLM traces on it, and compare measured makespans and per-event times with CHASE's predictions. If the mean absolute relative error exceeds the reported 5–10% band, or if the predicted ranking between the CHASE design and the fixed baseline flips on real hardware, the scale-out projection model is falsified.
Extended reading notes
Core claim
CHASE's central claim: a hardware candidate cannot be ranked without pairing it with a good mapping — the 'topology-mapping deadlock.' The paper resolves it by evaluating each hardware point at the performance of its best mapping, via a decoupled two-level loop: an inner mapper projects workload DAGs onto topology-aware event traces, a calibrated event-driven simulator scores them, and an outer GNN-based RL optimizer edits the hardware graph under physical-constraint masks. Reported results: 6.06% of exhaustive mapping optima, near-global hardware convergence within 64 iterations, and 6.20×/2.12× geomean speedups for sparse and LLM designs over El Capitan-like and NVL72-like baselines at low
Load-bearing premise
The load-bearing premise is that the scale-out projection model — a ridge regression fit to five systems spanning at most 16 GPUs and validated on a single 8-GPU held-out platform — extrapolates calibration knobs accurately to unbuilt rack- and cluster-scale topologies; if that extrapolation fails, the simulator rankings and the headline speedups lose their ground.
Editorial extensions
If this is right
- Architecture selection for mixed AI/HPC workloads can become an automated, workload-specific design step under physical constraints, rather than a choice among fixed vendor pods.
- Cross-workload comparisons are meaningful only when each hardware candidate is evaluated with its own best-effort mapping; single-platform comparisons may mis-rank designs.
- Sparse-computing systems should concentrate high-bandwidth resources on dependency-critical operations (such as getrf) and offload breadth work to cheaper CPU-GPU hosts.
- LLM inference systems should keep tensor-parallel groups within high-bandwidth scale-up islands and use the cluster fabric primarily for pipeline boundaries, avoiding excessively fine tensor parallelism.
- Simulator rankings calibrated on a few physical platforms can be projected to unbuilt rack/cluster-scale candidates, enabling pre-silicon architectural choices.
Reading between the lines
- The decoupled base-space/fiber structure generalizes beyond XHS: any co-design problem with a nested 'configuration → feasible plans' relationship — memory pooling vs. data placement, optical circuit switching vs. job routing — could use the same two-level organization.
- The reported speedups are geometric means over the chosen workload suites; a different portfolio weighting would likely change the optimal design, so the durable contribution may be the per-portfolio re-optimization capability, not the specific discovered topologies.
- A direct experimental test of the parallelism-granularity thesis ('more GPUs can hurt LLM inference') is possible without building CHASE: run the same model with TP=8/PP=5 vs. TP=64/PP=1 on systems matching those topologies and compare makespans.
- Held-out projection tests at larger scales (32–64 GPUs, multi-plane topologies) would turn the calibration-guardrail claim into a falsifiable engineering guarantee.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CHASE, a framework for architecture exploration of cross-layer heterogeneous systems. Candidate hardware is represented as hierarchical typed graphs, invalid designs are filtered by physical constraints, and a decoupled two-level search is used: an inner mapper and calibrated event-driven simulator evaluate workload mappings for each candidate, while an outer GNN/RL optimizer evolves the hardware graph. The authors claim mapper near-optimality (6.06% gap from exhaustive search on tractable instances), compute-model errors of 4.4–7.5%, communication errors below 10%, held-out projection error of 5.8%, optimizer convergence within 64 iterations, and workload-specific designs achieving 6.20x and 2.12x geomean speedups over El Capitan-like and NVL72-like baselines.
Significance. The paper's core decomposition is appealing and the component-level evidence is credible. The topology-mapping deadlock is made concrete, and the mapper/simulator/optimizer split is a reasonable way to make the joint search tractable. The near-optimality check against exhaustive mapping on small problems and the calibration errors on commercial platforms are concrete, falsifiable measurements. If the scale-out projection can be validated or its uncertainty bounded, the framework would be a worthwhile contribution to system-level DSE. The main unresolved issue is that the headline end-to-end speedups are simulator predictions for unbuilt systems, and the scale-out extrapolation on which those predictions rest has only in-family held-out support.
major comments (4)
- [§6.3, Table 3, §7.3] The scale-out projection in §6.3 is load-bearing. The ridge-regression training pool in Table 3 is five systems at 2–16 GPUs over NVLink, PCIe, and a single 4x200G IB path. The held-out L20 validation is an 8-GPU NVLink platform—same fabric family and within the training scale. The optimizer, however, is allowed to produce candidates with CXL pooling, UALink, optical circuit switches, fat-tree/DragonFly topologies, and 40-GPU islands, and Appendix A.2 states that the numerical ranges in Table 10 are 'experiment inputs' to be finalized. Since the §7.5–7.6 speedups are simulator outputs for unbuilt systems, an unquantified extrapolation error can change the ranking and the headline numbers. Please add (i) a sensitivity analysis of the final architecture and speedups to the projection-regression coefficients, and (ii) at least one held-out validation at larger scale or on an alternative fab
- [§7.5, Table 8] The sparse case-study baseline appears cost- and power-inconsistent. The 'El Capitan-like baseline' is described as 128 H100 SXM GPUs, 32 compute nodes, and 64 CPU ranks, while the CHASE topology is 8 H200 + 4 H100 + 4 L40S GPUs and 5 CPU ranks. Table 8 states 'Both systems are under the same cost and power constrains.' A 128-GPU system is an order of magnitude larger in GPU count and host count; it is difficult to see how both satisfy the same cost and power caps. If the constraints are caps, the cap values must be stated and the baseline shown to satisfy them. If not, the 6.20x number is largely a comparison against a much larger, unconstrained system and does not support the claimed improvement. Please use a cost/power-equivalent baseline or report cost- and power-normalized speedup.
- [§7.6, Table 9] The LLM speedup claim depends on the baseline mapping. The text says 'a use-all-GPU policy on the 64-GPU baseline' selects TP=64/PP=1 for Llama, causing overhead, while the 40-GPU design uses TP=8/PP=5. If the baseline is forced to use all GPUs rather than its best legal mapping, part of the 2.12x speedup is a mapping artifact. Please state the mapping search applied to each baseline and confirm that the NVL72-like baseline was evaluated at its best legal mapping (e.g., TP=8/PP=9 or another configuration that fits the 72-GPU system). If the baseline mapper cannot represent such mappings, that limitation should be stated and the speedup reinterpreted accordingly.
- [§7.4–§7.6] The end-to-end speedups are produced by the same calibrated simulator that the outer loop optimizes; there is no independent end-to-end measurement of any discovered architecture. Component-level calibration is necessary but not sufficient to validate system-level ranking for unbuilt heterogeneous topologies. Please state explicitly which numbers in §7.5–7.6 are simulated predictions, and either provide a sensitivity analysis showing that the final rank order is stable under the reported calibration errors and projection-parameter perturbations, or validate the ranking against a physical prototype or published measurements of a comparable system.
minor comments (6)
- [Tables 8 and 9] Captions: 'constrains' should be 'constraints.'
- [§2.3] Typo: 'presentative' should be 'representative.'
- [Figure 8] The y-axis label 'Optimized Percentage (%)' is unclear; the text describes a gap to exhaustive optimal, so the label should reflect 'makespan relative to exhaustive optimal' or similar.
- [Figure 12] The first x-axis label appears as 'assemb0,' which looks like a leftover token; clean up the figure label.
- [Appendix A.2] The sentence 'Concrete numerical ranges are experiment inputs and will be finalized' conflicts with the claim that physical constraints are enforced with concrete values. Provide the actual values used in the evaluation or state that the published evaluation uses a frozen working set.
- [Table 7] The column header 'Spd. (x)' is undefined; define it in the caption.
Circularity Check
No significant circularity: the calibration is measured, the L20 projection is genuinely held out, and the end-to-end speedups are simulator outputs that do not reduce by construction to the fitted inputs.
full rationale
The derivation chain is: (1) fit per-operator calibration knobs to measured commercial platforms (Section 6.1, Table 3); (2) validate the fitted model on a held-out L20 8-GPU NVLink platform (Section 7.3); (3) use the calibrated mapper/simulator/optimizer loop to search hardware (Section 5); (4) report case-study speedups (Tables 8 and 9). The only projection to unbuilt systems is the ridge-regression scale-out model of Section 6.3. That is an extrapolation, not a circular reduction: it maps hardware attributes to calibration knobs, and the reported speedups are computed by a calibrated event-driven simulator, not read off from any fitted parameter. The case-study speedups are the same geometric-mean objective used in the reward (Eqs. 14-16), but using the optimized objective as the reported result is standard self-consistent DSE evaluation, not a fitted input renamed as a prediction. External anchoring exists at the component level: compute errors average 4.4-7.5%, intra-machine communication errors average below 10%, and the held-out L20 projection has 5.8% mean error. The risk that the scale-out model extrapolates poorly to CXL/UALink/OCS/DragonFly candidates is a validation limitation — the paper itself says numerical ranges in Table 10 are 'experiment inputs' — but that is an unmeasured extrapolation error, not circularity. No load-bearing self-citations, imported uniqueness theorems, or ansatz-smuggling chains appear.
Assumptions & free parameters
free parameters (4)
- Per-operator compute calibration knobs φ_o, β_o, ℓ_o =
Not reported in text; fit per platform via weighted least squares
- Communication calibration parameters (protocol startup latency, link bandwidth derating) =
Not reported; separate parameters for P2P and collective transfers
- Scale-out projection regression coefficients (ridge regression) =
Not reported
- Optimizer objective/reward weights and telemetry-prior weight λ =
w_T, w_C, w_P, w_U, w_Q, w_R not specified; λ=1.0 default (App. B.4)
assumptions (6)
- domain assumption Calibration knobs fitted on three commercial platforms extrapolate to unbuilt XHS candidates via ridge regression.
- domain assumption Workload DAGs (HPCG, Trojan Horse solver traces, LLM inference) preserve performance-relevant dependencies.
- domain assumption Event-driven simulation with roofline compute bounds and serialization/queueing network models is sufficient for ranking candidate architectures.
- domain assumption Physical constraint predicates (power, thermal, radix, space, budget) are correct and complete for deployability.
- domain assumption Baseline systems (El Capitan-like, NVL72-like) are representative and subject to the same cost/power constraints as optimized designs.
- standard math Roofline upper bound max(flops/peak, bytes/BW) is a valid baseline for compute event time.
Cite this review
Pith. "Pith review of Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems." pith.science (2026). https://pith.science/paper/FAF5C4ZK
@misc{pith2026260723042,
author = {Pith},
title = {Pith review of: Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/FAF5C4ZK}},
note = {Machine review of arXiv:2607.23042}
}
abstract
AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives. Supporting these portfolios requires coordinated choices across accelerators, memory tiers, scale-up fabrics, and cluster networks. The resulting Cross-layer Heterogeneous System (XHS) design space is difficult to explore: hardware choices change legal task mappings, while rack power, switch radix, cabling, and cost constraints invalidate many candidates. We present CHASE, an application-driven framework that searches physically feasible XHS architectures through the workloads they must execute. CHASE represents candidates as hierarchical typed graphs and rejects designs that violate deployment constraints. It avoids intractable joint hardware-mapping search with a decoupled two-level loop: an inner mapper translates hardware-independent workload DAGs into topology-aware event traces, a calibrated event-driven simulator evaluates each mapping, and an outer telemetry-guided optimizer evolves the hardware graph. We evaluate CHASE on sparse-computing and LLM workloads. Its mapper remains within 6.06% of exhaustive optima while reducing mapping time by 60.5% on average relative to PEFT. Compute-model errors average 4.4-7.5%, and communication validation reproduces key trends across physical platforms. The outer search reaches near-global optima within 64 iterations. End-to-end case studies show that sparse workloads favor criticality-aware heterogeneous pods, whereas LLM inference favors scale-up islands; the resulting designs deliver 6.20$\times$ and 2.12$\times$ geomean speedups, respectively, while reducing cost and power relative to the baselines.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Sergi Abadal, Akshay Jain, Robert Guirado, Jorge López-Alonso, and Eduard Alarcón. 2021. Computing graph neural networks: A survey from algorithms to accelerators.ACM Computing Surveys (CSUR)54, 9 (2021), 1–38
2021
-
[2]
AMD. 2025. AMD Instinct MI350 Series GPUs. Product documenta- tion.https://www.amd.com/en/products/accelerators/instinct/mi350. htmlAccessed 2026-06-01
2025
-
[3]
Hamid Arabnejad and Jorge G. Barbosa. 2014. List Scheduling Algo- rithm for Heterogeneous Systems by an Optimistic Cost Table.IEEE Transactions on Parallel and Distributed Systems25, 3 (2014), 682–694. doi:10.1109/TPDS.2013.57
-
[4]
Grey Ballard, Erin Carson, James Demmel, Mark Hoemmen, Nicholas Knight, and Oded Schwartz. 2014. Communication Lower Bounds and Optimal Algorithms for Numerical Linear Algebra.Acta Numerica23 (2014), 1–155. doi:10.1017/S0962492914000038
-
[5]
Jehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim, and Minsoo Rhu. 2024. vtrain: A simulation framework for evaluating cost-effective and compute-optimal large language model training. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 153–167
2024
-
[6]
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov...
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[8]
Henri Casanova, Arnaud Giersch, Arnaud Legrand, Martin Quinson, and Frédéric Suter. 2014. Versatile, Scalable, and Accurate Simulation of Distributed Applications and Platforms.J. Parallel and Distrib. Comput.74, 10 (2014), 2899–2917. doi:10.1016/j.jpdc.2014.06.008
Show all 64 references
-
[9]
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th US...
2018
-
[10]
Compute Express Link Consor- tium.https://computeexpresslink.org/wp-content/uploads/2024/02/ CXL-3.1-Specification.pdf
Compute Express Link Consortium 2023.Compute Express Link Specification Revision 3.1. Compute Express Link Consor- tium.https://computeexpresslink.org/wp-content/uploads/2024/02/ CXL-3.1-Specification.pdf
2023
-
[11]
Jeffrey Dean and Luiz André Barroso. 2013. The Tail at Scale.Commun. ACM56, 2 (2013), 74–80. doi:10.1145/2408776.2408794
2013
-
[12]
Jack Dongarra, Michael A Heroux, and Piotr Luszczek. 2016. High- performance conjugate-gradient benchmark: A new metric for rank- ing high-performance computing systems.The International Jour- nal of High Performance Computing Applications30, 1 (2016), 3–
2016
-
[13]
arXiv:https://doi.org/10.1177/1094342015593158 doi:10.1177/ 1094342015593158
-
[14]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravanku- mar, Artem Korenev, A...
2024 arXiv
-
[15]
Charles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar, and Yakun Sophia Shao. 2023. Dosa: Differentiable model-based one- loop search for dnn accelerators. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture. 209–224
2023
-
[16]
Eliu A Huerta, Asad Khan, Edward Davis, Colleen Bushell, William D Gropp, Daniel S Katz, Volodymyr Kindratenko, Seid Koric, William TC Kramer, Brendan McGinty, et al. 2020. Convergence of artificial intel- ligence and high performance computing on NSF-supported cyberin- frastr...
2020 arXiv
-
[17]
Mikhail Isaev, Nic Mcdonald, Larry Dennison, and Richard Vuduc
-
[18]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Men- sch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-An...
2024 arXiv
-
[19]
Qi, and Alex Aiken
Zhihao Jia, Sina Lin, Charles R. Qi, and Alex Aiken. 2019. Beyond Data and Model Parallelism for Deep Neural Networks. InProceedings of Machine Learning and Systems. MLSys, Stanford, CA, 1–13
2019
-
[20]
Sheng-Chun Kao, Geonhwa Jeong, and Tushar Krishna. 2020. Confu- ciux: Autonomous hardware resource assignment for dnn accelerators using reinforcement learning. In2020 53rd Annual IEEE/ACM Interna- tional Symposium on Microarchitecture (MICRO). IEEE, 622–636
2020
-
[21]
Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Na- garajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andrew Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David A. Patterson. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machin...
2023
-
[22]
Katz, Jonathan Bachrach, and Krste Asanović
Sagar Karandikar, Howard Mao, Donggyu Kim, David Biancolin, Alon Amid, Dayeol Lee, Nathan Pemberton, Emmanuel Amaro, Colin Schmidt, Aditya Chopra, Qijing Huang, Kyle Kovacs, Borivoje Nikolic, Randy H. Katz, Jonathan Bachrach, and Krste Asanović. 2018. FireSim: FPGA-Accelerated...
2018
-
[23]
Sheng-Chun Kao and Tushar Krishna. 2020. Gamma: Automating the hw mapping of dnn models on accelerators via genetic algorithm. In Proceedings of the 39th International Conference on Computer-Aided Design. 1–9
2020
-
[24]
Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. 2019. Under- standing Reuse, Performance, and Hardware Cost of DNN Dataflows: A Data-Centric Approach. InProceedings of the 52nd Annual IEEE/ACM 15 International Symposium on Microarchitecture. IEEE, Columbus, OH, 754–768. do...
2019
-
[25]
Dally, Steve Scott, and Dennis Abts
John Kim, William J. Dally, Steve Scott, and Dennis Abts. 2008. Technology-Driven, Highly-Scalable Dragonfly Topology. InProceed- ings of the 35th Annual International Symposium on Computer Archi- tecture. IEEE, Beijing, China, 77–88. doi:10.1109/ISCA.2008.19
2008 doi
-
[26]
Leiserson
Charles E. Leiserson. 1985. Fat-Trees: Universal Networks for Hardware-Efficient Supercomputing.IEEE Trans. Comput.C-34, 10 (1985), 892–901. doi:10.1109/TC.1985.6312192
1985
-
[27]
2025.El Capitan System Readiness: L2 Milestone Summary
Matthew LeGendre and Adam Bertsch. 2025.El Capitan System Readiness: L2 Milestone Summary. Technical Report. Lawrence Liv- ermore National Laboratory (LLNL), Livermore, CA (United States). doi:10.2172/2584749
2025 doi
-
[28]
Yida Li, Siwei Zhang, Yiduo Niu, Yang Du, Qingxiao Sun, Zhou Jin, and Weifeng Liu. 2026. Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters. InProceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Progra...
2026
-
[29]
Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D
Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bian- chini. 2023. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. InPro...
2023
-
[30]
Thomas McSweeney, Neil Walton, and Mawussi Zounon. 2020. An Ef- ficient New Static Scheduling Heuristic for Accelerated Architectures. InComputational Science – ICCS 2020, Valeria V. Krzhizhanovskaya, Gábor Závodszky, Michael H. Lees, Jack J. Dongarra, Peter M. A. Sloot, Sérgi...
2020
-
[31]
Berger, Marie Nguyen, Xun Jian, Sam H
Jinshu Liu, Hamid Hadian, Yuyue Wang, Daniel S. Berger, Marie Nguyen, Xun Jian, Sam H. Noh, and Huaicheng Li. 2025. Systematic CXL Memory Characterization and Performance Analysis at Scale. In Proceedings of the 30th ACM International Conference on Architectural Support for Pr...
2025
-
[32]
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGres- ley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Efficient Large-Scale Language Model Training on GPU Cl...
2021
-
[33]
Diksha Moolchandani, Joyjit Kundu, Frederik Ruelens, Peter Vrancx, Timon Evenblij, and Manu Perumkunnil. 2023. Amped: An analytical model for performance in distributed training of transformers. In2023 IEEE International Symposium on Performance Analysis of Systems and Softwar...
2023
-
[34]
NVIDIA. 2025. CES 2025: AI Advancing at ’Incredible Pace, ’ NVIDIA CEO Says. NVIDIA Blog.https://blogs.nvidia.com/blog/ces-2025- jensen-huang/Accessed 2026-06-08
2025
-
[35]
NVIDIA. 2023. NVIDIA DGX SuperPOD: Next Generation Scalable Infrastructure for AI Leadership Reference Architec- ture Featuring NVIDIA DGX H100. Reference Architecture. https://docs.nvidia.com/dgx-superpod/reference-architecture- scalable-infrastructure-h100/latest/Accessed 2026-06-01
2023
-
[36]
NVIDIA. 2026. GB200 NVL72. Product documentation.https://www. nvidia.com/en-us/data-center/gb200-nvl72/Accessed 2026-06-01
2026
-
[37]
NVIDIA. 2026. AI Factories. NVIDIA Data Center Solutions.https: //www.nvidia.com/en-us/solutions/ai-factories/Accessed 2026-06-10
2026
-
[38]
Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5
2026
-
[39]
Shao, Yu-Hsin Chen, Victor A
Angshuman Parashar, Priyanka Raina, Yakun S. Shao, Yu-Hsin Chen, Victor A. Ying, Anurag Mukkara, Rangharajan Venkatesan, Brucek Khailany, Stephen W. Keckler, and Joel Emer. 2019. Timeloop: A Systematic Approach to DNN Accelerator Evaluation. In2019 IEEE International Symposium...
2019
-
[40]
Daniel Reed, Dennis Gannon, and Jack Dongarra. 2023. HPC forecast: Cloudy and uncertain.Commun. ACM66, 2 (2023), 82–90
2023
-
[41]
Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna. 2020. ASTRA-SIM: Enabling SW/HW Co-Design Exploration for Distributed DL Training Platforms. In2020 IEEE International Symposium on Performance Analysis of Systems and Software. IEEE, Boston, MA, 81–9...
2020
-
[42]
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos. 2022. Compute trends across three eras of machine learning. In2022 international joint conference on neural networks (IJCNN). IEEE, 1–8
2022
- [43]
-
[44]
1999.The topology of fibre bundles
Norman Earl Steenrod. 1999.The topology of fibre bundles. Vol. 14. Princeton university press
1999
- [45]
-
[46]
Haluk Topcuoglu, Salim Hariri, and Min-You Wu. 2002. Performance- Effective and Low-Complexity Task Scheduling for Heterogeneous Computing.IEEE Transactions on Parallel and Distributed Systems13, 3 (2002), 260–274. doi:10.1109/71.993206 16
2002 doi
-
[47]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex...
2024 arXiv
-
[48]
UnifiedBus. 2026. UnifiedBus SuperPoD Reference Architecture White Paper. Technical white paper.https://www.unifiedbus.com/enAc- cessed 2026-06-01
2026
-
[49]
Kevin Tran and Zachary W. Ulissi. 2018. Active learning across inter- metallics to guide discovery of electrocatalysts for CO2 reduction and H2 evolution.Nature Catalysis1 (2018), 696–703. doi:10.1038/s41929- 018-0142-1
2018 doi
-
[51]
UnifiedBus Open Ecosys- tem.https://www.openeuler.org/projects/ub-service-core/white- paper/UB-Service-Core-SW-Arch-RD-2.0-en.pdfVersion 2.0
UnifiedBus Open Ecosystem 2025.UnifiedBus™(UB) Service Core Software Architecture Reference Design. UnifiedBus Open Ecosys- tem.https://www.openeuler.org/projects/ub-service-core/white- paper/UB-Service-Core-SW-Arch-RD-2.0-en.pdfVersion 2.0
2025
-
[52]
Min Wang, Haoyuan Wang, Sibo Qiao, Jiawang Chen, Qin Xie, and Cuijuan Guo. 2025. Heterogeneous system list scheduling algorithm based on improved optimistic cost matrix.Future Generation Computer Systems164 (2025), 107576. doi:10.1016/j.future.2024.107576
2025
-
[53]
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, An- dreea Deac, et al. 2023. Scientific discovery in the age of artificial intelli- gence.Nature620, 7972 (2023), 47–60. doi:10.1038/s41586-023-06221-2
2023 doi
-
[54]
Wilke and Joseph P
Jeremiah J. Wilke and Joseph P. Kenny. 2015.Using Discrete Event Simu- lation for Programming Model Exploration at Extreme-Scale: Macroscale Components for the Structural Simulation Toolkit (SST). Technical Report SAND2015-1027. Sandia National Laboratories. doi:10.2172/ 1170619
2015
-
[55]
2023.{TopoOpt}: Co-optimizing network topology and parallelization strategy for distributed training jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Kewitsch. 2023.{TopoOpt}: Co-optimizing network topology and parallelization strategy for distributed training jobs. In20th USENIX Symposium on Networked Systems...
2023
-
[56]
Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: An Insightful Visual Performance Model for Multicore Ar- chitectures.Commun. ACM52, 4 (2009), 65–76. doi:10.1145/1498765. 1498785
2009 doi
-
[57]
Samuel Williams, Leonid Oliker, Richard Vuduc, John Shalf, Katherine Yelick, and James Demmel. 2009. Optimization of Sparse Matrix-Vector Multiplication on Emerging Multicore Platforms.Parallel Comput.35, 3 (2009), 178–194. doi:10.1016/j.parco.2008.12.006
2009 doi
-
[58]
William Won, Saeed Rashidi, Sudarshan Srinivasan, and Tushar Kr- ishna. 2024. LIBRA: Enabling workload-aware multi-dimensional network topology optimization for distributed training of large AI models. In2024 IEEE International Symposium on Performance Analysis of Systems and ...
2024
-
[59]
William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudar- shan Srinivasan, and Tushar Krishna. 2023. ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-Model Training at Scale. In2023 IEEE International Symposium on Perfor- mance Analys...
2023
-
[60]
Yannan Nellie Wu, Po-An Tsai, Angshuman Parashar, Vivienne Sze, and Joel S Emer. 2022. Sparseloop: An analytical approach to sparse tensor accelerator modeling. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1377–1395
2022
-
[61]
Emer, and Vivienne Sze
Yannan Nellie Wu, Joel S. Emer, and Vivienne Sze. 2019. Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator Designs. InProceedings of the 2019 IEEE/ACM International Conference on Computer-Aided Design. IEEE, Westminster, CO, 1–8. doi:10.1109/ ICCAD4...
2019
-
[62]
Xing, Joseph E
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Automat- ing Inter- and Intra-Operator Parallelism for Distributed Deep Learn- ing. InProceed...
2022
-
[63]
Gonzalez, and Ion Stoica
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica. 2020. Ansor: Generating High-Performance Tensor Programs for Deep Learning. InProceed- ings of the 14th USENIX Symp...
2020
-
[65]
Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun. 2020. Graph Neural Networks: A Review of Methods and Applications.AI Open1 (2020), 57–81. doi:10.1016/j.aiopen.2021.01.001 A Detailed Hardware Description S...
2020 doi
-
[2023]
InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Calculon: a methodology and tool for high-level co-design of systems and large language models. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1–14
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.