REVIEW 5 cited by
Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Benchmarking and co-design are essential for driving optimizations and innovation around ML models, ML software, and next-generation hardware. Full workload benchmarks, e.g. MLPerf, play an essential role in enabling fair comparison across different software and hardware stacks especially once systems are fully designed and deployed. However, the pace of AI innovation demands a more agile methodology to benchmark creation and usage by simulators and emulators for future system co-design. We propose Chakra, an open graph schema for standardizing workload specification capturing key operations and dependencies, also known as Execution Trace (ET). In addition, we propose a complementary set of tools/capabilities to enable collection, generation, and adoption of Chakra ETs by a wide range of simulators, emulators, and benchmarks. For instance, we use generative AI models to learn latent statistical properties across thousands of Chakra ETs and use these models to synthesize Chakra ETs. These synthetic ETs can obfuscate key proprietary information and also target future what-if scenarios. As an example, we demonstrate an end-to-end proof-of-concept that converts PyTorch ETs to Chakra ETs and uses this to drive an open-source training system simulator (ASTRA-sim). Our end-goal is to build a vibrant industry-wide ecosystem of agile benchmarks and tools to drive future AI system co-design.
Forward citations
Cited by 5 Pith papers
-
MoX: Efficient MoE Routing on Direct-Connect Topologies
Static, demand-oblivious routing with token-aware multicast trees and precomputed per-link weights brings MoE traffic on direct-connect fabrics close to ideal switch performance.
-
Opus: Photonic Rail-Optimized Fabric in ML Datacenters
Opus time-multiplexes a single photonic rail fabric across parallelism phases in ML training, achieving up to 23x network power reduction and 4x cost savings at under 6.7% training overhead in simulation.
-
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
STAGE synthesizes high-fidelity Chakra-format execution graphs for distributed LLM workloads from symbolic tensor definitions, validated against real 128-GPU H100 traces and scaled to 32K GPUs.
-
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
Mycroft adds collective-communication-level tracing to NCCL so that slow or stuck data transfers in LLM training can be detected and traced to likely faulty ranks in seconds.
-
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.
Discussion (0). Continue with ORCID to comment.