PHOENIX recovers permanent node failures in LLM training via hot-swapping of spares using zero-overhead per-step in-memory optimizer-state replication, finishing recovery in under 40 s on up to 512 GPUs.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis,
4 Pith papers cite this work, alongside 564 external citations. Polarity classification is still indexing.
years
2026 4representative citing papers
MATCHA optimizes DNN deployment on heterogeneous multi-accelerator edge SoCs via constraint programming for memory and scheduling plus pattern matching for parallel execution, cutting latency up to 35% versus the MATCH compiler on MLPerf Tiny.
NEURON-Fabric provides a profile-guided runtime for controlled low-bit gradient communication that preserves accuracy near full-precision levels while reducing modeled communication traffic across vision, transformer, and language model workloads.
MRC extends RoCEv2 with multipath primitives, sender congestion control, decoupled delivery, accelerated loss recovery, and path resilience for AI/ML training over best-effort Ethernet.
citing papers explorer
-
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint
PHOENIX recovers permanent node failures in LLM training via hot-swapping of spares using zero-overhead per-step in-memory optimizer-state replication, finishing recovery in under 40 s on up to 512 GPUs.
-
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
MATCHA optimizes DNN deployment on heterogeneous multi-accelerator edge SoCs via constraint programming for memory and scheduling plus pattern matching for parallel execution, cutting latency up to 35% versus the MATCH compiler on MLPerf Tiny.
-
NEURON-Fabric: Architecture-Runtime Co-Design for Controlled Low-Bit Gradient Communication
NEURON-Fabric provides a profile-guided runtime for controlled low-bit gradient communication that preserves accuracy near full-precision levels while reducing modeled communication traffic across vision, transformer, and language model workloads.
-
The Multipath Reliable Connection (MRC) Transport
MRC extends RoCEv2 with multipath primitives, sender congestion control, decoupled delivery, accelerated loss recovery, and path resilience for AI/ML training over best-effort Ethernet.