Link failures cap LEO capacity scalability at O(1/n)
Under a Markov model of ISL states, scalability rises to an optimum then falls toward zero as satellite count grows.
· “Capacity Scalability of LEO Constellations With Dynamic Link Failures”
Distributed, Parallel, and Cluster Computing
Covers fault-tolerance, distributed algorithms, stabilility, parallel computation, and cluster computing. Roughly includes material in ACM Subject Classes C.1.2, C.1.4, C.2.4, D.1.3, D.4.5, D.4.7, E.1.
sort pith recommended most recent
Under a Markov model of ISL states, scalability rises to an optimum then falls toward zero as satellite count grows.
· “Capacity Scalability of LEO Constellations With Dynamic Link Failures”
HyperParallel-MoE converts serial operator runs into a fixed heterogeneous schedule, cutting Dispatch-to-Combine latency up to 1.58x inside
· “HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs”
A stability score and early-exit choices expand the scheduler's options to lower deadline misses and P95 latency on shared GPUs.
· “EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge”
OpenG2G connects real AI service measurements to grid models so users can compare controllers and measure impacts of model choices.
· “OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination”
Flexible per-contract endorsement is added by folding checks into consensus, avoiding extra messages and abort spikes in contested workloads
· “Back to the Future: Rethinking Endorsement in Order-Execute Blockchains”
BCSR and WCSR formats overlap TMA transfers with WGMMA computation to beat cuSPARSE on benchmarks and deliver 2.66x end-to-end LLM prefill.
· “AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures”
T-RBFT runs efficient Raft inside groups and TEE-BFT across them, raising throughput for permissioned chains.
EdgeFlow lowers precision on less critical weights to speed loading from flash while preserving accuracy on phone NPUs.
Separating adapter execution from base models cuts memory costs and lets more requests meet strict response time targets.
· “InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models”
Nesterov-accelerated projection with subspace embeddings preserves accuracy while cutting full-model runtime from over a minute to under 3 s
Hybrid architecture uses MDI star topology to integrate information-theoretic security into distributed consensus.
Projective-geometry adversary families lift the Ω(Ln) baseline to Ω(L n^{1+1/d}) bits for agreement and broadcast.
· “Multivalued Consensus: General Adversaries Require More Communication”
Near-resolution of the tradeoff conjecture shows how larger neighborhoods reduce label sizes in proof labeling schemes for general and minor
· “Near-Resolution of the Tradeoff Conjecture in Distributed Proof Labeling Schemes”
Extending delay-thresholding to LMO momentum handles heterogeneous workers and matches best known bounds in smooth cases.
· “Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method”
Oblivious nodes or polynomial bounds on n let randomness help more problems and create unnatural complexity classes.
· “The Distributed Complexity Landscape on Trees Depends on the Knowledge About the Network Size”
For message-success probability p, no one-round algorithm can beat 2p²q+q³ for p≤2/3, nor q above it.
Grouping peers by reliability and coding within and across clusters cuts latency 10-23% and raises retention up to 30%.
Trained on synthetic workloads, it migrates 7.7x fewer slices than reactive balancing at similar runtimes.
· “Scalable Multi-GPU Simulation of 3D Multicellular Growth with RNN-Based Workload Balancing”
Server and client agents jointly learn aggregation weights and personalization without pretraining.
A neurosymbolic pipeline reads provider documentation, generates state-machine code, and aligns it against the real cloud.
A lightweight model maps sharded and geo-distributed transformer training to deployment plans, cutting cost 21.7% at 99% of peak throughput.
· “ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork”
Demanded weights beat the 34 GiB envelope; all 64 logit rows still match a zero-cache oracle byte for byte.
· “Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency”
A federated graph GAN that scores molecules on user-defined metrics keeps datasets private while targeting drug-likeness.
On chips with fast and slow cores, letting the OS scheduler move threads beats pinning them for uneven workloads.
· “Effects of Hybrid CPU and Cache Architectures on Parallel HPC and Cloud Applications”
Local delays in service-based vehicles spread through shared resources and control loops that timing models overlook.
· “A Survey of Timing Variability in Microservice-Based Software-Defined Vehicles”
Coloring, independent sets, and matchings each need only O(Delta) or O(Delta^2) awake rounds per node.
Even biased Top-k sparsification keeps the centralized O(σ/√(NK)) rate, after a transient N^3/(δ^4Δ^4).
· “CED-EF: Compressed Exact Diffusion with Error Feedback for Multi-Agent Learning”
Nekbone's AX kernel hits 242.97 GFLOPS on Tenstorrent's Wormhole at about 7x less power than the Xeon CPU.
· “Exploring spectral element methods on the Tenstorrent RISC-V accelerator”
Models that exceed resident memory fetch only the newly active neurons per token, holding 92–96% of dense accuracy.
· “NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching”
GCA matches FedAvg accuracy while making server-side data extraction measurably harder.
Consensus on which updates to aggregate removes divergence and blunts poisoning without a central server.
· “Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions”
Barrier waits from matrix-math timing variation grow with cluster size and cap the payoff of faster GPU links.
· “Understanding the Synchronization Tax in GPU Scale-Up Domains”
Epistemic logic sets the bar: bounded registers stay leak-free; unbounded max registers cannot.
A quadtree pruning scheme finds the optimal virtual payment path in milliseconds without sacrificing exactness.
· “Scalable Exact Path Selection via Structure-Aware Search for Virtual Payment Channels”
Prefix and ENI preallocation match overlay speed while keeping VPC-native pod addresses and flow-log visibility.
· “Fleet-Scale Pod Deployment with VPC-Native Networking in Managed Kubernetes”
The cut layers that maximize speed and privacy are exactly where model quality collapses, for every aggregation method.
· “Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMs”
Safety proofs dominate formal verification; liveness, code-level assurance, and scale remain the field's weak points.
· “Systematization of Knowledge: Formal Verification of Consensus Protocols”
The method picks the cheapest model and CPU allocation per request while meeting deadlines and quality targets.
· “PRISM: Predictive Runtime In-place Scaling and Model Selection for Edge Microservices”
Two benchmark numbers separate uniform trust from best-route luck on tested quantum hardware.
Broker-queue design over NVSHMEM also lifts Bellman-Ford to 3.03x max speedup on 10 graphs.
· “A Concurrent Queue System for Multi-GPU Platforms: Application to Bellman-Ford SSSP”
FlashReg keeps recall on par with TurboReg while halving peak tensor memory on embedded GPUs.
FFT convolution on 64 GPUs cuts the source-to-waveform map to two matvecs, making 10,000-scenario ensembles a 4-minute job.
A KKT-driven runtime spends contracted latency slack on memory-bound stages, keeping tails near 1.3x nominal.
· “PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response”
By grouping experts per reasoning stage, SAEM cuts GPU-CPU transfers and boosts throughput 1.33x on average.
Device heat throttles edge AI; Thermo-FL trims workload when hot and filters poisoned sparse updates.
· “Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI”
GrAND updates Vamana and CAGRA in place, beating SVFusion and FreshDiskANN-GPU by 2.2x–8.7x and up to 25x.
· “GrAND: GPU-based Dynamic Graph Indexes for Approximate Nearest Neighbour Search”
Workload-aware choice among PyTorch ops, CUDA libraries, and custom CUDA beats fixed spaces under an 18-candidate budget.
· “HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization”
GT4Py and DaCe generate optimized GPU code from one Python source, with zero copy overhead in the legacy Fortran model.
A 13-defense benchmark finds Byzantine-robust methods fail on heterogeneous data and depend heavily on network topology.
· “BackDFL: A Unified Benchmark For Backdoor Attacks and Defenses In Decentralized Federated Learning”
Why half a year of CubeSat telemetry says orbital schedulers should plan around heat, sunlight, and battery state
For gated-delta hybrids, verifying a draft tree no longer needs a state snapshot per node, and freed HBM buys throughput where memory binds.
· “TreeWY: Speculative Verification for Gated DeltaNet Hybrids”
POT3D scales across multi-GPU nodes using only standard 'do concurrent' loops and unified memory.
Same servers at 1, 10 and 100 Gbps show throughput tracks link rate below 10 Gbps, then memory topology takes over.
A 6.4M-token study shows liquidity and trade features alone carry early warning signal—no contract code needed.
· “Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning”
A committee's private scores are bound into digests any client can replay to check ranking.
· “TrustRAG: Blockchain-Enhanced RAG via Committee-Based Credibility Scoring”
CPU workers reach 11.6x at 16 threads; WebGPU tops 260x on compute-bound workloads.
· “ParaWeb: Parallel Programming Patterns for Web Development”
KV-cache hits jump from 64% to 93% at a 3.5-second tail SLO without changing model execution.
· “CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving”
FleetSieve measures only cells that could shift the allocation, stopping once conservative and optimistic plans agree.
· “FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration”
A controlled test shows the LLM layer gains a small but significant edge only when a mid-run surge creates headroom
Measurements show generation dominates edge-RAG cost and mild compression loses energy, so rates should be set at runtime.
· “From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG”
A 196-second window hides real read failures behind Ready and health checks; cache sizing is the fix.
Sharding a model across idle Intel laptops plus speculative decoding beats the unsplit baseline on the same hardware.
· “Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets”
A proposed architecture splits each safety-message verdict among the car, roadside edge, and cloud, biasing ambiguity toward escalation.
· “Autonomous Cyber Defense in Connected Vehicles: A Multi-Agent Approach to V2X Security”
Shared transformations and cached intermediates reduce duplication, energy, and wall-clock time in a federated trial.
· “Enabling Reuse for Data-Sharing Pipelines in Federated Environments”
A server-only screen of normalization-weight changes beats six baselines, cutting test perplexity by up to 18.6%.
Gating scaling decisions on forecast confidence and bottleneck location trims replica counts by up to a third.
· “SLO-Scaler: Uncertainty-Aware SLO-Driven Autoscaling for Microservices”
For one large environmental workflow, microservices beat the monolith on energy despite running longer.
· “When Do Microservices Save Energy? Evidence from Environmental Simulation Workflows”
Probing the full prefix, installing in fixed windows: local restore staging stays flat as external state grows 32x.
· “Bounded-State Restoration: Decoupling Local Restore Capacity from External LLM State”
In overloaded Wi-Fi tests, uncompressed DMPC received 0.00% of messages; the LSTM code received 98% and converged.
· “Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs”
A strict evidence rule shows the real edge is persistence at close range, not altitude.
A hash-chained telemetry ledger also cuts wire payloads 96.4% and lets agents prove an event never happened.
· “Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations”
FreeToken divides expert cache misses between PCIe and CPU by two measured bandwidths, reaching 1.5–2.3x faster edge decode.
· “FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution”
Single-file scheduler matches static sharding at zero skew and recovers all 2,000 tasks when half the workers die.
Deferring dense recurrent-state updates to periodic merges speeds decode kernels up to 1.86x without changing the model.
· “DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding”
One kernel reuses activation tiles on chip, beating unfused outlier-aware baselines by up to 1.5x.
· “FlashQuant: Sparse-Dense Fusion for Memory-Efficient Outlier-Aware LLM Inference”
Reading it one feed-forward early overlaps both devices — at no measurable training cost.
· “Q-First: Most of Attention Needs Only the Query in Disaggregated LLM Decoding”
Each node drops one of its two fragments on its own once every node has confirmed, and the message stays recoverable.
· “eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning”
Planner-critical objects keep their recall; the rest loses compute, not accuracy.
· “MM-BEV: Enhancing Timeliness by Computing Where and When it Matters”
Hybrid parallelism and rollout-aware memory offloading cut peak GPU memory by up to 52 percent.
P-PAS uses a large prefill budget under light load and a small one under pressure, beating static choices.
· “P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving”