Pith. sign in

cs.PF

Performance

Covers performance measurement and evaluation, queueing, and simulation. Roughly includes material in ACM Subject Classes D.4.8 and K.6.2.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Node-local NVMe staging removes two shared-fabric penalties in DL training

DYAD's lazy byte-range caching cuts DataLoader stalls from 28% to 2.9% of iterations and reduces all-reduce contention 145×.

· “Sharing a Fabric with Collective Communication: Two Storage Penalties in Deep Learning Training”

open re-runnable review →
Figure from the paper

RAGMark is a modular benchmarking framework that measures per-stage latency

RAGMark enables fine-grained, reproducible benchmarking of RAG pipelines across retrievers, vector databases, reranking, compression, and…

· “RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems”

open re-runnable review →
Figure from the paper

RASER is a user-space framework that runs LLM-based agentic workflows on existing HPC…

RASER uses Slurm job arrays with shared-filesystem work stealing to dynamically balance agent workloads, achieving near-39% faster makespan…

· “RASER: Resilient Agent Scheduling and Execution Runtime for HPC Clusters”

open re-runnable review →
Figure from the paper

Scale analysis shows 19% SMT throughput loss in SPEC CPU 2026 on Zen 5

First microarchitectural characterization of the new suite reveals SMT contention, L3 interference, and three workload clusters invisible…

· “Performance Characterization of SPEC CPU 2026 on AMD EPYC 9755 Processor”

open re-runnable review →
Figure from the paper

LLM inference energy per token falls as request energy rises

A fixed-plus-step energy model explains why longer outputs and larger batches hide growing total GPU energy while lowering per-token cost.

· “Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms”

open re-runnable review →
Figure from the paper

Launch-bound and substitutable: why three MoE optimizations fail

Fused kernels, INT4 quantization, and graph compilation underperform because MoE is launch-bound and experts are interchangeable.

· “Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models”

open re-runnable review →
Figure from the paper

O-RAN's CU and DU show opposite processor responses to traffic

Profiling under matched hardware reveals the CU shifts strongly under load while the DU stays dominated by IQ-sample processing, guiding…

· “PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed Units”

open re-runnable review →
Figure from the paper

Fluxion speeds long-context inference 1.5x-3.7x via CPU-GPU hybrid sparse attention

Dynamic output-aware budgeting and priority scheduling keep quality within 0.26 of full attention for CPU-resident KV caches

· “An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference”

open re-runnable review →
Figure from the paper

Network bandwidth rules Earth-observation storage up to 10 Gbps

Same servers at 1, 10 and 100 Gbps show throughput tracks link rate below 10 Gbps, then memory topology takes over.

· “Design and Empirical Evaluation of a Network-Centric, On-Premises Architecture for Earth Observation Data Access”

open re-runnable review →
Figure from the paper

Falcon-512 fits C-V2X, fails 90% delivery above lightest traffic

The only NIST post-quantum signature that fits today's sidelink spec delivers 90% only at lightest traffic, LOS.

· “Deployment Feasibility Analysis of Post-Quantum Digital Signatures in Safety-Critical C-V2X Communication for Urban Mobility Scenario”

open re-runnable review →

browse all of cs.PF → full archive · search · sub-categories