Pith. sign in

REVIEW 1 cited by

TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08415 v2 pith:AM5BO454 submitted 2025-03-11 cs.DC

classification cs.DC
keywords systemstokensiminferenceservingexplorationhardwarelanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The increasing demand for large language model (LLM) serving has necessitated significant advancements in the optimization and profiling of LLM inference systems. As these models become integral to a wide range of applications, the need for efficient and scalable serving solutions has grown exponentially. This work introduces TokenSim, a comprehensive hardware and software exploration system designed specifically for LLM inference. TokenSim is characterized by its support for extensible system optimizations including scheduling and memory management. We validate the results with systems running with realworld datasets, achieving an error rate of less than 1%. Furthermore, TokenSim facilitates various insightful explorations into the performance and optimization of LLM serving systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HBF Sucks! A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving

    cs.AR 2026-08 conditional novelty 6.0 of 10

    Replacing an SSD KV-offload tier with High-Bandwidth Flash in an SSD-style LLM serving stack raises average end-to-end latency 2 to 5.5 times and cuts SLO goodput, because transient KV is write-heavy and off the criti...

Pith tools