Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T08:20:01.011625Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 100 inbound Pith citation observations for arXiv:2312.07104.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T08:20:01.011625Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:54.657298Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
62 of 62 outbound references displayed
External citation measurements
10
pith, observed 2026-08-05T02:28:24.338817Z
Observation 6a74ddd4-1c8b-4d04-85ea-4a396ab5641c · outbound
SGLang: Efficient Execution of Structured Language Model Programs InferCept: Efficient Intercept Support for Augmented Large Language Model Inference
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b347b4f8-7709-4be5-9d16-9102b1c22db3 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfa5f099-8e66-4b48-8788-f110d5523b09 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Deepspeed- inference: enabling efficient inference of transformer models at unprecedented scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 222e3b06-ed59-4657-9946-75ac28b2a895 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Prompting is programming: A query language for large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6225a0e-e2ac-4848-b0ac-42a84a4bb767 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Language models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a2bf951c-e3c7-438e-a045-66f735e760f9 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c510580-7f54-49ab-8e5e-829538ee0a74 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Gonzalez, Ion Stoica, and Eric P
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6666ca9-1e4f-4bdb-b702-6f659369aceb · outbound
SGLang: Efficient Execution of Structured Language Model Programs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b114185-21e5-47b0-af92-dfe57e4809d8 · outbound
SGLang: Efficient Execution of Structured Language Model Programs PaLM: Scaling Language Modeling with Pathways
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a9d5276-b026-4565-9745-586e13278345 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8161cbd0-a3ce-4ff1-bcb4-8f19c6ee6dc4 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Model tells you what to discard: Adaptive kv cache compression for llms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd159a32-1bc2-4141-8132-23950eee6907 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3156d99-88b5-4a7f-8d5e-220dd2014b0f · outbound
SGLang: Efficient Execution of Structured Language Model Programs A guidance language for controlling large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 02314b4d-771e-4e98-95cd-fd7a80d63454 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Measuring massive multitask language understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0ce2c05-fd48-4cd7-8700-7480987707f1 · outbound
SGLang: Efficient Execution of Structured Language Model Programs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22dbdc37-18d0-4486-a3d7-f15d598bbfd0 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Text generation inference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e88c493e-4ea4-44f2-a79e-04395ba29cc6 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Mixtral of Experts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8dc0df40-83bd-4002-b1a2-e7021f3f8de3 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Hydragen: High-Throughput LLM Inference with Shared Prefixes
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c274a29-48b0-4191-a303-036c206199ef · outbound
SGLang: Efficient Execution of Structured Language Model Programs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 102c2646-90ea-498c-95fb-b32e4e104da7 · outbound
SGLang: Efficient Execution of Structured Language Model Programs DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa6b9cdd-d472-4b54-830b-b2b035656591 · outbound
SGLang: Efficient Execution of Structured Language Model Programs An LLM Compiler for Parallel Function Calling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 467bc4d9-5eeb-4dd8-ab7d-7288bc48faba · outbound
SGLang: Efficient Execution of Structured Language Model Programs Validating large language models with relm
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd4e539b-a689-46ef-8cbb-807e5a8d3822 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Efficient memory management for large language model serving with pagedattention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6ae0b63-a0e2-4f5e-b038-e68abab52220 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Langchain
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93dea5ef-e549-421f-812b-9e58b5adca8d · outbound
SGLang: Efficient Execution of Structured Language Model Programs Competition-level code generation with alphacode
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11f75dab-de9a-4207-a686-38d025a71289 · outbound
SGLang: Efficient Execution of Structured Language Model Programs AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bef4281-47d3-4d0b-a34a-3135424c8f37 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Improved Baselines with Visual Instruction Tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97d09804-584a-42d1-94db-9eaecdbba861 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2fe388f-f18b-4bd7-9039-26cb1fcc4d43 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Optimizing LLM Queries in Relational Data Analytics Workloads
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 119edb1b-f6fc-4628-8d29-d3ec447e9e46 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Prompting Frameworks for Large Language Models: A Survey
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a53583b1-1e75-4360-80df-7418980dded1 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Scissorhands: Exploiting the persistence of impor- tance hypothesis for llm kv cache compression at test time
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0735568f-dbb2-4efc-ad5b-90524c111cb0 · outbound
SGLang: Efficient Execution of Structured Language Model Programs KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ed3cef9-2a5a-4609-8810-538bcadcb404 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Skeleton- of-thought: Prompting LLMs for efficient parallel generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9e6dfef-c1e3-4282-833e-29fc99778ec7 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Tensorrt-llm
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe9e2e16-12c9-4755-886d-040f99da480a · outbound
SGLang: Efficient Execution of Structured Language Model Programs Gpt-4 technical report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5ed99bc-5455-44d8-95c2-bc8ba52ea017 · outbound
SGLang: Efficient Execution of Structured Language Model Programs O’Brien, Carrie J
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7b9b901-6dde-44a2-b694-c3e5adef8cf2 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Pytorch: An imperative style, high-performance deep learning library
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4c379df-e753-4d3e-a163-ec6de16abe71 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Gorilla: Large Language Model Connected with Massive APIs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f5f4433-affe-49a3-a4e5-34b4640bff3c · outbound
SGLang: Efficient Execution of Structured Language Model Programs Efficiently scaling transformer inference
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee83fe8c-3d29-44d9-8141-754db48372c6 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29cfe3c7-7df4-46dc-aa8e-6f7cbf7fd161 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d50df4e9-57fb-45d3-a95e-ca3f44752149 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Fairness in Serving Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db7e1274-624e-4de8-a56f-f16fe85fe872 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Flexgen: high-throughput generative inference of large language models with a single gpu
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9d0f233-c377-4ed9-8dfa-7df0dfe79328 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40c2f736-b5ee-4e78-8293-49a484249d2e · outbound
SGLang: Efficient Execution of Structured Language Model Programs Preble: Efficient distributed prompt scheduling for llm serving
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5bddc7cd-152d-4182-a952-ad1bd8e58035 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Cognitive architec- tures for language agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3fd41751-957f-4c58-8d78-565e7ead93f6 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Gemini: A Family of Highly Capable Multimodal Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc261552-7ec8-47db-b99d-d54239b727d0 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Triton: an intermediate language and compiler for tiled neural network computations
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81c43e3d-5c12-4c9d-bcb2-9a7a86abf7da · outbound
SGLang: Efficient Execution of Structured Language Model Programs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 635e28f4-f792-4512-9142-8d833297cc99 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Fast, high-fidelity llm decoding with regex constraints
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a590169-463e-4d0d-b548-b48f07c1d48c · outbound
SGLang: Efficient Execution of Structured Language Model Programs Attention is all you need
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa982f5a-1680-4661-b7d7-86df00d044bd · outbound
SGLang: Efficient Execution of Structured Language Model Programs Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d131e241-e0a1-45ad-bd2e-3af1639e5821 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Self-consistency improves chain of thought reasoning in language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edf0df46-2270-4f2e-aaf1-a2aaa2880b3b · outbound
SGLang: Efficient Execution of Structured Language Model Programs Efficient guided generation for large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6b32f6d-5a96-421c-ba41-0df032b1ed7f · outbound
SGLang: Efficient Execution of Structured Language Model Programs AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bc1570a-30a3-4a00-91ed-eea3e61fb91e · outbound
SGLang: Efficient Execution of Structured Language Model Programs Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1144386f-bf4b-473f-b9dc-d2c6938a4606 · outbound
SGLang: Efficient Execution of Structured Language Model Programs React: Synergizing reasoning and acting in language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15a4b487-9e26-478c-9c62-a13b250ca0c4 · outbound
SGLang: Efficient Execution of Structured Language Model Programs ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f022368-f493-4cc7-863c-594ca2e0963c · outbound
SGLang: Efficient Execution of Structured Language Model Programs Accelerating self-attentions for llm serving with flashinfer, February 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13e324ef-f3f6-4ab6-b7ae-664bf3524902 · outbound
SGLang: Efficient Execution of Structured Language Model Programs Orca: A distributed serving system for {Transformer-Based} generative models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb642d8f-19a2-4f52-9d32-21eb415a57ea · outbound
SGLang: Efficient Execution of Structured Language Model Programs Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc89d572-cdb3-4a1c-8ad9-2c4b34fb3883 · outbound
SGLang: Efficient Execution of Structured Language Model Programs prefill"). It then sequentially decodes output tokens, with each token depending on prior tokens (this process is called
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d65f1508-ed36-4284-a44a-59fa6e2af223 · inbound
A Survey on Efficient Inference for Large Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c34b8ce4-76bf-4cc0-b6fe-4eca1d791962 · inbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2611a68e-8f2e-4ccf-8031-76d4f7472c8f · inbound
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling SGLang: Efficient Execution of Structured Language Model Programs
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69661b31-10ff-46b6-bc7c-c34df8a156de · inbound
A Survey on LLM-as-a-Judge SGLang: Efficient Execution of Structured Language Model Programs
Reference 222
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b756769-b3f8-4ab6-aac6-c6f91020a592 · inbound
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fad1018-e2de-4510-be2b-8afa984b15a2 · inbound
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec42c426-1a3b-4ee4-a42c-26f989f42211 · inbound
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production SGLang: Efficient Execution of Structured Language Model Programs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ebf05aa-7325-4941-b15a-c360b46e99c5 · inbound
Hermes 4 Technical Report SGLang: Efficient Execution of Structured Language Model Programs
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403bab14-b320-418e-8c40-0c84715ee33a · inbound
Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics SGLang: Efficient Execution of Structured Language Model Programs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11e8fe6-25f0-43d0-9d44-4fe37e879ddd · inbound
Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution SGLang: Efficient Execution of Structured Language Model Programs
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496c1c4f-4c06-4eaa-8077-690f696a7307 · inbound
CacheClip: Accelerating RAG with Effective KV Cache Reuse SGLang: Efficient Execution of Structured Language Model Programs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d2e4ea5-9283-43a2-a1a0-ad804922f90c · inbound
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a713be5b-4cec-48fd-b760-e990a3df4b86 · inbound
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators SGLang: Efficient Execution of Structured Language Model Programs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9810cff1-4b96-47b4-9d9d-6f7b22f0db94 · inbound
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework SGLang: Efficient Execution of Structured Language Model Programs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b06a2ab-482a-481d-9eb2-7c1c7a9ee975 · inbound
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ed898ed-3df2-4af9-8e02-a7ca12b3ed72 · inbound
Trust Region Masking for Long-Horizon LLM Reinforcement Learning SGLang: Efficient Execution of Structured Language Model Programs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b80fe2-8206-4506-8f36-e2330aecb0f6 · inbound
XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs SGLang: Efficient Execution of Structured Language Model Programs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e83e2a4-75d9-405d-af51-6fd6f86898b4 · inbound
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 089459dc-6f11-4905-bada-132ad7b20679 · inbound
GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0798b9c7-430c-4aa0-9cbc-93e281776b1e · inbound
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System SGLang: Efficient Execution of Structured Language Model Programs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5860a4ef-4582-4754-ad0b-b8ad3d42278e · inbound
Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization SGLang: Efficient Execution of Structured Language Model Programs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a9b69302-04e4-41e5-8070-6a5ff93f2c06 · inbound
Knowledge Packs: Zero-Token Knowledge Delivery via KV Cache Injection SGLang: Efficient Execution of Structured Language Model Programs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a31155bc-1873-4467-bb67-ac2813eea3a0 · inbound
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding SGLang: Efficient Execution of Structured Language Model Programs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8583487d-5047-4bb3-88b7-c142ed5f2737 · inbound
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding SGLang: Efficient Execution of Structured Language Model Programs
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b28353f-f5b0-4002-b37d-a61a505906d6 · inbound
MEMENTO: Teaching LLMs to Manage Their Own Context SGLang: Efficient Execution of Structured Language Model Programs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc9a5121-b1e8-44ef-be3a-2fd6f7c83ed8 · inbound
CodeComp: Structural KV Cache Compression for Agentic Coding SGLang: Efficient Execution of Structured Language Model Programs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e29febe5-7e34-4270-a0c3-f8dbc3898095 · inbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale SGLang: Efficient Execution of Structured Language Model Programs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a753f19f-1917-4b37-8513-bfe471f63066 · inbound
ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems SGLang: Efficient Execution of Structured Language Model Programs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9994de8-27e4-4ba1-838d-5ea605b04250 · inbound
TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3b18b48-1254-4a5d-a25a-529de8129fd7 · inbound
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs SGLang: Efficient Execution of Structured Language Model Programs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7fff1fc6-8867-4895-a2f6-03f312d6bdb6 · inbound
When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis SGLang: Efficient Execution of Structured Language Model Programs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c12f2007-9cf2-4714-83c0-83c01da4da5e · inbound
enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways SGLang: Efficient Execution of Structured Language Model Programs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa6e9a18-a6fd-4710-97ba-828e14b38205 · inbound
enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways SGLang: Efficient Execution of Structured Language Model Programs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e24beab-0a2f-488f-81ce-33a052b932af · inbound
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention SGLang: Efficient Execution of Structured Language Model Programs
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3212af09-1c2c-4a0e-b030-0fbd6925d848 · inbound
LLM StructCore: Schema-Guided Reasoning Condensation and Deterministic Compilation SGLang: Efficient Execution of Structured Language Model Programs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92f9ecf4-7edd-4ae6-8b4f-ce6037819f62 · inbound
Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study SGLang: Efficient Execution of Structured Language Model Programs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ebbbd2b-962b-419c-88b9-dd31f90b1926 · inbound
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts SGLang: Efficient Execution of Structured Language Model Programs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88acee04-583d-4909-b9fb-572ad13a5c49 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db2e29a4-53d6-4d7c-871b-ede04c5fac02 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bc78a40-bdc5-4dfa-9653-382725e49b8e · inbound
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data SGLang: Efficient Execution of Structured Language Model Programs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d13f8ec-8dce-48cb-9396-bbf153318aad · inbound
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs SGLang: Efficient Execution of Structured Language Model Programs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ff193e2-f2f3-4fc6-b274-bc496e382089 · inbound
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs SGLang: Efficient Execution of Structured Language Model Programs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7eba79e-a926-4e8c-bb02-8ab96f49d555 · inbound
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd91c405-cf97-4b68-9aff-eb202e10c2b4 · inbound
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7d4c32f-9706-4f7f-ae37-3ff1f82f4ea3 · inbound
Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use SGLang: Efficient Execution of Structured Language Model Programs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad73c77c-d20c-4fd6-b94a-5547dbb74204 · inbound
VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading SGLang: Efficient Execution of Structured Language Model Programs
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd1b0b98-b4b8-47e6-9803-ec22097f3e1a · inbound
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? SGLang: Efficient Execution of Structured Language Model Programs
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3dc821e-22e1-4b0c-8972-16b2b54f21eb · inbound
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment SGLang: Efficient Execution of Structured Language Model Programs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e5b66506-6dcb-4c0c-9d23-dbe0aefb4728 · inbound
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation SGLang: Efficient Execution of Structured Language Model Programs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28060f05-4488-40bc-9fcd-36cec22f71c6 · inbound
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation SGLang: Efficient Execution of Structured Language Model Programs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c06c480b-8f8d-4fe8-b9b3-33386d6b83fb · inbound
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89ebd297-0578-4958-b9c5-6d581bdfcf51 · inbound
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8398f1d4-a84e-4003-aa33-507fe504b73f · inbound
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c531d5e-010a-446c-9463-834072b0d43d · inbound
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents SGLang: Efficient Execution of Structured Language Model Programs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 17fb4adf-ee85-4319-9dae-e3a30ef57a59 · inbound
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents SGLang: Efficient Execution of Structured Language Model Programs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5123a9c-2b52-4052-bc31-966a2c8ab010 · inbound
An Executable Benchmarking Suite for Tool-Using Agents SGLang: Efficient Execution of Structured Language Model Programs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1f9dcf4-d8f7-4f17-8687-e1c7bf3ff911 · inbound
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding SGLang: Efficient Execution of Structured Language Model Programs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72295d26-52dc-47e1-9d7f-f27904a2e39e · inbound
Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers SGLang: Efficient Execution of Structured Language Model Programs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3aa5c586-3078-4fef-b59c-fcc6aa720414 · inbound
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts SGLang: Efficient Execution of Structured Language Model Programs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cfe0c73-b624-47cf-8e24-8021eb71098b · inbound
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts SGLang: Efficient Execution of Structured Language Model Programs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d3f47e6-ac61-498a-8977-6ada1b1d7dee · inbound
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts SGLang: Efficient Execution of Structured Language Model Programs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4173c3bc-c9bf-48a7-999a-7f97326069e5 · inbound
OpenJarvis: Personal AI, On Personal Devices SGLang: Efficient Execution of Structured Language Model Programs
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21034799-bddb-4951-b43e-ccb6c92e8072 · inbound
Format-Constraint Coupling in Knowledge Graph Construction from Statistical Tables SGLang: Efficient Execution of Structured Language Model Programs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff640e11-23b6-41c9-9f00-e404c7919708 · inbound
Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 83f8063c-b2d8-4ce0-b475-5edcb3c41c8a · inbound
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks SGLang: Efficient Execution of Structured Language Model Programs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9652e40-f9f6-4e5d-8a82-71ac2f4f709e · inbound
Polar: Agentic RL on Any Harness at Scale SGLang: Efficient Execution of Structured Language Model Programs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e38ec5f-ca75-48b3-b722-84089385d332 · inbound
The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e71b475-38ef-4ec8-b547-19e578e0aaa2 · inbound
Stateful Inference for Low-Latency Multi-Agent Tool Calling SGLang: Efficient Execution of Structured Language Model Programs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c334fdb-00fa-4dd5-a7bf-0386f8199c22 · inbound
REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization SGLang: Efficient Execution of Structured Language Model Programs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6fe9a0f9-2032-40cb-b238-589ffeebb5f9 · inbound
A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc15d38b-2c6f-41c7-a2c1-676b57deea6f · inbound
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 884b777e-3abd-4dea-b920-bb4dc0c0956d · inbound
Draft-OPD: On-Policy Distillation for Speculative Draft Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 891319e3-785e-43f7-aed2-e2e2fc9c80cb · inbound
Schedule-Level Shared-Prefix Reuse for LLM RL Training SGLang: Efficient Execution of Structured Language Model Programs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de2fdcd6-e2d9-4f62-b903-82666e52ffb9 · inbound
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads SGLang: Efficient Execution of Structured Language Model Programs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 083c2371-c491-4faa-91e4-952d5f754bab · inbound
DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c80b046-d9a5-4f30-87cd-91ba8bd118ef · inbound
MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU SGLang: Efficient Execution of Structured Language Model Programs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 56210994-c691-4877-bf4b-b9f6d58aa3b6 · inbound
QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f18438d6-ef0c-4909-8003-5f0d4364358d · inbound
Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3fa253e3-bc3e-4ad9-8031-d10c99e06bd7 · inbound
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents SGLang: Efficient Execution of Structured Language Model Programs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ffeb3a49-abcc-4cef-bd56-69957ec9a217 · inbound
Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0146a87-7f23-42cd-a6fb-01535021a337 · inbound
Harnessing Streaming Video in the Wild SGLang: Efficient Execution of Structured Language Model Programs
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6918d03b-d60f-4c51-9979-55617cbe470e · inbound
Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy SGLang: Efficient Execution of Structured Language Model Programs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 159460cc-065c-4e64-935f-8cf432cba2df · inbound
RKSC: Reasoning-Aware KV Cache Sharing and Confident Early Exit for Multi-Step LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d76ef9d0-6300-4f96-a88f-2753b53f9754 · inbound
UltraQuant: 4-bit KV Caching for Context-Heavy Agents SGLang: Efficient Execution of Structured Language Model Programs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25fb1a25-d229-4ec0-a606-2fd3588df483 · inbound
Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9bfadf86-ecfc-4ac3-86b1-237a449e3be3 · inbound
Human-Less LLM Serving: Quantifying the Human Tax on Throughput SGLang: Efficient Execution of Structured Language Model Programs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d5c4e72-1c64-4877-91ba-d608f4bf6484 · inbound
Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 364ee36a-212e-4a44-a757-f59d9f9cfa5e · inbound
The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing SGLang: Efficient Execution of Structured Language Model Programs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd90b1b6-c4c0-4a32-a506-6bfaa07c1a4f · inbound
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2bce52ad-96ec-44e1-ba55-529da006ac69 · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator SGLang: Efficient Execution of Structured Language Model Programs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22ab3e59-a354-4cd7-a124-3bf2c9707e67 · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator SGLang: Efficient Execution of Structured Language Model Programs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60d3431c-d0c7-40eb-86ea-3db7f5baae15 · inbound
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM SGLang: Efficient Execution of Structured Language Model Programs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c30ca5d4-15f2-4a6f-815b-b3bd07ab9313 · inbound
Speculative Pre-Positioning: Decoding Stateful Sessions to the Next Decision Point Off the Critical Path SGLang: Efficient Execution of Structured Language Model Programs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b1f8f9e-290f-43b5-8d04-e15c9c36155f · inbound
Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6298436d-31b8-4141-a5da-aa19cdb95452 · inbound
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a3fabcc-50c9-4279-930f-575aa58c5c33 · inbound
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving SGLang: Efficient Execution of Structured Language Model Programs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e9fe7b7-2041-460d-8eb5-2de22be94f75 · inbound
OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters SGLang: Efficient Execution of Structured Language Model Programs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f56283a-de44-4c79-811a-022269d6b8dc · inbound
MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems SGLang: Efficient Execution of Structured Language Model Programs
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf6d3df-2fdb-4167-8fb6-c62bf9d2021e · inbound
A Workflow-Aware Serving Layer for Agentic Applications SGLang: Efficient Execution of Structured Language Model Programs
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 303e54f1-e3b4-4172-9751-c76ff02f98a7 · inbound
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling SGLang: Efficient Execution of Structured Language Model Programs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.