Pith. sign in

Paper Citation Record · LEDGER

SGLang: Efficient Execution of Structured Language Model Programs

As of 5 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 100 inbound Pith citation observations for arXiv:2312.07104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.07104 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T08:20:01.011625Z

measured 162 of 162 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 126 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:54.657298Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact22
  • verified fuzzy35
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

10
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6a74ddd4-1c8b-4d04-85ea-4a396ab5641c · outbound

This paper cites InferCept: Efficient Intercept Support for Augmented Large Language Model Inference.

SGLang: Efficient Execution of Structured Language Model Programs InferCept: Efficient Intercept Support for Augmented Large Language Model Inference

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.189749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:fbdd4cf47db2a0e0c82e8efa661be502030da2b00dcc56d18a874c4e4dc25e18

Observation b347b4f8-7709-4be5-9d16-9102b1c22db3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SGLang: Efficient Execution of Structured Language Model Programs Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.229584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:d3b75960d370c6e5e2d472a17e3c46df861ef59f95000752c2869fa05e6e5d3d

Observation cfa5f099-8e66-4b48-8788-f110d5523b09 · outbound

This paper cites Deepspeed- inference: enabling efficient inference of transformer models at unprecedented scale.

SGLang: Efficient Execution of Structured Language Model Programs Deepspeed- inference: enabling efficient inference of transformer models at unprecedented scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.232947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:f2c749c56fac6b25470194b7b4b19a87fedf1194802421267ac173e7439f211e

Observation 222e3b06-ed59-4657-9946-75ac28b2a895 · outbound

This paper cites Prompting is programming: A query language for large language models.

SGLang: Efficient Execution of Structured Language Model Programs Prompting is programming: A query language for large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.235949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:16409db99ef6414a262544b2f38b2e2bc38a67ab140ae08bb237968dee33a811

Observation d6225a0e-e2ac-4848-b0ac-42a84a4bb767 · outbound

This paper cites Language models are few-shot learners.

SGLang: Efficient Execution of Structured Language Model Programs Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.239435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:c5aa798eb59829f8e5942f68a509f1c21b7c0b7a2931d0098088abaaab9f65e6

Observation a2bf951c-e3c7-438e-a045-66f735e760f9 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

SGLang: Efficient Execution of Structured Language Model Programs Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.095155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:eadee99af4ef57bf59419bb701334e4ae2d7b83d433d919c2f6f5324746a76f3

Observation 5c510580-7f54-49ab-8e5e-829538ee0a74 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

SGLang: Efficient Execution of Structured Language Model Programs Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.242215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:a667c80af9644b5ea9c2227512a115d1e8907eedbc24f7132c70dc4442d4dd5a

Observation f6666ca9-1e4f-4bdb-b702-6f659369aceb · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SGLang: Efficient Execution of Structured Language Model Programs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:31e1e67a0bd26d602cc776852ab8e5771d4df13b0863e27d6b8ab80c7621d1ef

Observation 5b114185-21e5-47b0-af92-dfe57e4809d8 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

SGLang: Efficient Execution of Structured Language Model Programs PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.071915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:20ea73f142bbce40d7ec33f39f489fa5e26bf6111867e399202d88bc43803a1b

Observation 9a9d5276-b026-4565-9745-586e13278345 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359.

SGLang: Efficient Execution of Structured Language Model Programs Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.245282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:f7830f4bc33a0c4b2f0af6f78093dfee7e5933800a93d5d255294ce601aa4cf2

Observation 8161cbd0-a3ce-4ff1-bcb4-8f19c6ee6dc4 · outbound

This paper cites Model tells you what to discard: Adaptive kv cache compression for llms.

SGLang: Efficient Execution of Structured Language Model Programs Model tells you what to discard: Adaptive kv cache compression for llms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.248190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:9c94a5b88a37d3b2759e2a13cee2dd04240da3b2c35e880aedff2623dfdd176f

Observation dd159a32-1bc2-4141-8132-23950eee6907 · outbound

This paper cites Prompt Cache: Modular Attention Reuse for Low-Latency Inference.

SGLang: Efficient Execution of Structured Language Model Programs Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.060678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:da926bcd90bc334ef122fe633af2d059cbc338023d8bcd229e048553b1bcc670

Observation a3156d99-88b5-4a7f-8d5e-220dd2014b0f · outbound

This paper cites A guidance language for controlling large language models.

SGLang: Efficient Execution of Structured Language Model Programs A guidance language for controlling large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.251034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:28ff1a66bed0db529b82bbbd2bf63013d5c8d6f09d8aca2a17f50ccdb549cd54

Observation 02314b4d-771e-4e98-95cd-fd7a80d63454 · outbound

This paper cites Measuring massive multitask language understanding.

SGLang: Efficient Execution of Structured Language Model Programs Measuring massive multitask language understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.253854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:732c5f78458fa77f10250e5e024c17e90d5d6417ef9a63fad07f7473bf2a8f5d

Observation f0ce2c05-fd48-4cd7-8700-7480987707f1 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

SGLang: Efficient Execution of Structured Language Model Programs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.106533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:09679588d68010f818a84b05f318311813a9257f15160b573e9b0f16f37bf14c

Observation 22dbdc37-18d0-4486-a3d7-f15d598bbfd0 · outbound

This paper cites Text generation inference.

SGLang: Efficient Execution of Structured Language Model Programs Text generation inference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.256958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:8d2342f56272be05bea146634ff7d5d48c7f796881aa951268f2e844880edacc

Observation e88c493e-4ea4-44f2-a79e-04395ba29cc6 · outbound

This paper cites Mixtral of Experts.

SGLang: Efficient Execution of Structured Language Model Programs Mixtral of Experts

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.149376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:13834e206ba7764fb336fa99c67ce8b9c04b167c7d3d6dc83b5f5c0e70529436

Observation 8dc0df40-83bd-4002-b1a2-e7021f3f8de3 · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

SGLang: Efficient Execution of Structured Language Model Programs Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.176768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:3d2850df4a6532626c9dca429bfcd5cdd135dd1cf4def8c3172991790dfdd696

Observation 4c274a29-48b0-4191-a303-036c206199ef · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

SGLang: Efficient Execution of Structured Language Model Programs GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.181115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:18bf59e6e1c5a4f5c84f28437299e49c3cb8df38aad104ad127fde53f70a672d

Observation 102c2646-90ea-498c-95fb-b32e4e104da7 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

SGLang: Efficient Execution of Structured Language Model Programs DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:20:01.184879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:27ac86146ca9355952b23d544052496fc1b1b21b3b06eb62998c813aa2209644

Observation aa6b9cdd-d472-4b54-830b-b2b035656591 · outbound

This paper cites An LLM Compiler for Parallel Function Calling.

SGLang: Efficient Execution of Structured Language Model Programs An LLM Compiler for Parallel Function Calling

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.051787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:c7369f46a70ddc50d82f96824721c2f4d043bd990c162f2de20d0aefa86aab0b

Observation 467bc4d9-5eeb-4dd8-ab7d-7288bc48faba · outbound

This paper cites Validating large language models with relm.

SGLang: Efficient Execution of Structured Language Model Programs Validating large language models with relm

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.259766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:8cfa72dc19962b6c3635362a1a39b5974cfccb8cf4df3472c68780eb4a2dfcda

Observation cd4e539b-a689-46ef-8cbb-807e5a8d3822 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

SGLang: Efficient Execution of Structured Language Model Programs Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.262617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:2a80b4fb02389836bf3403b7c47a062e0688bfed1d6f6c137cd11842dabd6c87

Observation b6ae0b63-a0e2-4f5e-b038-e68abab52220 · outbound

This paper cites Langchain.

SGLang: Efficient Execution of Structured Language Model Programs Langchain

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.265246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:cc1f8fe2adbef2a1b6f9cecdc69ca378905e9f44a28e6ce881ed3de52b0e8df9

Observation 93dea5ef-e549-421f-812b-9e58b5adca8d · outbound

This paper cites Competition-level code generation with alphacode.

SGLang: Efficient Execution of Structured Language Model Programs Competition-level code generation with alphacode

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.267951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:c3f1ddf5576b473124e76272a42b87cdc3528607b9772e75f46399b247c26e33

Observation 11f75dab-de9a-4207-a686-38d025a71289 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

SGLang: Efficient Execution of Structured Language Model Programs AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.111415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:bcbb03717996202cfca3d6710a0ce41e7820449ac66a8ce2abc67dce8a62511d

Observation 6bef4281-47d3-4d0b-a34a-3135424c8f37 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

SGLang: Efficient Execution of Structured Language Model Programs Improved Baselines with Visual Instruction Tuning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:34.143871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:28aab564677dad9716c647f5a1d9f10f85f81e5aa8b83571877589e9f08224fe

Observation 97d09804-584a-42d1-94db-9eaecdbba861 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

SGLang: Efficient Execution of Structured Language Model Programs Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.270855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:47ed5f545859916ceb9423dcba171eb251de122dbc39ae7e56dada492033d34c

Observation e2fe388f-f18b-4bd7-9039-26cb1fcc4d43 · outbound

This paper cites Optimizing LLM Queries in Relational Data Analytics Workloads.

SGLang: Efficient Execution of Structured Language Model Programs Optimizing LLM Queries in Relational Data Analytics Workloads

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.136725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:4047f0c81cf3e5b91eab1718f932d9c8213d933bc4e6452d37e984db67a0e838

Observation 119edb1b-f6fc-4628-8d29-d3ec447e9e46 · outbound

This paper cites Prompting Frameworks for Large Language Models: A Survey.

SGLang: Efficient Execution of Structured Language Model Programs Prompting Frameworks for Large Language Models: A Survey

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.146006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:effd4fce87b74ba65c0b5622a0f5db72f38ab471c77cb2b5ec4bff71c15057eb

Observation a53583b1-1e75-4360-80df-7418980dded1 · outbound

This paper cites Scissorhands: Exploiting the persistence of impor- tance hypothesis for llm kv cache compression at test time.

SGLang: Efficient Execution of Structured Language Model Programs Scissorhands: Exploiting the persistence of impor- tance hypothesis for llm kv cache compression at test time

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.274141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:e6062e72538035124ecd634beb1eb3b176e248f9b10713364b26e708ac1d2ea9

Observation 0735568f-dbb2-4efc-ad5b-90524c111cb0 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

SGLang: Efficient Execution of Structured Language Model Programs KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:53:12.376678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:4bfc045c209c2ff780fb566969164946dfd0a5d5cb0c1aa6ddde441fde519a95

Observation 9ed3cef9-2a5a-4609-8810-538bcadcb404 · outbound

This paper cites Skeleton- of-thought: Prompting LLMs for efficient parallel generation.

SGLang: Efficient Execution of Structured Language Model Programs Skeleton- of-thought: Prompting LLMs for efficient parallel generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.277308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:cf19ab8a1e61839a8c62a8a6941f3abef0ad6fa2e78cc8352db8cf1ce74b44a4

Observation b9e6dfef-c1e3-4282-833e-29fc99778ec7 · outbound

This paper cites Tensorrt-llm.

SGLang: Efficient Execution of Structured Language Model Programs Tensorrt-llm

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.280050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:6ed6b9654b231f19c33b8d36bac61ea17e27df21a798692bd2da8eec43de5fba

Observation fe9e2e16-12c9-4755-886d-040f99da480a · outbound

This paper cites Gpt-4 technical report.

SGLang: Efficient Execution of Structured Language Model Programs Gpt-4 technical report

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.282769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:10b99c72c6545e0d93732edbcd9741eff46ced3fda773c28e541fcace37080c0

Observation b5ed99bc-5455-44d8-95c2-bc8ba52ea017 · outbound

This paper cites O’Brien, Carrie J.

SGLang: Efficient Execution of Structured Language Model Programs O’Brien, Carrie J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.285365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:f5a9a6a570732a02c624e7e6c9eb1696c14709ade415e0e421f1c0eef601d47f

Observation d7b9b901-6dde-44a2-b694-c3e5adef8cf2 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

SGLang: Efficient Execution of Structured Language Model Programs Pytorch: An imperative style, high-performance deep learning library

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.288995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:f2971fee56c59d9f49cef4e09617c31c0a235ac199d48e1b7a025fd39e2dfc42

Observation e4c379df-e753-4d3e-a163-ec6de16abe71 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

SGLang: Efficient Execution of Structured Language Model Programs Gorilla: Large Language Model Connected with Massive APIs

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:20:01.065788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:1c3d5a6ff8a4d5e0fd8a09d760e5271eda3107b9d549b816da31a7c9e63041a3

Observation 3f5f4433-affe-49a3-a4e5-34b4640bff3c · outbound

This paper cites Efficiently scaling transformer inference.

SGLang: Efficient Execution of Structured Language Model Programs Efficiently scaling transformer inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.291869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:1262aa6649df6124fc860aa021a15ad0d3bb1443a017fef13a2486a7b80eae7e

Observation ee83fe8c-3d29-44d9-8141-754db48372c6 · outbound

This paper cites Branch-Solve-Merge Improves Large Language Model Evaluation and Generation.

SGLang: Efficient Execution of Structured Language Model Programs Branch-Solve-Merge Improves Large Language Model Evaluation and Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.078104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:83646891e4ed25e7dc147a6218f94a0ff1953a93f996c46aff8836049086313e

Observation 29cfe3c7-7df4-46dc-aa8e-6f7cbf7fd161 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

SGLang: Efficient Execution of Structured Language Model Programs Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.083748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:cdfa730a7edbf5c9b3703db9d11c6062fcba415742fd0df74cbcc37f9070b019

Observation d50df4e9-57fb-45d3-a95e-ca3f44752149 · outbound

This paper cites Fairness in Serving Large Language Models.

SGLang: Efficient Execution of Structured Language Model Programs Fairness in Serving Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.089977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:9e2ae4cd6269d6e998b91ededd8b84a4c788f009e40741d489175cdd82a63c2d

Observation db7e1274-624e-4de8-a56f-f16fe85fe872 · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu.

SGLang: Efficient Execution of Structured Language Model Programs Flexgen: high-throughput generative inference of large language models with a single gpu

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.294768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:aae22945afd56deac956351e3f3635edee5aa0745aae37e4ee8fd6b59c9e4245

Observation b9d0f233-c377-4ed9-8dfa-7df0dfe79328 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SGLang: Efficient Execution of Structured Language Model Programs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.100447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:bad9cd40142c73f97839b4f657fa3f480aafa5242aefdcf37050a215f8e7321d

Observation 40c2f736-b5ee-4e78-8293-49a484249d2e · outbound

This paper cites Preble: Efficient distributed prompt scheduling for llm serving.

SGLang: Efficient Execution of Structured Language Model Programs Preble: Efficient distributed prompt scheduling for llm serving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.297834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:fcaad80c0571f787628a2a63c38a605c04afef0a61aff06a193020699c7ea541

Observation 5bddc7cd-152d-4182-a952-ad1bd8e58035 · outbound

This paper cites Cognitive architec- tures for language agents.

SGLang: Efficient Execution of Structured Language Model Programs Cognitive architec- tures for language agents

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.300984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:f39ed97c8d7428b19f439a57fab2f69ad42bb3ac787d0d927bd2beaa1f292164

Observation 3fd41751-957f-4c58-8d78-565e7ead93f6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SGLang: Efficient Execution of Structured Language Model Programs Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.116096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:03f55c0a53278f375f9ef7874005f3bf8602e333042f4582b145c9ce476a758d

Observation dc261552-7ec8-47db-b99d-d54239b727d0 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations.

SGLang: Efficient Execution of Structured Language Model Programs Triton: an intermediate language and compiler for tiled neural network computations

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.304576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:1bfce10b4c7358def1701663688e931380d72aefb1ca63f20bc26da4bc575625

Observation 81c43e3d-5c12-4c9d-bcb2-9a7a86abf7da · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SGLang: Efficient Execution of Structured Language Model Programs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.126803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:4b0a6d73b8ac053bfd5b12693bf0bcf0453bab0f4974d43576cefe7f9bca5689

Observation 635e28f4-f792-4512-9142-8d833297cc99 · outbound

This paper cites Fast, high-fidelity llm decoding with regex constraints.

SGLang: Efficient Execution of Structured Language Model Programs Fast, high-fidelity llm decoding with regex constraints

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.194759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:bfa1359c8033b53a0adfd6944e458cf5addf7c9b37320bd4beffe4bd160f6ef6

Observation 3a590169-463e-4d0d-b548-b48f07c1d48c · outbound

This paper cites Attention is all you need.

SGLang: Efficient Execution of Structured Language Model Programs Attention is all you need

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.200070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:3ae7af133c26ed3debecf876f12ed71c47380420d691d0fcdf3dd4ac1eb4299e

Observation fa982f5a-1680-4661-b7d7-86df00d044bd · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SGLang: Efficient Execution of Structured Language Model Programs Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.141382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:8d216a31f4e3097b275db2c8b088e52b1d640f18fccbd569cc963a7293a8af60

Observation d131e241-e0a1-45ad-bd2e-3af1639e5821 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

SGLang: Efficient Execution of Structured Language Model Programs Self-consistency improves chain of thought reasoning in language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.204425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:40d7320de3a8536d13bf9e55f165c144e49a33900e4072beba2bc44adbad4354

Observation edf0df46-2270-4f2e-aaf1-a2aaa2880b3b · outbound

This paper cites Efficient guided generation for large language models.

SGLang: Efficient Execution of Structured Language Model Programs Efficient guided generation for large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.207700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:3798ea3f493fde038eb24978f8822d5d63fe733d5b7f95550158255bb8e22b71

Observation b6b32f6d-5a96-421c-ba41-0df032b1ed7f · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

SGLang: Efficient Execution of Structured Language Model Programs AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.158301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:f1d96574fbacfb606fb8dba76f0fc0ef41df1b1314f1c5db3dbbc9fd834a452b

Observation 8bc1570a-30a3-4a00-91ed-eea3e61fb91e · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

SGLang: Efficient Execution of Structured Language Model Programs Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.162196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:78d838d9643b0a589ef57c1806ccfa18f00b0935e33a56d1f0a57d4de7619d4e

Observation 1144386f-bf4b-473f-b9dc-d2c6938a4606 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

SGLang: Efficient Execution of Structured Language Model Programs React: Synergizing reasoning and acting in language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.211481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:b889879be895e945d6a2574c58d8bc21804196dcf74b9d94a71e798078425eec

Observation 15a4b487-9e26-478c-9c62-a13b250ca0c4 · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

SGLang: Efficient Execution of Structured Language Model Programs ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.172272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:2d788c335677fd04461f715ac6d6a5ac6cf9d595be702f84977ac5d1f5adc5ce

Observation 8f022368-f493-4cc7-863c-594ca2e0963c · outbound

This paper cites Accelerating self-attentions for llm serving with flashinfer, February 2024.

SGLang: Efficient Execution of Structured Language Model Programs Accelerating self-attentions for llm serving with flashinfer, February 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.215360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:96c965e8fbd33dc2874f631e70dbe3f3d677eb02bd201181f29e9f1704abed17

Observation 13e324ef-f3f6-4ab6-b7ae-664bf3524902 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

SGLang: Efficient Execution of Structured Language Model Programs Orca: A distributed serving system for {Transformer-Based} generative models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.218909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:d75cd306d2ffc218e159098734c4fa997ac017d540d506ffcb5396ccbb23296a

Observation cb642d8f-19a2-4f52-9d32-21eb415a57ea · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800.

SGLang: Efficient Execution of Structured Language Model Programs Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.222798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:4f4be57e567974866d77778bbc4c44bea47321af9851f3c781f88ba7072f18dc

Observation cc89d572-cdb3-4a1c-8ad9-2c4b34fb3883 · outbound

This paper cites prefill"). It then sequentially decodes output tokens, with each token depending on prior tokens (this process is called.

SGLang: Efficient Execution of Structured Language Model Programs prefill"). It then sequentially decodes output tokens, with each token depending on prior tokens (this process is called

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T08:20:01.226418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:181751ae6b8765bd2341d683eea41bf40362dbe634117f8ace69a64997546ec8

Pith citing papers

Observation d65f1508-ed36-4284-a44a-59fa6e2af223 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:39:33.414625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:fa5fdb01f11742c5711f2c651038b580a188f3a6cd866fcfc1e108b298eb59de

Observation c34b8ce4-76bf-4cc0-b6fe-4eca1d791962 · inbound

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference cites this paper.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.126000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:569e7a50d3b7c201b26f2ae6bb59e74ddafd1b964e877d42d60e104dfe57211f

Observation 2611a68e-8f2e-4ccf-8031-76d4f7472c8f · inbound

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling cites this paper.

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling SGLang: Efficient Execution of Structured Language Model Programs

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:42:23.297389Z digest=sha256:d3a34b056417b189b0e6434f15bf8fb4b36185ab67afa2bd1a426328f0ab09e6

Observation 69661b31-10ff-46b6-bc7c-c34df8a156de · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge SGLang: Efficient Execution of Structured Language Model Programs

Reference 222

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T17:35:43.579524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:b297b7c4622d6ccf92b4d732c824ab568d7cfcce0e7408140d37966c1a949b2b

Observation 1b756769-b3f8-4ab6-aac6-c6f91020a592 · inbound

FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving cites this paper.

FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 16

Resolution
malformed identifier
local_arxiv, observed 2026-05-16T13:26:34.815784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T13:26:34.747336Z digest=sha256:579152c57098cafe148f76104647bfdf35afa21415b4e630d0d0e28a7974c2f1

Observation 4fad1018-e2de-4510-be2b-8afa984b15a2 · inbound

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference cites this paper.

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T21:17:07.965750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T21:16:31.655330Z digest=sha256:d6571932426145c82ff111dcce5c3960ce271f9ee1b877c455cb101fecd977d5

Observation ec42c426-1a3b-4ee4-a42c-26f989f42211 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production SGLang: Efficient Execution of Structured Language Model Programs

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T15:44:58.129988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:a86c9c0b577b62dad69f241f31fdb6635f259bd26877280524a57d5dc9923d0a

Observation 3ebf05aa-7325-4941-b15a-c360b46e99c5 · inbound

Hermes 4 Technical Report cites this paper.

Hermes 4 Technical Report SGLang: Efficient Execution of Structured Language Model Programs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:54.657298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:32:54.657298Z digest=sha256:485594b0c66c74177f595f08063555da1fef509795bf4cd34f495dd2019d624b

Observation 403bab14-b320-418e-8c40-0c84715ee33a · inbound

Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics cites this paper.

Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics SGLang: Efficient Execution of Structured Language Model Programs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:02.319006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:02.319006Z digest=sha256:659ac9a312a9d6fc48b748a0615a5ebd8c6c798ce863e97748624f4254bf344f

Observation c11e8fe6-25f0-43d0-9d44-4fe37e879ddd · inbound

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution cites this paper.

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution SGLang: Efficient Execution of Structured Language Model Programs

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:13.051506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:13.051506Z digest=sha256:78b4d83a8d74fcbc1a642f8cf9a73ad7d91d3fb685a51f4477925c549984cb32

Observation 496c1c4f-4c06-4eaa-8077-690f696a7307 · inbound

CacheClip: Accelerating RAG with Effective KV Cache Reuse cites this paper.

CacheClip: Accelerating RAG with Effective KV Cache Reuse SGLang: Efficient Execution of Structured Language Model Programs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:41:33.796275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T12:36:41.630618Z digest=sha256:c63116e84c2500b7dd8cbc8bb953102eca3b692b5e51fe45d46d093b6021f297

Observation 2d2e4ea5-9283-43a2-a1a0-ad804922f90c · inbound

DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference cites this paper.

DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 20

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T04:32:23.059538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T04:31:40.360872Z digest=sha256:5a17f0f6f8d583484ddab386934c9b8aba7cb0346a22e6003a7fb1d7fa13fcbe

Observation a713be5b-4cec-48fd-b760-e990a3df4b86 · inbound

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators cites this paper.

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators SGLang: Efficient Execution of Structured Language Model Programs

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:00:39.397829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T01:59:30.583027Z digest=sha256:db3c9c29101dfe2da6283c294e5f18381f51c83505936ce7faf53f1f46ecfb57

Observation 9810cff1-4b96-47b4-9d9d-6f7b22f0db94 · inbound

Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework cites this paper.

Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework SGLang: Efficient Execution of Structured Language Model Programs

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T04:24:00.343515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:23:03.393566Z digest=sha256:40e8f62e489beb2a04c9ad9e98366382f3f0b0e31ce99193fe818e2cb3857093

Observation 9b06a2ab-482a-481d-9eb2-7c1c7a9ee975 · inbound

Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving cites this paper.

Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:21:18.696657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:20:48.239186Z digest=sha256:ed09a9b3b4b46a1a795a9b3101de0b731e615cec9201720234ad3c84bbe36f79

Observation 2ed898ed-3df2-4af9-8e02-a7ca12b3ed72 · inbound

Trust Region Masking for Long-Horizon LLM Reinforcement Learning cites this paper.

Trust Region Masking for Long-Horizon LLM Reinforcement Learning SGLang: Efficient Execution of Structured Language Model Programs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T13:48:01.890708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T13:48:01.890708Z digest=sha256:9c005a18c165823a9da20f2c0bb3f3f794d765b088230cec1b1606df595c6c52

Observation 53b80fe2-8206-4506-8f36-e2330aecb0f6 · inbound

XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs cites this paper.

XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs SGLang: Efficient Execution of Structured Language Model Programs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:07:34.766774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:07:34.766774Z digest=sha256:307c9ab63b31dc74b31ac47deeff582f108ad8bc9c7619c932d94340b9745c85

Observation 4e83e2a4-75d9-405d-af51-6fd6f86898b4 · inbound

Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference cites this paper.

Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:27:55.914398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T13:23:16.969774Z digest=sha256:cc7972857f138e40f60406772cfbac32fa5ebf3ffa08872aa93dcae2e012e841

Observation 089459dc-6f11-4905-bada-132ad7b20679 · inbound

GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving cites this paper.

GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:06:25.206499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:06:25.206499Z digest=sha256:d8d67696846346ce2a21b2ebcc8324f74460a70a7d98badb81585ad7cbc3b4a8

Observation 0798b9c7-430c-4aa0-9cbc-93e281776b1e · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System SGLang: Efficient Execution of Structured Language Model Programs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:15.224685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:15.224685Z digest=sha256:2e87d8ff1670c5a544cd4813d67db9334088cb677066e2aa072054889232cfc9

Observation 5860a4ef-4582-4754-ad0b-b8ad3d42278e · inbound

Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization cites this paper.

Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization SGLang: Efficient Execution of Structured Language Model Programs

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:18:04.293566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T22:17:13.431164Z digest=sha256:30cbc53707cb89ff61798e42da9440c2bb54b2a2210dead2d5f9a34db72eb898

Observation a9b69302-04e4-41e5-8070-6a5ff93f2c06 · inbound

Knowledge Packs: Zero-Token Knowledge Delivery via KV Cache Injection cites this paper.

Knowledge Packs: Zero-Token Knowledge Delivery via KV Cache Injection SGLang: Efficient Execution of Structured Language Model Programs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:15:12.068501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:11:13.865969Z digest=sha256:bc57ba77cfa2772dcd15e376c889cde9dd7874e281d7918602f7a36933ee1b72

Observation a31155bc-1873-4467-bb67-ac2813eea3a0 · inbound

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding cites this paper.

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding SGLang: Efficient Execution of Structured Language Model Programs

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T03:17:12.246827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T03:13:37.202362Z digest=sha256:6b127e8d698645f421b7ef10e104783bacda19659fdc4a017a9531944f63b26b

Observation 8583487d-5047-4bb3-88b7-c142ed5f2737 · inbound

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding cites this paper.

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding SGLang: Efficient Execution of Structured Language Model Programs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T02:40:59.456464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:40:59.456464Z digest=sha256:4c3d4cac912b5d8883a16744730af6b4a602c9cc9a09acbd7a379a5636da4e57

Observation 1b28353f-f5b0-4002-b37d-a61a505906d6 · inbound

MEMENTO: Teaching LLMs to Manage Their Own Context cites this paper.

MEMENTO: Teaching LLMs to Manage Their Own Context SGLang: Efficient Execution of Structured Language Model Programs

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T17:16:36.784237Z digest=sha256:9ef5fe704fe20dafcf33cfb8e7da60e799a0bf2200f622e4bcde452d124fba83

Observation dc9a5121-b1e8-44ef-be3a-2fd6f7c83ed8 · inbound

CodeComp: Structural KV Cache Compression for Agentic Coding cites this paper.

CodeComp: Structural KV Cache Compression for Agentic Coding SGLang: Efficient Execution of Structured Language Model Programs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:44.149481Z digest=sha256:b8c18fca86258ce965de9fbdead0447b95df43fd3700a7c232eb5283d5ebfab1

Observation e29febe5-7e34-4270-a0c3-f8dbc3898095 · inbound

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale cites this paper.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale SGLang: Efficient Execution of Structured Language Model Programs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:447badcdf8a7a50a53eade36da10274e57a36ea5faccc190dbd31f42f44b4d31

Observation a753f19f-1917-4b37-8513-bfe471f63066 · inbound

ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems cites this paper.

ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems SGLang: Efficient Execution of Structured Language Model Programs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:27:18.594869Z digest=sha256:e41c7a7b3c8c2b5d2db8c922fcfbb3a32d20a42a216969b8f834a454ba643782

Observation f9994de8-27e4-4ba1-838d-5ea605b04250 · inbound

TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models cites this paper.

TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:45:07.257159Z digest=sha256:e1f8f83707772026a7630362e8bed6349878a961daba2b5b82bcc1d4e3dcb38c

Observation d3b18b48-1254-4a5d-a25a-529de8129fd7 · inbound

Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs cites this paper.

Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs SGLang: Efficient Execution of Structured Language Model Programs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:35:58.213556Z digest=sha256:0ffd43ad96c05534b672ca18b2dfed1474e374dfccc4d2212f07970b0e1fd695

Observation 7fff1fc6-8867-4895-a2f6-03f312d6bdb6 · inbound

When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis cites this paper.

When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis SGLang: Efficient Execution of Structured Language Model Programs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:50:54.128824Z digest=sha256:0a0abb23ef8da1b02b875d454a70df81b8f3a191ec9f7b4f7f3dd25839400ea1

Observation c12f2007-9cf2-4714-83c0-83c01da4da5e · inbound

enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways cites this paper.

enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways SGLang: Efficient Execution of Structured Language Model Programs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:11:06.389107Z digest=sha256:a1c93c61b38a66e52e5dcd18f8d150a790e89ea7169af1f00e6a073e1bf5a0ee

Observation fa6e9a18-a6fd-4710-97ba-828e14b38205 · inbound

enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways cites this paper.

enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways SGLang: Efficient Execution of Structured Language Model Programs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:50:36.175721Z digest=sha256:4e7547909ad1887594cc5fb6220a835cf06662bde214c2e7b3e6b816c87ea0fe

Observation 6e24beab-0a2f-488f-81ce-33a052b932af · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention SGLang: Efficient Execution of Structured Language Model Programs

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:11422f3c613d0aa588e44e0414c737b50bb282aa30fe9203d5bc400ab396599a

Observation 3212af09-1c2c-4a0e-b030-0fbd6925d848 · inbound

LLM StructCore: Schema-Guided Reasoning Condensation and Deterministic Compilation cites this paper.

LLM StructCore: Schema-Guided Reasoning Condensation and Deterministic Compilation SGLang: Efficient Execution of Structured Language Model Programs

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:43:22.473104Z digest=sha256:bed061e773fa625810f68029667ddc2a86944303dd432e41953b51508dc47c5b

Observation 92f9ecf4-7edd-4ae6-8b4f-ce6037819f62 · inbound

Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study cites this paper.

Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study SGLang: Efficient Execution of Structured Language Model Programs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:19:35.815472Z digest=sha256:557b11807ab58258acb7bde283458d1a0d162b0a4a9241759fcdb82a9d992889

Observation 3ebbbd2b-962b-419c-88b9-dd31f90b1926 · inbound

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts cites this paper.

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts SGLang: Efficient Execution of Structured Language Model Programs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:39:35.128879Z digest=sha256:0819bb281d9da372753283a77c10e6a07c8e6b390a5b554a03f8fd0b4ec0efab

Observation 88acee04-583d-4909-b9fb-572ad13a5c49 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:01:27.694910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:30:35.264804Z digest=sha256:906dfd2cddff116b26b4cd95a546889ce3eefdc1dda6ab13d04e1010e9170f65

Observation db2e29a4-53d6-4d7c-871b-ede04c5fac02 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:55:34.747200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:54:41.845820Z digest=sha256:4751bd7f5127ded9e876fa47bc23e9fede19fccf7b0634841d641ded2d72f753

Observation 1bc78a40-bdc5-4dfa-9653-382725e49b8e · inbound

SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data cites this paper.

SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data SGLang: Efficient Execution of Structured Language Model Programs

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T18:16:03.307938Z digest=sha256:fd4328a5b72fa9f533d0e95a5c6b70f268bc0cdc44c2dc996e0c0606c78e65b1

Observation 0d13f8ec-8dce-48cb-9396-bbf153318aad · inbound

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs cites this paper.

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs SGLang: Efficient Execution of Structured Language Model Programs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T01:27:54.164991Z digest=sha256:27328ab0f84bdd3a40dab28df193aac99d80973878672652cc06e77c40d70c11

Observation 3ff193e2-f2f3-4fc6-b274-bc496e382089 · inbound

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs cites this paper.

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs SGLang: Efficient Execution of Structured Language Model Programs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T16:57:26.334739Z digest=sha256:8e447bbc44ea7d54412335d4b71d336c5d144710031580a03bfd67b142e6ed28

Observation b7eba79e-a926-4e8c-bb02-8ab96f49d555 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:530b4320acffcc614b521b9b52fc6ffd828af3bff8ba68f1643e2b143285193b

Observation cd91c405-cf97-4b68-9aff-eb202e10c2b4 · inbound

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving cites this paper.

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:49:24.880528Z digest=sha256:78a5c818e37545139ebd7b236135de70dbb171a3354ebd1ba460a16f684b1584

Observation b7d4c32f-9706-4f7f-ae37-3ff1f82f4ea3 · inbound

Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use cites this paper.

Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use SGLang: Efficient Execution of Structured Language Model Programs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T16:38:04.876480Z digest=sha256:3ba4c6e46c8dbb692e573c577253ba5ba48000c733848d41eec387c2e3ec2085

Observation ad73c77c-d20c-4fd6-b94a-5547dbb74204 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading SGLang: Efficient Execution of Structured Language Model Programs

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:a4ee5f0a332f4fc83ea1b4e7dd2fa47e05fbe3b310faf4d23bfe7c2f3f98e460

Observation fd1b0b98-b4b8-47e6-9803-ec22097f3e1a · inbound

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? cites this paper.

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? SGLang: Efficient Execution of Structured Language Model Programs

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:44:55.943351Z digest=sha256:0941bf30884bfc701ac3dac5388584698fdaad20c66baa995d3568ca4c4bf1b2

Observation b3dc821e-22e1-4b0c-8972-16b2b54f21eb · inbound

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment cites this paper.

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment SGLang: Efficient Execution of Structured Language Model Programs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:25:30.458137Z digest=sha256:0e71e9c4dc14b1c40c249b35ca5741b9e742bf63748fdeb77ef9875fc3662346

Observation e5b66506-6dcb-4c0c-9d23-dbe0aefb4728 · inbound

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation cites this paper.

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:42:32.739352Z digest=sha256:f4ca99c76595ebf2e0c0463ff1adaaebe236e16f9b1260653d57fd13a81a22a9

Observation 28060f05-4488-40bc-9fcd-36cec22f71c6 · inbound

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation cites this paper.

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:21:23.945417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T10:20:45.375375Z digest=sha256:b755b6eedfd97b088d926c91f710d7900815658064d675aa42516e09bc419eb6

Observation c06c480b-8f8d-4fe8-b9b3-33386d6b83fb · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:57:30.104609Z digest=sha256:56a9eb568d7d87278540105c13ea899370473d158589e810584eb3fa4d5a47cd

Observation 89ebd297-0578-4958-b9c5-6d581bdfcf51 · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:15:32.464109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:05:44.256565Z digest=sha256:201bdf2fd6edc1e128c87db989e21f9b34349c9f46b1b8c642f147be825343f3

Observation 8398f1d4-a84e-4003-aa33-507fe504b73f · inbound

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference cites this paper.

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:30:58.729425Z digest=sha256:8cf160ef3d6e3dae0ba5ea1b11b5e05c1ceeb4db0f0257f007f8e98e12205542

Observation 4c531d5e-010a-446c-9463-834072b0d43d · inbound

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents cites this paper.

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents SGLang: Efficient Execution of Structured Language Model Programs

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:15:40.042348Z digest=sha256:4dfb014f507405351201508c34f8d40099595a724eb4b3302e2125eab5a207b6

Observation 17fb4adf-ee85-4319-9dae-e3a30ef57a59 · inbound

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents cites this paper.

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents SGLang: Efficient Execution of Structured Language Model Programs

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T13:55:46.345891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:30:28.649803Z digest=sha256:3ca506cbe751266c5609b9090874263b10f80f5dbfc94c3b876e471e95682a52

Observation f5123a9c-2b52-4052-bc31-966a2c8ab010 · inbound

An Executable Benchmarking Suite for Tool-Using Agents cites this paper.

An Executable Benchmarking Suite for Tool-Using Agents SGLang: Efficient Execution of Structured Language Model Programs

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:52:06.208342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:08:46.020900Z digest=sha256:6f8b27a8f605de3df5a8be3978373242dd69b48cad73c3733825407f4e54ad8e

Observation a1f9dcf4-d8f7-4f17-8687-e1c7bf3ff911 · inbound

NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding cites this paper.

NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding SGLang: Efficient Execution of Structured Language Model Programs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:37:11.445363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T03:32:28.187814Z digest=sha256:b6ed6f41c44e1f1d4a76a4fd2ad97485b76162302f7018980ef4355527bb2e23

Observation 72295d26-52dc-47e1-9d7f-f27904a2e39e · inbound

Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers cites this paper.

Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers SGLang: Efficient Execution of Structured Language Model Programs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:22:51.296567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T19:19:43.718149Z digest=sha256:eb16b01c5eb770b08f9d0465dff5e8c8b36505d2119d97dd12a54cc51d39c901

Observation 3aa5c586-3078-4fef-b59c-fcc6aa720414 · inbound

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts cites this paper.

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts SGLang: Efficient Execution of Structured Language Model Programs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.385803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T15:48:59.800756Z digest=sha256:97c3356d9bf8111e2143ef15a1ca06a627f8e40944c45cc881de2072a77cd81b

Observation 5cfe0c73-b624-47cf-8e24-8021eb71098b · inbound

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts cites this paper.

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts SGLang: Efficient Execution of Structured Language Model Programs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.569758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T20:58:51.141871Z digest=sha256:11514776f2f4ef6d2d78a532c2aa3a3a38c88f4799cf16193a6295bb0cfa3228

Observation 6d3f47e6-ac61-498a-8977-6ada1b1d7dee · inbound

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts cites this paper.

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts SGLang: Efficient Execution of Structured Language Model Programs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T14:03:10.541975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:03:10.541975Z digest=sha256:da087534927e7c983c345a22c4fa0e8d2056b3e112e7a2c85938fce8e0f47dea

Observation 4173c3bc-c9bf-48a7-999a-7f97326069e5 · inbound

OpenJarvis: Personal AI, On Personal Devices cites this paper.

OpenJarvis: Personal AI, On Personal Devices SGLang: Efficient Execution of Structured Language Model Programs

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:28:21.232464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T14:26:16.480948Z digest=sha256:dee870a36f549c2d5de8af491b7667d35d93b8743a7c343510bd7069d2bd06c2

Observation 21034799-bddb-4951-b43e-ccb6c92e8072 · inbound

Format-Constraint Coupling in Knowledge Graph Construction from Statistical Tables cites this paper.

Format-Constraint Coupling in Knowledge Graph Construction from Statistical Tables SGLang: Efficient Execution of Structured Language Model Programs

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T06:44:41.999964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T06:43:31.886617Z digest=sha256:c4858b4bc2b4b85e81431a58ef8123e069b58efc3a2e203e7a5fdc4337e1cfcc

Observation ff640e11-23b6-41c9-9f00-e404c7919708 · inbound

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference cites this paper.

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T07:31:14.081205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T07:28:35.020478Z digest=sha256:eafa4158e746c81476413e5c2802198ca706b4c0a8795ddec2e4e0c6d81d139c

Observation 83f8063c-b2d8-4ce0-b475-5edcb3c41c8a · inbound

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks cites this paper.

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks SGLang: Efficient Execution of Structured Language Model Programs

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T15:44:48.303225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T15:42:21.405913Z digest=sha256:1c599eed4bf4696a61e56bac4515d6b1dff7ce9042fdc53d0e1f4773676c93a2

Observation b9652e40-f9f6-4e5d-8a82-71ac2f4f709e · inbound

Polar: Agentic RL on Any Harness at Scale cites this paper.

Polar: Agentic RL on Any Harness at Scale SGLang: Efficient Execution of Structured Language Model Programs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.311780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T14:38:14.073340Z digest=sha256:15945c2ddd58961b702089ee2733c386500b066625acef95e337f7758ff54c9a

Observation 2e38ec5f-ca75-48b3-b722-84089385d332 · inbound

The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models cites this paper.

The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:44:57.797981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:39:16.469049Z digest=sha256:48416d4b545d6a4e962639655bac163a928a96a6aaf76e1670add4a8346689d0

Observation 4e71b475-38ef-4ec8-b547-19e578e0aaa2 · inbound

Stateful Inference for Low-Latency Multi-Agent Tool Calling cites this paper.

Stateful Inference for Low-Latency Multi-Agent Tool Calling SGLang: Efficient Execution of Structured Language Model Programs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.523426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T22:39:27.600051Z digest=sha256:c24410626c3daaca6cc3f95929c613321a89da35f8088fe5ae871557c827b00e

Observation 0c334fdb-00fa-4dd5-a7bf-0386f8199c22 · inbound

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization cites this paper.

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization SGLang: Efficient Execution of Structured Language Model Programs

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:23:50.724064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:19:10.346738Z digest=sha256:e36831a8de6349d22e7b34737e4ac361abf70ddab48ff524f5ebedf9160dd788

Observation 6fe9a0f9-2032-40cb-b238-589ffeebb5f9 · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:03:47.821304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:2b071645c46da12e491394a3094a12085abfd3fb5826d10407e018ab99ff75dd

Observation fc15d38b-2c6f-41c7-a2c1-676b57deea6f · inbound

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference cites this paper.

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.821608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T10:07:06.141581Z digest=sha256:e03ac63acd1a119d53d754b8890b507eff5298742f576c92b096aabbf88fa6a2

Observation 884b777e-3abd-4dea-b920-bb4dc0c0956d · inbound

Draft-OPD: On-Policy Distillation for Speculative Draft Models cites this paper.

Draft-OPD: On-Policy Distillation for Speculative Draft Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:43:13.246143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:33:39.119309Z digest=sha256:509179db81903ba070676695f414567ff3073b7310557ca82ca495d1003d973d

Observation 891319e3-785e-43f7-aed2-e2e2fc9c80cb · inbound

Schedule-Level Shared-Prefix Reuse for LLM RL Training cites this paper.

Schedule-Level Shared-Prefix Reuse for LLM RL Training SGLang: Efficient Execution of Structured Language Model Programs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:36:15.136976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T16:43:39.796020Z digest=sha256:5293bf35d49bf7a404cee7c9cf1ca23e4d6605e524013a888edee1ffb1a1ce13

Observation de2fdcd6-e2d9-4f62-b903-82666e52ffb9 · inbound

Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads cites this paper.

Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads SGLang: Efficient Execution of Structured Language Model Programs

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:06:24.189836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T12:48:02.761437Z digest=sha256:e4e3e4f804ba647cf70af8f90e928534c74841493d20108b5d82cb10289a285e

Observation 083c2371-c491-4faa-91e4-952d5f754bab · inbound

DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference cites this paper.

DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:40.720499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:47:26.074448Z digest=sha256:47c965f63d82b00602206de216cef888a1ebc43b6392f25090a7205657c96b54

Observation 8c80b046-d9a5-4f30-87cd-91ba8bd118ef · inbound

MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU cites this paper.

MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU SGLang: Efficient Execution of Structured Language Model Programs

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:26:45.835943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:58:31.351774Z digest=sha256:4172d8304dfaeb37b7436896eee60a25072ed3398e52ddf94617c3f1868028df

Observation 56210994-c691-4877-bf4b-b9f6d58aa3b6 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:51:29.010070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:fbed72848e1c961d4925495ef478832f8a0be771767888ac6e2d14e7e0ee9b80

Observation f18438d6-ef0c-4909-8003-5f0d4364358d · inbound

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving cites this paper.

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.102545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:25:46.234576Z digest=sha256:c1f8f74487762e6e9f729e553d2bcb9612e11297f444b0116168f726dd62ecdf

Observation 3fa253e3-bc3e-4ad9-8031-d10c99e06bd7 · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents SGLang: Efficient Execution of Structured Language Model Programs

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:36:59.419341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:fbb715845761bf359709240767a80c65c372bfb4fe865be94755f7521bf11890

Observation ffeb3a49-abcc-4cef-bd56-69957ec9a217 · inbound

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models cites this paper.

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:07:26.529373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T19:10:53.882876Z digest=sha256:94cdc1b00d5f22bb646e7cb838d3030542bae79ee8c2e84aa594719d0f253545

Observation d0146a87-7f23-42cd-a6fb-01535021a337 · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild SGLang: Efficient Execution of Structured Language Model Programs

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:37:25.724728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:227f6a0d37f02699c9a3945a155895625145f9179f3e7bf3600e6f83e0be2ecb

Observation 6918d03b-d60f-4c51-9979-55617cbe470e · inbound

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy cites this paper.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy SGLang: Efficient Execution of Structured Language Model Programs

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T17:11:05.710367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:44c2006f6bac6e87a1621dc38b5266d5a23e665a59b90fc4b071d8cd58727fec

Observation 159460cc-065c-4e64-935f-8cf432cba2df · inbound

RKSC: Reasoning-Aware KV Cache Sharing and Confident Early Exit for Multi-Step LLM Inference cites this paper.

RKSC: Reasoning-Aware KV Cache Sharing and Confident Early Exit for Multi-Step LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:07:26.749247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:29:16.153111Z digest=sha256:d150b7efb302e417f55a6d8963e6dcd26d94817145dba038a1626325c331f001

Observation d76ef9d0-6300-4f96-a88f-2753b53f9754 · inbound

UltraQuant: 4-bit KV Caching for Context-Heavy Agents cites this paper.

UltraQuant: 4-bit KV Caching for Context-Heavy Agents SGLang: Efficient Execution of Structured Language Model Programs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:39:30.380619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T17:49:02.835708Z digest=sha256:579923948142d830a609dc424d329560d3181156b6a6c50131df4fa22ec7eb95

Observation 25fb1a25-d229-4ec0-a606-2fd3588df483 · inbound

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving cites this paper.

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:49:29.926167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:39:47.477338Z digest=sha256:652c641bcf459c3cd31e6fa8093150bd4ce476f1d247d74fa2f2a3d390dce7a1

Observation 9bfadf86-ecfc-4ac3-86b1-237a449e3be3 · inbound

Human-Less LLM Serving: Quantifying the Human Tax on Throughput cites this paper.

Human-Less LLM Serving: Quantifying the Human Tax on Throughput SGLang: Efficient Execution of Structured Language Model Programs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:45:11.809066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T00:44:55.517266Z digest=sha256:4ec48680facda5a75cdef95dfa29b1def2bf93ca1d585731c74af680baf80f9e

Observation 1d5c4e72-1c64-4877-91ba-d608f4bf6484 · inbound

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models cites this paper.

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:49:37.573661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T14:13:13.678245Z digest=sha256:63ea639193f1633c7ca09034fb70fd8b1f9959ba80ff3c27e6b384982f01cacc

Observation 364ee36a-212e-4a44-a757-f59d9f9cfa5e · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing SGLang: Efficient Execution of Structured Language Model Programs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-26T22:10:09.425558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T06:39:04.161585Z digest=sha256:0314a0a4869c79024c43f24a51708851ca7f8b16adb91b4eaedad5e306fe8e0c

Observation fd90b1b6-c4c0-4a32-a506-6bfaa07c1a4f · inbound

FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models cites this paper.

FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T20:03:56.533297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T04:37:08.966401Z digest=sha256:c8ea3221c4c1a54730e2e16dedb9650358562c89c490d7a22f18a1c65037ae94

Observation 2bce52ad-96ec-44e1-ba55-529da006ac69 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator SGLang: Efficient Execution of Structured Language Model Programs

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:54:06.239888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T00:48:19.207465Z digest=sha256:c037ca04f046e4e9b4b32bd4ceba94e747374750ea4e9c670e6ff3af005d59bf

Observation 22ab3e59-a354-4cd7-a124-3bf2c9707e67 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator SGLang: Efficient Execution of Structured Language Model Programs

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:19:02.250445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T23:09:38.092583Z digest=sha256:758fbc360b560e1ea252ba97b1590cfd98bc4f028b836e1c52844b9de6df5037

Observation 60d3431c-d0c7-40eb-86ea-3db7f5baae15 · inbound

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM cites this paper.

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM SGLang: Efficient Execution of Structured Language Model Programs

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.333448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:19:48.530272Z digest=sha256:8998f54eb3eb2d6c8f63ee70048dc54bea943539d0691fb194fdefb66f582790

Observation c30ca5d4-15f2-4a6f-815b-b3bd07ab9313 · inbound

Speculative Pre-Positioning: Decoding Stateful Sessions to the Next Decision Point Off the Critical Path cites this paper.

Speculative Pre-Positioning: Decoding Stateful Sessions to the Next Decision Point Off the Critical Path SGLang: Efficient Execution of Structured Language Model Programs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.336468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:32:54.738720Z digest=sha256:e0707dbbede1be18935b79ba58228f106673887fbe28f5a9df6cf91f2118dd12

Observation 3b1f8f9e-290f-43b5-8d04-e15c9c36155f · inbound

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference cites this paper.

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:35:43.616162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T04:13:53.722933Z digest=sha256:9881f787f4f168d48a1bb6ab1f0ae2fcf07906f7f10dd6f4372eac7f1963187b

Observation 6298436d-31b8-4141-a5da-aa19cdb95452 · inbound

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving cites this paper.

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:56:43.949731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T06:40:53.505889Z digest=sha256:8efea53c2f9b3f37ebee334fd325b12f4f960e7840b04ecb16702ceb40efa87e

Observation 1a3fabcc-50c9-4279-930f-575aa58c5c33 · inbound

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving cites this paper.

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving SGLang: Efficient Execution of Structured Language Model Programs

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T18:58:50.498142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T18:58:44.830766Z digest=sha256:839fa81ea8fe2690e8b1dd4e96e496eabfa7744d123dd38970e6ab2b114ad3f3

Observation 2e9fe7b7-2041-460d-8eb5-2de22be94f75 · inbound

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters cites this paper.

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters SGLang: Efficient Execution of Structured Language Model Programs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:37:42.205666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T06:30:30.308713Z digest=sha256:308269f646abfd10bcb50c9849adc914b3822dc2ef6cc84f71eaec89b5d1b9bf

Observation 2f56283a-de44-4c79-811a-022269d6b8dc · inbound

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems cites this paper.

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems SGLang: Efficient Execution of Structured Language Model Programs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-12T11:05:56.233115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T11:05:56.233115Z digest=sha256:ecadf903920d3993a2d0aa889f0ee5e238f57f6318a691a6831581bc6e0e8ffe

Observation 1cf6d3df-2fdb-4167-8fb6-c62bf9d2021e · inbound

A Workflow-Aware Serving Layer for Agentic Applications cites this paper.

A Workflow-Aware Serving Layer for Agentic Applications SGLang: Efficient Execution of Structured Language Model Programs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T05:56:18.797766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:56:18.797766Z digest=sha256:dd380749098ab8f0f571a12ae7c7cb0bb21b037a13823982f04257f237b0a65f

Observation 303e54f1-e3b4-4172-9751-c76ff02f98a7 · inbound

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling cites this paper.

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling SGLang: Efficient Execution of Structured Language Model Programs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T05:40:56.557646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:40:56.557646Z digest=sha256:d663cd234777d4e72e6421e107452224ca00ec8c1f6604afcadcf0df4bbbf28f