Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:43:32.575076Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 4 inbound Pith citation observations for arXiv:2512.04013.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:43:32.575076Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:46.161659Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T00:35:10.221606Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 67c303cd-ba22-4a7b-86ad-464690d13be6 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Infercept: efficient intercept support for augmented large language model inference
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7a8620-12a5-494a-a017-a9093cbe60bc · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd4a6a13-d068-4778-b4e5-fc8fa69a3ca3 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Revisiting service level objectives and system level metrics in large language model serving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9316851-8943-49d3-9131-6e569b986d96 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Model context protocol (mcp)
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f111267d-5bb3-4e45-a2ed-af70dde5d220 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving From good to great: Improving math reasoning with tool-augmented interleaf prompting
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff33d1e1-b9f0-4e2a-a8d3-751b46dbf341 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Advancing tool-augmented large language models: Inte- grating insights from errors in inference trees.Advances in Neural Information Processing Systems, 37:106555– 106581, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b363ce5-31f8-477c-b737-9ecb9df569cc · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f543e4a-108e-424c-a310-63199f4d8aa5 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625– 630, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0bfac4b-a9e2-4e46-b71d-eaf42f765e6d · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50c05bb9-0fe4-4866-bb0b-8cde0a358e20 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Efficient llm scheduling by learning to rank.Advances in Neural Information Processing Systems, 37:59006–59029, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066f641a-bbf7-4a02-a9f5-7308e1fa448d · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Jetcheva, and Hardi Trivedi
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7557ef2-f386-4d82-86f8-bcdf36f92130 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Apt-serve: Adaptive request scheduling on hybrid cache for scalable llm inference serving.Proceedings of the ACM on Management of Data, 3(3):1–28, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd7e3ae0-b221-4779-996c-64e3fa541e64 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Asynchronous LLM Function Calling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f327b7-3238-4f4f-8c35-82e0b97cb622 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving A study on classi- fication based concurrent api calls and optimal model combination for tool augmented llms for ai agent.Sci- entific Reports, 15(1):20579, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf67f45b-3f9e-4fbf-8534-75b8eedb9a7c · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems, 36:45870–45894, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59bc1e72-68c0-48f4-a4f0-72f2ad15655a · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Shuffleinfer: Disaggregate llm inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02c14a8-c891-432f-9f95-b01ee43d2528 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Tightllm: Maximizing throughput for llm inference via adaptive offloading policy.IEEE Transactions on Computers, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24de9e02-25b6-469b-82f7-bc4c0ccd6b7b · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Accelerating llm serving for multi-turn dialogues with efficient resource management
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d071d6-57ed-4f8f-8dd5-3ce9ab223eb4 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving S3: increasing gpu utilization during generative inference for higher throughput
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8832e6e-db31-4eb7-8d6e-95543d8339ec · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Optimizing goodput through sharing for batch analytics with deadlines
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 314f4266-187f-4d0c-a260-22b32d2659c0 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Efficient memory manage- ment for large language model serving with pagedatten- tion
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc265919-a298-41fc-9975-27457dc032d4 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17fea0fe-b551-4dd8-a24a-e3a6176f7ac5 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Augmented Language Models: a Survey
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2596146-f0f2-4e60-97d9-32de7bad0821 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Introducing function calling in chatgpt
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63ab32b7-83fa-424d-880e-4ae40400bbf6 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Splitwise: Efficient generative llm inference using phase splitting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56e7b59-8ed3-4411-bdbd-5068a2497485 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving WebRL: Training LLM web agents via self-evolving online curriculum reinforcement learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be7f0d62-8ea7-4f7c-b5be-da4706076009 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Tool learning with foun- dation models.ACM Computing Surveys, 57(4):1–40, 2024
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a85e1ebb-75a7-461b-9321-65bb6eb50b70 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving ToolLLM: Facilitating large lan- guage models to master 16000+ real-world APIs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b06a4ee4-7246-419b-afb6-62c7d28dbc26 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Tool- former: Language models can teach themselves to use tools.Advances in Neural Information Processing Sys- tems, 36:68539–68551, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6650787b-9992-4f2d-9483-c76ed4992d1f · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving DON’t STOP ME NOW: EMBEDDING BASED SCHEDULING FOR LLMS
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4333ffe-36c9-4718-8fcf-4bca5319d1ed · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Fast Inference for Augmented Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11fd6a3-c04c-497f-8dd2-a1f9f578d4e2 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57485fd-c763-4d24-aca8-c80aab7ebf60 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aa51e48-48a6-48e3-bb39-5646e1755402 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Wang and A
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82bd4dbd-0214-47f3-b765-0934590afe9b · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Fast Distributed Inference Serving for Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656bd126-cf33-4f5a-968f-a9d99d15cf4a · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4ca4a5-0cef-42e8-86a2-46bcad201e81 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Orca: A distributed serving system for {Transformer-Based} generative models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21731362-e518-4e15-8f5b-d16505317704 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving {SHEPHERD}: Serving {DNNs} in the wild
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4d822e-3346-4f3d-8fa1-a1fc2168bb5d · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving OPT: Open Pre-trained Transformer Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3093174-0517-4fa4-b929-d3879215e197 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Webpilot: a versatile and au- tonomous multi-agent system for web task execution with strategic exploration
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9b5d6f-6bb7-494e-afd6-e67a180a207f · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Response length perception and sequence scheduling: an llm-empowered llm infer- ence pipeline
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fa7e3a-34fc-44d3-8ed9-f0a0e70fc139 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8dfccfb-94c0-42b7-b120-56c57d07f7f7 · outbound
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d157a2be-c120-415b-95a0-a02d4a444cc1 · inbound
Efficient Multi-round LLM Inference over Disaggregated Serving AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d5cf38-73a0-406b-97d0-83e3602788ed · inbound
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f48c869-091a-4c5a-8dbd-9b2d05ec1801 · inbound
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ba4fdb8-0140-4635-b7dc-25b138f82528 · inbound
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.