Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T13:58:45.830152Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 12 inbound Pith citation observations for arXiv:2502.01976.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T13:58:45.830152Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:55:06.664723Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T08:15:32.127896Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8379fe4b-7862-47dd-b711-cd1edc4956a2 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Deep Learning using Rectified Linear Units (ReLU)
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f91f56-f8e5-47a8-8eaf-08f3ddf7ccdf · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Concrete Problems in AI Safety
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a396dda-4596-493d-97e6-661675bc099b · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088bec0a-dbb1-4904-af1e-b4664a4ebd5f · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Dynamic Programming
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7f1e64-d4c0-4928-be17-4bae0ac13bf8 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Token-Level Adaptation of LoRA Adapters for Downstream Task Generalization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2abcd8ac-4e7f-44c2-a606-6e08b53cde8c · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Speculative Streaming: Fast LLM Inference without Auxiliary Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c31b60b-3ed9-4c58-bd32-d3fa7b2ecd32 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Rank analysis of incomplete block designs: I
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d7174b2-2b0b-471c-9fa3-98dedda0ab54 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 431d1a99-ccc9-4edb-8137-c3dec221408c · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Accelerating Large Language Model Decoding with Speculative Sampling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f87e93-2712-45d1-90d7-50f2d84fbffb · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0896544-42fc-4a78-b075-1b2c8e9cd1f7 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ce0ee4-1c8f-4fda-92fe-3a28b33cd8b2 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8019aa24-573e-4083-966d-88d642b1e396 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing DAM: Dynamic Adapter Merging for Continual Video QA Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de607fb-8dbd-4dc1-9972-e0cf5bf53464 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6907a17-7f96-4430-876f-295ed6b36bd1 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b127db8-30e6-4dca-b0b9-941e9dd6255a · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Training Verifiers to Solve Math Word Problems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933ed9bb-e91b-4bc0-8cd4-d2f437108547 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing LLM-Assisted Rule Based Machine Translation for Low/No-Resource Languages
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6fcd52c0-b5e4-4281-be51-af0b81469285 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084a8d86-641c-4891-a9b3-f1f1c1fa9670 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models Memories
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a732764-f1be-4d51-913d-1c3427b92698 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268fd119-6753-4045-97e3-b45d02615cc3 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Towards Translating Real-World Code with LLMs: A Study of Translating to Rust
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251bd8a1-2454-4603-9e1c-7b738f151477 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Language Model Cascades: Token-level uncertainty and beyond
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6e4cbb-595a-4e1a-99e5-a21ff9d221d5 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Measuring massive multitask language understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d4acb4-6397-4317-935f-71b9e8a001fe · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Measuring mathematical problem solving with the math dataset
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7afff9-9f17-4ce5-8f4f-4f0d4d17b645 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67663a97-86b0-4113-91bd-41a2b4774bb6 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Exploring the benefits of training expert language models over instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6f9dc153-5b08-4829-bd01-cc8b6e1b1c31 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Towards robust qa evaluation via open llms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f549a1-7534-4f4a-9de1-3f9d8849a1b0 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Adam: A Method for Stochastic Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68d61f4-0031-4b62-911f-332547c73ac5 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing CLLMs: Consistency Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b20d33c-8138-4f41-8b31-1f98aaf39ffe · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Efficient memory management for large language model serving with pagedattention
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db58ce78-8ba4-40e2-82d2-db70552a5e99 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Specification gaming examples in ai
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b5cdf383-a428-40ca-b2ba-972dc08263e4 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Fast Inference from Transformers via Speculative Decoding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ca6e8b-007c-4e4d-8b93-4a380184e658 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f7539f-8d74-40ba-9805-3df51fc7522f · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1eb9bdb-54a8-49ff-b8d0-3a03a9fd1420 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd29c9b-8f2f-4158-aa4a-7ea2147c4e6a · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Routoo: Learning to Route to Large Language Models Effectively
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5845137-7153-458b-99e5-0ff272ba5d02 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Learning to Route Among Specialized Experts for Zero-Shot Generalization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff3495b-4e86-4a6f-b395-2f9207bb1cdb · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Faster Cascades via Speculative Decoding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ede4813-0d4d-498b-8736-3838cf9652a1 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing RouteLLM: Learning to Route LLMs with Preference Data
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a652f26c-b1cb-4b0b-a80b-24eb071c407c · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Towards Modular LLMs by Building and Reusing a Library of LoRAs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e2d0d6-0888-4dc1-8944-f3e4711f11ab · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing A dapter F usion: Non-destructive task composition for transfer learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60dde419-03f6-4f5a-b565-cf80433439f4 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Direct preference optimization: Your language model is secretly a reward model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 29be8268-b0ff-4e74-a651-8b67835586fb · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Direct preference optimization: Your language model is secretly a reward model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e9ff71-4472-4d56-abc8-3b4b60fee956 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3b78a4-639f-4299-8619-f8b53334dda9 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Learning to Decode Collaboratively with Multiple Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63a1616-dc9e-4c72-b71c-05bacd70de32 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Harnessing the power of multiple minds: Lessons learned from LLM routing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a5228ab-35f8-4d50-bfd2-2884eef4c08a · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing T ensor O pera router: A multi-model router for efficient LLM inference
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b0b922-42ef-4aaf-a09b-a1bc4743654a · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849b2f77-0e41-4ccc-a3f6-3846f87fa424 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing C ommonsense QA : A question answering challenge targeting commonsense knowledge
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4253d66-dd04-4325-8f19-6a0551226d68 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e054da66-5fa1-43b3-8adf-8f5d4c7ba0aa · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Fusing models with complementary expertise
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0cd43e0b-5f51-4d0d-819e-e96c6e051865 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a4c6fc-11a5-4240-b16a-ab4c775e91a0 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Mixture of LoRA Experts
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a3195f7-70d5-4e33-9d87-26a316baf7dd · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea3c5454-cfb1-484f-8c08-2a9ae584c9c0 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing write newline
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f516930-4983-4d08-84f5-a24963237040 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing @esa (Ref
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90dde09-24b8-43bb-9035-a990f82d3806 · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497cb05e-257f-4978-97bc-5d02ce1ef09e · outbound
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d945a911-9fa0-43ee-bd74-a244d66d22fa · inbound
Harnessing Multiple Large Language Models: A Survey on LLM Ensemble CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 88a66c46-e69a-4a11-9d25-92d03cfd6088 · inbound
Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ee61f5-4280-4ede-bbbe-f0fdf462b4c1 · inbound
RLAE: Reinforcement Learning-Assisted Ensemble for LLMs CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb73bfbc-065e-45c1-8575-0053ad623da8 · inbound
Sampling from Your Language Model One Byte at a Time CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 127465ed-f961-4ed3-8f79-c835b9b7616f · inbound
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 183
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df8c9ca-e25c-4918-8ad8-efbfd166a930 · inbound
NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9482e410-3884-4ffd-b979-89f5c33f1c2a · inbound
Rethinking LLM Ensembling from the Perspective of Mixture Models CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1e596ca0-c9d6-4902-8889-de3a477d5455 · inbound
Rethinking LLM Ensembling from the Perspective of Mixture Models CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 28555106-1416-4f58-b7f3-19a2b0208d0c · inbound
Accelerating Heterogeneous Agent Collaboration in Dynamic Edge Networks CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9791e570-20ef-4b91-9855-395ec8fb6379 · inbound
PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07fd0cd4-105c-4dbf-bb4d-824d58d5dc52 · inbound
TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2049d7d-e5e1-44e9-9f02-3fe40069efe6 · inbound
Divergence Decoding: Training-Free Capability Fusion CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.