Pith. sign in

Paper Citation Record · LEDGER

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

As of 24 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 12 inbound Pith citation observations for arXiv:2502.01976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01976 v6

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:58:45.830152Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:55:06.664723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:15:32.127896Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved50
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8379fe4b-7862-47dd-b711-cd1edc4956a2 · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.848993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.848993Z digest=sha256:ba5f706541f1e94d7d5c731c3e15442f78c8ec47263d40103182b0e916286725

Observation 69f91f56-f8e5-47a8-8eaf-08f3ddf7ccdf · outbound

This paper cites Concrete Problems in AI Safety.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Concrete Problems in AI Safety

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.855270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.855270Z digest=sha256:b5430a898a0b8322ec9f21fc8cc7afa42c906609d329cd506bdea8e476abf9c9

Observation 0a396dda-4596-493d-97e6-661675bc099b · outbound

This paper cites Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.860751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.860751Z digest=sha256:ee92c28dfe3969c4139cd5d3dd6556d8ee49c0e875c8ce77098c9a6829ba3265

Observation 088bec0a-dbb1-4904-af1e-b4664a4ebd5f · outbound

This paper cites Dynamic Programming.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Dynamic Programming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.865987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.865987Z digest=sha256:af789e7247733e63f237c44ef55b3be3f9483b813b0de3a5e5d6cc9362c59a49

Observation ce7f1e64-d4c0-4928-be17-4bae0ac13bf8 · outbound

This paper cites Token-Level Adaptation of LoRA Adapters for Downstream Task Generalization.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Token-Level Adaptation of LoRA Adapters for Downstream Task Generalization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:58:46.525251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:44.870574Z digest=sha256:c819e7d1b87b3a3fe5148f48f8d791c06fe4d9af90195d20b4efe308b1bf96e1

Observation 2abcd8ac-4e7f-44c2-a606-6e08b53cde8c · outbound

This paper cites Speculative Streaming: Fast LLM Inference without Auxiliary Models.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.876800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.876800Z digest=sha256:9e9f533196993ca949f256cbef03785cc7fd5fa6d7337482299b70c1c557dde7

Observation 9c31b60b-3ed9-4c58-bd32-d3fa7b2ecd32 · outbound

This paper cites Rank analysis of incomplete block designs: I.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Rank analysis of incomplete block designs: I

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.882631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.882631Z digest=sha256:523038a0688b980389aeda12d1ed6753f6e6dab372ec18651fd235c6ba9bcc04

Observation 8d7174b2-2b0b-471c-9fa3-98dedda0ab54 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.887464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.887464Z digest=sha256:3d8bb07260cea42b526ea6726a96da4f8bea70efc894bc36c06977f6a45f5e1a

Observation 431d1a99-ccc9-4edb-8137-c3dec221408c · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Accelerating Large Language Model Decoding with Speculative Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.892064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.892064Z digest=sha256:fb4c4ba43159aa7df35320a84770d187c47b44d76c860f044441e3565afb9c4f

Observation 93f87e93-2712-45d1-90d7-50f2d84fbffb · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.897101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.897101Z digest=sha256:f8ce2e205ad95d2697b91e54faa2eff21bcdc609947a16a7f472997dfa466ad5

Observation c0896544-42fc-4a78-b075-1b2c8e9cd1f7 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.902016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.902016Z digest=sha256:2a6a07e3bfd9f5d568fc1ded1bfc17bfc8ea6176e2bdfed0ffd36abea26b8d97

Observation 71ce0ee4-1c8f-4fda-92fe-3a28b33cd8b2 · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.961855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.961855Z digest=sha256:a83458a387e9df5d58a9d625a7317aa2341ab9fe814900fa76debb7d5906a529

Observation 8019aa24-573e-4083-966d-88d642b1e396 · outbound

This paper cites DAM: Dynamic Adapter Merging for Continual Video QA Learning.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing DAM: Dynamic Adapter Merging for Continual Video QA Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.050713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.050713Z digest=sha256:e096aace503d3b68255b80305779d8624b682abdfc24a4997c93dad221a2bc9f

Observation 8de607fb-8dbd-4dc1-9972-e0cf5bf53464 · outbound

This paper cites AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.122333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.122333Z digest=sha256:9e74b392da8e120c9b415c1cda000b22466f1d4be827d955be2fd91e0c075c8f

Observation c6907a17-7f96-4430-876f-295ed6b36bd1 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.242020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.242020Z digest=sha256:a66024aa514dd019797d705ee1a0dc476cf65821607e1a44897a4624b5367758

Observation 0b127db8-30e6-4dca-b0b9-941e9dd6255a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.288112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.288112Z digest=sha256:5f0ff9c3d1131c0f1ef88fd57df124350d9ae75e190ffb421a3b1f26a1388436

Observation 933ed9bb-e91b-4bc0-8cd4-d2f437108547 · outbound

This paper cites LLM-Assisted Rule Based Machine Translation for Low/No-Resource Languages.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing LLM-Assisted Rule Based Machine Translation for Low/No-Resource Languages

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:58:46.355811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:45.292988Z digest=sha256:c71413436334d0bb93694e8c4a586b805afd9e7aefada53ab8958979c10ad38d

Observation 6fcd52c0-b5e4-4281-be51-af0b81469285 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.298483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.298483Z digest=sha256:2b18d64b84f8401374b281e95a88505c63ce16a51296262f5e2be43a7330cb1b

Observation 084a8d86-641c-4891-a9b3-f1f1c1fa9670 · outbound

This paper cites Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models Memories.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models Memories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.302879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.302879Z digest=sha256:35d7a2f94061d68f38912022e06564b3a1f1a3e5ddfb36245cad58ba975da8f3

Observation 5a732764-f1be-4d51-913d-1c3427b92698 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.307926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.307926Z digest=sha256:8839bbc855daa7154df4c0b677f0f1e0b2c572d2e4b5f62882cf0f9db1d3880f

Observation 268fd119-6753-4045-97e3-b45d02615cc3 · outbound

This paper cites Towards Translating Real-World Code with LLMs: A Study of Translating to Rust.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Towards Translating Real-World Code with LLMs: A Study of Translating to Rust

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.313458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.313458Z digest=sha256:86b3ca0b31752ec7d21d0f62bd7fa27e748b751e667a8c40d8f2cccdccd150e9

Observation 251bd8a1-2454-4603-9e1c-7b738f151477 · outbound

This paper cites Language Model Cascades: Token-level uncertainty and beyond.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Language Model Cascades: Token-level uncertainty and beyond

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.318248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.318248Z digest=sha256:a6c63e07cfd8a82630d7af4bebb47f1958090185e4c13a01656d18493cc6b7d9

Observation 6c6e4cbb-595a-4e1a-99e5-a21ff9d221d5 · outbound

This paper cites Measuring massive multitask language understanding.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Measuring massive multitask language understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.323235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.323235Z digest=sha256:7fc46c5becca721fc713339aa92e142e31a45dd7722f7d02468105024ff33efc

Observation 83d4acb4-6397-4317-935f-71b9e8a001fe · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Measuring mathematical problem solving with the math dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.328358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.328358Z digest=sha256:1af0571783b18fe99b60e07afd55693a2b6076aeb3720ed3b1fa98aa2d5c6315

Observation bc7afff9-9f17-4ce5-8f4f-4f0d4d17b645 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.332991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.332991Z digest=sha256:742d4179bd1e8739caaa90c2c667001c81ce301a9099901dd6fcd910cce97fa5

Observation 67663a97-86b0-4113-91bd-41a2b4774bb6 · outbound

This paper cites Exploring the benefits of training expert language models over instruction tuning.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Exploring the benefits of training expert language models over instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:58:46.702470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:45.338991Z digest=sha256:eea7e9e95b0bd0e6bbdaa7013e8d0bd31bef7062e656ee397ac9796f9fc987e5

Observation 6f9dc153-5b08-4829-bd01-cc8b6e1b1c31 · outbound

This paper cites Towards robust qa evaluation via open llms.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Towards robust qa evaluation via open llms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.344008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.344008Z digest=sha256:c3227f2cec2de815f0e3ffcbb4149c06606f56aacc59829dfa9c3afe69bde322

Observation 78f549a1-7534-4f4a-9de1-3f9d8849a1b0 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.348995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.348995Z digest=sha256:89dc70c51b94ee4b298c0c17b0a0099d89fe2dca1419c9173f35e5711277386c

Observation f68d61f4-0031-4b62-911f-332547c73ac5 · outbound

This paper cites CLLMs: Consistency Large Language Models.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing CLLMs: Consistency Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.354104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.354104Z digest=sha256:e46a6b200167894c24e4826acc0a717067f3969f67c807aac9d69c0dff56b624

Observation 6b20d33c-8138-4f41-8b31-1f98aaf39ffe · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Efficient memory management for large language model serving with pagedattention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.359067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.359067Z digest=sha256:1b8d36289e4b88fd9a09212689b90bbbfa81b1d7121e9b7712a01e0fb99ca017

Observation db58ce78-8ba4-40e2-82d2-db70552a5e99 · outbound

This paper cites Specification gaming examples in ai.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Specification gaming examples in ai

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:58:46.678920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:45.363913Z digest=sha256:3074e1d8902ca43c1cc9ca326fe55ad90e83c54baac021011458970abe2d754b

Observation b5cdf383-a428-40ca-b2ba-972dc08263e4 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Fast Inference from Transformers via Speculative Decoding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.369148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.369148Z digest=sha256:5d9d4ae2739646a1f99aa419629ace2b584dc664d3ad0cff657431da75802798

Observation 16ca6e8b-007c-4e4d-8b93-4a380184e658 · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.374081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.374081Z digest=sha256:fb109f4306a2ebc0e3f62b543600a2b7d3c4978e2509ea54d4045142bbb0c877

Observation 34f7539f-8d74-40ba-9805-3df51fc7522f · outbound

This paper cites Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.555563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.555563Z digest=sha256:76c31898e4b5fa63e53a0f0cea4047756d2ae2623d03fa369aa174654639eb66

Observation e1eb9bdb-54a8-49ff-b8d0-3a03a9fd1420 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.603607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.603607Z digest=sha256:04f273b134d58b6f9953d7873a044f5d18a2ee0044721237b36a847f7bc2787f

Observation 6fd29c9b-8f2f-4158-aa4a-7ea2147c4e6a · outbound

This paper cites Routoo: Learning to Route to Large Language Models Effectively.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Routoo: Learning to Route to Large Language Models Effectively

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.656575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.656575Z digest=sha256:8b1a7bc67b6dd64f28b98f70aa7943635df2ca3905c7dacdfa44b1a525ae7765

Observation a5845137-7153-458b-99e5-0ff272ba5d02 · outbound

This paper cites Learning to Route Among Specialized Experts for Zero-Shot Generalization.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Learning to Route Among Specialized Experts for Zero-Shot Generalization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.664477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.664477Z digest=sha256:f30a820b544c5a9769f4c20338b43cec9ac19926d3209b02dfac7eeb4b4a9712

Observation eff3495b-4e86-4a6f-b395-2f9207bb1cdb · outbound

This paper cites Faster Cascades via Speculative Decoding.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Faster Cascades via Speculative Decoding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.704596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.704596Z digest=sha256:1d091ef7be7e916a414cba1c0ef93798a432f9fe33a5bd427f2e490e6ef5ebd3

Observation 6ede4813-0d4d-498b-8736-3838cf9652a1 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing RouteLLM: Learning to Route LLMs with Preference Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.709420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.709420Z digest=sha256:7d0684efd2f27d65f0d73ed6b1ba617da6c90ab281d2cae26dd6406db83fac95

Observation a652f26c-b1cb-4b0b-a80b-24eb071c407c · outbound

This paper cites Towards Modular LLMs by Building and Reusing a Library of LoRAs.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Towards Modular LLMs by Building and Reusing a Library of LoRAs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.714214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.714214Z digest=sha256:f41a8bb087e7efd87d575be044e4c41d1c9861bc00c72cad8f13f7a6e2093ac7

Observation 91e2d0d6-0888-4dc1-8944-f3e4711f11ab · outbound

This paper cites A dapter F usion: Non-destructive task composition for transfer learning.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing A dapter F usion: Non-destructive task composition for transfer learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.718700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.718700Z digest=sha256:103eafd17e57ac298f24928e56061f0b0004becb5818cb89fb6d49ef10179c9e

Observation 60dde419-03f6-4f5a-b565-cf80433439f4 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:58:46.665206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:45.724005Z digest=sha256:a39f74025ab2368cc9cb0e9151f1a75d97702d03d21fece352e1944361983fae

Observation 29be8268-b0ff-4e74-a651-8b67835586fb · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Direct preference optimization: Your language model is secretly a reward model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.728917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.728917Z digest=sha256:b0e9e89303f1e2f6951d2dedc454b733971895f97904cee70b4f1afbe0522dc6

Observation e5e9ff71-4472-4d56-abc8-3b4b60fee956 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.734215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.734215Z digest=sha256:d907e3d32d374d2d95c291a863d88f03f052700a9ba58268525ecc6f1031d349

Observation 8a3b78a4-639f-4299-8619-f8b53334dda9 · outbound

This paper cites Learning to Decode Collaboratively with Multiple Language Models.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Learning to Decode Collaboratively with Multiple Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.740288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.740288Z digest=sha256:0765252eb86cb0fef479b8043682c21874840e55710c7d9539238bd332e78693

Observation f63a1616-dc9e-4c72-b71c-05bacd70de32 · outbound

This paper cites Harnessing the power of multiple minds: Lessons learned from LLM routing.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Harnessing the power of multiple minds: Lessons learned from LLM routing

Reference 46

Resolution
verified exact
doi, observed 2026-08-09T13:58:45.880940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:45.745179Z digest=sha256:c99d505e2b80e82e5013536d20539320565a10f64469bd40abe8aa2d67989bf2

Observation 4a5228ab-35f8-4d50-bfd2-2884eef4c08a · outbound

This paper cites T ensor O pera router: A multi-model router for efficient LLM inference.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing T ensor O pera router: A multi-model router for efficient LLM inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.776889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.776889Z digest=sha256:0d4f4b13b16ee56d711f3bf00439ee63860e1c172e45f86f335b4a297bc9085b

Observation 38b0b922-42ef-4aaf-a09b-a1bc4743654a · outbound

This paper cites Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.781449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.781449Z digest=sha256:af3321e912683039ce4daad7ec0276108eb1e7b33a59ca5037f9fcf84d46630a

Observation 849b2f77-0e41-4ccc-a3f6-3846f87fa424 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-09T13:58:45.786463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.786463Z digest=sha256:a9b226544d7d2ebfb884bcc40a8fa66a40897e5d6cf39eb6095b8e82de5fbcd2

Observation b4253d66-dd04-4325-8f19-6a0551226d68 · outbound

This paper cites LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.791531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.791531Z digest=sha256:ee3bb3a431b02d37cae8b514f561546eb6c503c2fa3a2102a66907f5130ba924

Observation e054da66-5fa1-43b3-8adf-8f5d4c7ba0aa · outbound

This paper cites Fusing models with complementary expertise.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Fusing models with complementary expertise

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:58:46.631980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-09T13:58:45.796689Z digest=sha256:2a656d30159e42398a26680d97e2b9dc2086c21762fe40b6bca7a4fbe297e58b

Observation 0cd43e0b-5f51-4d0d-819e-e96c6e051865 · outbound

This paper cites an unresolved cited work.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.801184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.801184Z digest=sha256:b1c3f36bacedf2ffaca068a518093020f9f963e1edc386b6a686791a10d9430c

Observation 66a4c6fc-11a5-4240-b16a-ab4c775e91a0 · outbound

This paper cites Mixture of LoRA Experts.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Mixture of LoRA Experts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.805889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.805889Z digest=sha256:eee9111374ad20eba3ae640f7735507e48e0adf72919a66ff3c42dfc5a3cf02d

Observation 0a3195f7-70d5-4e33-9d87-26a316baf7dd · outbound

This paper cites MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.810626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.810626Z digest=sha256:411c92773afeb9526ba94ceb04b5ad9e772a45a76383bbc92680c3a708506667

Observation ea3c5454-cfb1-484f-8c08-2a9ae584c9c0 · outbound

This paper cites write newline.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.815498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.815498Z digest=sha256:28fbcc4a7509736ef07a3f22e643284611c39affbd1a9d1ca5df141e1539bcea

Observation 5f516930-4983-4d08-84f5-a24963237040 · outbound

This paper cites @esa (Ref.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing @esa (Ref

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.821054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.821054Z digest=sha256:a6e5c00f526976127f56912b3e0bb82ec9318ebeb908afaa34990dab489f1216

Observation e90dde09-24b8-43bb-9035-a990f82d3806 · outbound

This paper cites an unresolved cited work.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.825862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.825862Z digest=sha256:c1ddedf042a9f92eeaa773b0719e7e1bb284aa0cf6b2fa77dc45aa37445b487b

Observation 497cb05e-257f-4978-97bc-5d02ce1ef09e · outbound

This paper cites an unresolved cited work.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:45.830152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:45.830152Z digest=sha256:0c1b83b74c5ff5a75a906f6f7ce4e91fe210f0997a48b265477e7ceb117c5376

Pith citing papers

Observation d945a911-9fa0-43ee-bd74-a244d66d22fa · inbound

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble cites this paper.

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.669400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T02:22:28.649071Z digest=sha256:3d771cb245f9dd55a3de8fd6c88ecfe88ffea306122a490e324a55e47b6b7c63

Observation 88a66c46-e69a-4a11-9d25-92d03cfd6088 · inbound

Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission cites this paper.

Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:55:06.664723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:55:06.664723Z digest=sha256:f9e93a7f69d9ab238e4701a491484c5f81287c84090e1e2d12cf84cbba01b73c

Observation 85ee61f5-4280-4ede-bbbe-f0fdf462b4c1 · inbound

RLAE: Reinforcement Learning-Assisted Ensemble for LLMs cites this paper.

RLAE: Reinforcement Learning-Assisted Ensemble for LLMs CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:19.253109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:08:19.253109Z digest=sha256:1e39130f12dc4d87c8322e1bc5639366bffecae21cef17ba24db64dc3d18947a

Observation bb73bfbc-065e-45c1-8575-0053ad623da8 · inbound

Sampling from Your Language Model One Byte at a Time cites this paper.

Sampling from Your Language Model One Byte at a Time CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:53:02.078562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T09:52:20.923710Z digest=sha256:3df3988c36f6fbd42fbabc0396206eafdef10ba361d638ead5ee3ea5c18a4adb

Observation 127465ed-f961-4ed3-8f79-c835b9b7616f · inbound

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges cites this paper.

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 183

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:48.536427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:48.536427Z digest=sha256:bcbbe7bdbd4e8cb29a8479012820efbd4acd0f491e91e06368f4a40fc4560b7b

Observation 2df8c9ca-e25c-4918-8ad8-efbfd166a930 · inbound

NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization cites this paper.

NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:01:13.633017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:56:52.205196Z digest=sha256:dd1fac6b5a45a40f2a29f9a27ec86bbf9862c8345bd4844abee7d97f42d5fb08

Observation 9482e410-3884-4ffd-b979-89f5c33f1c2a · inbound

Rethinking LLM Ensembling from the Perspective of Mixture Models cites this paper.

Rethinking LLM Ensembling from the Perspective of Mixture Models CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:06.339051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T20:06:12.248439Z digest=sha256:f1dd05600e1c3cc93b4a18bad0498828eda506f8c7dea32b7e51a8744c80ccfd

Observation 1e596ca0-c9d6-4902-8889-de3a477d5455 · inbound

Rethinking LLM Ensembling from the Perspective of Mixture Models cites this paper.

Rethinking LLM Ensembling from the Perspective of Mixture Models CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:32.130199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T08:07:09.032028Z digest=sha256:289279670f9b5838a73d6b63704214557745b2e44722bc8b6f31685b11e98939

Observation 28555106-1416-4f58-b7f3-19a2b0208d0c · inbound

Accelerating Heterogeneous Agent Collaboration in Dynamic Edge Networks cites this paper.

Accelerating Heterogeneous Agent Collaboration in Dynamic Edge Networks CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T15:18:54.607262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:18:54.607262Z digest=sha256:a9ca085752a14ccd025e0c5602b5b5b31964c7e561fa3657ce3a69cfd8fe8e1f

Observation 9791e570-20ef-4b91-9855-395ec8fb6379 · inbound

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference cites this paper.

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T10:14:15.660311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:14:15.660311Z digest=sha256:50a2c93dc546a189d766a469c652aace5db99921163c889d6e0df923ef67a9c3

Observation 07fd0cd4-105c-4dbf-bb4d-824d58d5dc52 · inbound

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI cites this paper.

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:42:15.122849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:42:15.122849Z digest=sha256:1d727d0087675f7a6f90486439ec092a2ceb33590115ad490ebdf24e87c5462c

Observation c2049d7d-e5e1-44e9-9f02-3fe40069efe6 · inbound

Divergence Decoding: Training-Free Capability Fusion cites this paper.

Divergence Decoding: Training-Free Capability Fusion CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:47:11.302112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:47:11.302112Z digest=sha256:dfcf242efb0970bfdf100791803319eafa352ceb032115ac7e4decb524b50db8