Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

As of 21 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 6 inbound Pith citation observations for arXiv:2506.06579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06579 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:56:57.070193Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:10:53.372941Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 18d0f5d8-50b3-4731-9b87-671596bc7b41 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.676293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.676293Z digest=sha256:e45ab81f5eb2be060562b17b548bfe3e72825e111d7f4545863460149942ff0a

Observation 81b4d2ff-f2a3-422a-9949-b79583aae3c4 · outbound

This paper cites Gpt-3: What is it good for,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Gpt-3: What is it good for,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.681421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.681421Z digest=sha256:3d01386bbd9583567150a70e15609125a086334dedf67bb6a36c560da6695eca

Observation 05ee312a-aeb1-4c58-afe7-62044f8759ee · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.685256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.685256Z digest=sha256:2b6dac7744abf75faa3c0f56d67b9f5f7e46424313a4728b7e162a862d0ada86

Observation b3796b1a-f0bc-4e92-a908-bdbe0898d0e7 · outbound

This paper cites Mobile Edge Intelligence for Large Language Models: A Contemporary Survey.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Mobile Edge Intelligence for Large Language Models: A Contemporary Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.689467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.689467Z digest=sha256:e51fdaae19e6dc2f6afee04f3a671b1d592d0eae06760f7783c9e77f66ab5b87

Observation 46ff31e9-dd1b-46b2-96a3-52481f9c4c92 · outbound

This paper cites Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.693866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.693866Z digest=sha256:1c28a2b8770fcfed93e3556e0064f0eb01b279ffd27d8903b2cbd8c2a52f4809

Observation 33ce2435-327e-49b8-953f-09acee7540f3 · outbound

This paper cites A Survey of Resource-efficient LLM and Multimodal Foundation Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey of Resource-efficient LLM and Multimodal Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.698136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.698136Z digest=sha256:a9a02bac335f92cf37f334c60afd8c7d9f3913b6f6bb6b377a8eccd18e4b171c

Observation 9fab6bad-1562-4300-b2ef-c1200ed5900d · outbound

This paper cites Efficient compressing and tuning methods for large language models: A systematic literature review,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Efficient compressing and tuning methods for large language models: A systematic literature review,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.703061Z digest=sha256:ac49b866af13466179b352d13a03e820aced73734352885163a2efee36b48a2f

Observation 61647dc1-bbbe-496c-bfb1-6c23b4989b34 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.706799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.706799Z digest=sha256:9444061fe3234f9731745d42d6949fd0b3db98011b375b70d2a89550fbd46046

Observation 3e1f6fc4-ce73-4bd6-988b-01bc1c177bbd · outbound

This paper cites Evaluation of the phi-3-mini slm for identification of texts related to medicine, health, and sports injuries,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Evaluation of the phi-3-mini slm for identification of texts related to medicine, health, and sports injuries,

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:56:57.788517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.710904Z digest=sha256:eeaa54c5819b2feca26d13d4f53fcf4a9252fa50cc7f54636af8676d9908dd8d

Observation e5184609-c34b-4f89-b9d8-454290dbc810 · outbound

This paper cites Harnessing moderate- sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing moderate- sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.714951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.714951Z digest=sha256:e846f441fa3be49dcff07df625944c257d27869d9afb9f140051be745da00e75

Observation ef61aed4-6660-40d4-a88f-3ed439004692 · outbound

This paper cites Overview of small language models in practice,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Overview of small language models in practice,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.718751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.718751Z digest=sha256:ba0a8a1b343a3c0182e8f85f3fe1dad914f18df36cc088e6328720e292629fc8

Observation 25ca84e2-fe07-4ddb-9284-b90423a8e8de · outbound

This paper cites MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.722420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.722420Z digest=sha256:645b23566e2a845f5175ba8585c8b1329009121cb779c5c96f55e3290b33982b

Observation 5cf3d3c2-0771-478c-a869-9735cf428956 · outbound

This paper cites EcoAssistant: Using LLM Assistant More Affordably and Accurately.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques EcoAssistant: Using LLM Assistant More Affordably and Accurately

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.726464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.726464Z digest=sha256:b9a31728d1a8d219e6c84a7d2ebedb645b969c40335f94701a029a95c6bbfec2

Observation ebb02bfe-3add-451a-adfe-b3185ef308e7 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.730282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.730282Z digest=sha256:b46fe34616a40653bc1e400b7274f6e24043eebc836dda472f2cc42ac847ed82

Observation 1adff952-db0e-4c63-9f30-5560021ae50e · outbound

This paper cites Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.734663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.734663Z digest=sha256:4fb564b2437745af5dc3f5fac979471d44dfb8178ac3e995258b0083990d016a

Observation c53fac9d-a16e-49b1-9c65-099db0fc7b9b · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.738323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.738323Z digest=sha256:afa0c9926b15981caf2b1c84a0eb318a33e958c2f967c4591b0c4f5531072283

Observation e13a777d-61db-46c8-bd3e-3388bac034aa · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.742261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.742261Z digest=sha256:29266b8578d45b6edc0619f6b4659aa8cac02bb2bf8b009154cccca360ea43c7

Observation 109d4bb6-2963-4afc-a98d-8f293c457236 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouterBench: A Benchmark for Multi-LLM Routing System

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.746355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.746355Z digest=sha256:3f7f53efd876a5b4c2be55eedd92c6d968e677f8a6e752339bb65df975a47453

Observation e14414f2-d12c-4f9d-85c1-466be764662a · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Efficient Inference for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.750392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.750392Z digest=sha256:da931bcb58a963e5627092a88a31b71fde1c4a552bce667f46c7203a3c7723f5

Observation 5e8ed8fa-1544-40d7-b95d-3d65bd8a07dd · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.754256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.754256Z digest=sha256:0ef5b353a062920b8fd23622685e27b6b8cb97f3824c2e4615ee02e6cebafd54

Observation 0c823151-99af-4d27-9d87-0b6c60cb9593 · outbound

This paper cites Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.758394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.758394Z digest=sha256:f6c183919f76a6fc54e505039518d3af94ef851ae8bffca96cc3fd158fb7ccc7

Observation 23b5ca1b-7de2-48db-b85d-84b7783cced2 · outbound

This paper cites LLM Inference Serving: Survey of Recent Advances and Opportunities.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Serving: Survey of Recent Advances and Opportunities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.762293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.762293Z digest=sha256:94e9a201140bd120194111df874c5e43fe7a1ef64f8cd9728134817bb5041207

Observation 8a93df01-9df9-49f3-a6fd-8babc8be4a46 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.766216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.766216Z digest=sha256:575fd90408b645ad9f750147cda7283909167082db721ab7fdca6abe32b734a5

Observation 1d8ad172-a5ce-46a3-ad55-399e14c4ea5e · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.770100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.770100Z digest=sha256:47512745ebc93a0bf923efc1bdab4b03166c20e0a785f3bd9f9244a9ac5ea6ee

Observation 5a9e363e-7f67-4f63-9e44-a51864a34603 · outbound

This paper cites The Efficiency Spectrum of Large Language Models: An Algorithmic Survey.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques The Efficiency Spectrum of Large Language Models: An Algorithmic Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.774041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.774041Z digest=sha256:440ae4b040d3ed63a69f2599165b9bfae21533418ea15056aff4ff621c381f1b

Observation eb037d63-9b43-4127-a192-d9fc29db83ea · outbound

This paper cites Harnessing Multiple Large Language Models: A Survey on LLM Ensemble.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.778198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.778198Z digest=sha256:f325a0592b22d0cd581b8521c4f0cfc3c5b2305437e1d0704d924599a7e9aebe

Observation 6039e81e-d05b-4f0a-ae3f-a52ab2991376 · outbound

This paper cites A Survey on Collaborative Mechanisms Between Large and Small Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Collaborative Mechanisms Between Large and Small Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.786194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.786194Z digest=sha256:a17cc3cd26c7d499be4fe4eacedf4bed30ba12664483f86af490189726e31f11

Observation 5de78550-fd1b-4748-9cb0-1c7e64efeb56 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Mixture of Experts in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.789923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.789923Z digest=sha256:cefc22f2f6a0475bcda8feb0bf8ea2ccec2dc80e6d1890d8427c87ddbfe96b50

Observation 3577678f-29a4-4934-b91a-f0345d91bd84 · outbound

This paper cites Attention is all you need,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Attention is all you need,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.793800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.793800Z digest=sha256:c21f4a4df5c888dc2795299e921ca2418bf5837f7816544cc4acffd5fbc5808b

Observation 0bbbec51-4dd0-4c8b-a035-89a8f7eb3720 · outbound

This paper cites Improving text classification with transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Improving text classification with transformer,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.797365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.797365Z digest=sha256:b97b2498d0f2011ca285bbf3fec67a9a331e8728ebd53219bab6717c9ece6cb2

Observation 978ff6f8-5e7e-437d-be5b-5011412a4460 · outbound

This paper cites Scaling Neural Machine Translation.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Scaling Neural Machine Translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.800726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.800726Z digest=sha256:ce2adf80019041889f8b28a0f5013c1ab518f5189d77bcc166b8100f9d7b179d

Observation be4655b7-aea5-4364-b2f2-dd5ae28b5a34 · outbound

This paper cites Block-skim: Efficient question answering for transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Block-skim: Efficient question answering for transformer,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.804693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.804693Z digest=sha256:2570941c579db0f9d67e7cbea68ce97cd69e3a0872665341c4fd10330a89c847

Observation b2660a75-897a-4d33-bb3f-35ccf3443b7f · outbound

This paper cites An attentive survey of attention models,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques An attentive survey of attention models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.808452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.808452Z digest=sha256:834086e99e6ec601562ec103b17491e96fabc3eb952a2ad9d5041eba0196067b

Observation c6c1e688-a417-42a8-b552-c22a9386b9f3 · outbound

This paper cites Attention, please! a survey of neural attention models in deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Attention, please! a survey of neural attention models in deep learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.811824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.811824Z digest=sha256:80a84c58e0aa33c6907baf2b1bf6cd153567dc158aa6dda11e078bc5f506e563

Observation 6710cb2b-7a60-47f6-bbc8-cc316fb1a6b0 · outbound

This paper cites Show and tell: A neural image caption generator,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Show and tell: A neural image caption generator,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.815532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.815532Z digest=sha256:4b0a923180ea068341391d481da80275972083aae94804189880b046ef27b6c2

Observation 3ae31ecd-3f5e-4c04-b5b3-662fbbc0d25d · outbound

This paper cites Deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Deep learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.819461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.819461Z digest=sha256:a05a310e2c98567ed48e635f2f4de4ed4613ff1713f6fda64aac22037435bfad

Observation 19dd7d3c-cf54-4f03-9153-8cece3bdd3da · outbound

This paper cites Deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Deep learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.823248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.823248Z digest=sha256:51015dbf3b663c866ba8c157483d062efa9600cd25e30bc2e69c879c227b39c4

Observation ed2ed15a-de27-4b25-a446-5ba96e48fe8b · outbound

This paper cites Long short-term memory,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Long short-term memory,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.827223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.827223Z digest=sha256:7936de309fc921b9f040c71cb0395f0b3e9a906197a9c165426933d28457dca1

Observation 91a31ed1-48ed-4645-aca7-f84d249d3ae8 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Expert-Level Medical Question Answering with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.830717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.830717Z digest=sha256:303a02dcf23718bd847bcf3ac9d2075ed1a83266dd4e231cdf552458976d1bbf

Observation 61ad6437-42e6-421f-a37d-6989b2f3066d · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Training data-efficient image transformers & distillation through attention,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.834768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.834768Z digest=sha256:61e43bab79f3bfa3ed7fbce7c62ce7170eb41b6b6db075c6cba5e46e81a02aec

Observation 026b1338-c47b-486d-8f6f-1b19e7c41a99 · outbound

This paper cites SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.838942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.838942Z digest=sha256:6e1e49efed46f02fa11ace14777077b3a78d7c71e3c1ce91688bb030dac0667c

Observation 17d69191-bd0e-4a14-bc3c-c4a18dc13238 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.843455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.843455Z digest=sha256:6ad36bf593ac31d84d88ea8e1fbf00507fee5487f6dae27dcc9b18efbd548127

Observation 022ef8eb-2881-47dd-8cf3-a1b946ccdd12 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.847663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.847663Z digest=sha256:be37bd06010b2c2a7d2ebe3f405723457a24c551b7014abab959d004a62bb116

Observation f37dcbe8-5666-47ea-9135-c094306895a8 · outbound

This paper cites Gpt-4 is here: what scientists think,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Gpt-4 is here: what scientists think,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.851308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.851308Z digest=sha256:bb06a9c7d2cb4c07a5f58cdbbfdf98a364f17331a8a88a44e5dc29626d903ffb

Observation 7d9c7891-7d51-4166-a825-b38212f263f3 · outbound

This paper cites Language mod- els are few-shot learners,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Language mod- els are few-shot learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.854998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.854998Z digest=sha256:02cbc9e62235a4bd017fcf05360fb437a9a8da6072dee8808a4efe50274f1329

Observation 5296042b-7465-4dcb-acdb-81f0fe42a09e · outbound

This paper cites Multimodal deep learning.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal deep learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.858659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.858659Z digest=sha256:d216f444881e1ae407deaddc23cff0b69ee744bf74e9bf1e533fddf2aaa73563

Observation 22078948-1812-47f7-aa95-707bf650affd · outbound

This paper cites Multimodal machine learning: A survey and taxonomy,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal machine learning: A survey and taxonomy,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.862350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.862350Z digest=sha256:21f0519276be9f359db2c6d4bc4e6798c2af6ee099f6cbd6d7c6fa93ad082766

Observation 81b85cd3-5be3-4683-a409-c9a9637d77f2 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal transformer for unaligned multimodal language sequences,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.866015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.866015Z digest=sha256:9e0762fe23acb1d744ba882f7572faa12b7044dcf6990ceb12dcd3d6f81ba209

Observation c9922b25-8d4c-44a1-a3b3-5ced6bfeb68f · outbound

This paper cites GPT-4 Technical Report.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.869757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.869757Z digest=sha256:c348a69ffc2e4ff3a2bbe699cc5604b3329c15f22d587257d2b978060ad97b4b

Observation afed7e32-5d07-4dbe-837e-0300bfc2d845 · outbound

This paper cites Crossner: Evaluating cross-domain named entity recognition,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Crossner: Evaluating cross-domain named entity recognition,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.873536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.873536Z digest=sha256:369d1c991ea6476d39fb9b775b8e3adfc185320bb95d1d6c0c4abf22ecf87681

Observation 770a5b30-5cab-4377-97c3-9a1bac24c3d7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Learning transferable visual models from natural language supervision,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.877137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.877137Z digest=sha256:7209b0ecbde7b7f09928f697cea85e0b12907373c4fa2ef19f1e59f155d2e46d

Observation 087f0a8a-8a25-44d6-bc85-4db173ac2277 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques On the Opportunities and Risks of Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.880936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.880936Z digest=sha256:6b495abdedf1a513d9a88d8125212071283c2edf24a9d3acd11989a1a7af98b6

Observation bade9839-b308-4cfe-bf10-2a8dc2d847d5 · outbound

This paper cites Xtext language engineering for everyone,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Xtext language engineering for everyone,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.885110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.885110Z digest=sha256:d0f8df63293dc0a8b9940a105045562768d17dbb47e57332ccd3d38371b3ddcb

Observation ec150ed6-967c-40c8-a0fa-2959ee49851e · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Align before fuse: Vision and language representation learning with momentum distillation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.888633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.888633Z digest=sha256:99093b8dee14d5d1ce44cc5a41b3174a17b15cca8d76c925ab55366796b5a9cf

Observation 1a5c4180-e441-4f2a-9a48-1367d5f4289d · outbound

This paper cites Training language models to follow instructions with human feedback,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Training language models to follow instructions with human feedback,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.892388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.892388Z digest=sha256:62074d3c1256f400207b7b46dc6f6c87c12f1264b3195467e0be7c76a79cdf89

Observation 3b22685c-af14-40c0-8907-133b90e7a683 · outbound

This paper cites Hymba: A Hybrid-head Architecture for Small Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hymba: A Hybrid-head Architecture for Small Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.896718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.896718Z digest=sha256:5b988bdcd80036ab713dd95cf5b629bca2151d036f1c4ef0584d8ce126918a13

Observation b3a4de7e-c2fd-4b95-a2d4-19d0257026f4 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.900912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.900912Z digest=sha256:be16a2dffa94a188ab38efac3ac5fd1a11ac4212ff5fb9e560793b4c5a0e082d

Observation b66c68af-5e7d-4b6a-b49d-f98e1538b848 · outbound

This paper cites Small Language Models: Survey, Measurements, and Insights.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Small Language Models: Survey, Measurements, and Insights

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.904953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.904953Z digest=sha256:3a16bddebee539fde90f1655f42cf8aade93f00a5bcd6a6b22d0c79334514683

Observation ec4eb572-b1a4-4129-b4ec-aee6e84b0c0c · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.908957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.908957Z digest=sha256:f7ab90913079e4948dc48d51de1d878144ce0a85ed48fa8697e2cb7b3026dab3

Observation d1c0e700-7c5e-4bb3-ad41-ee82a0f844d6 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.074869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.912852Z digest=sha256:bc779eedc030e08fb51e559cfb4664103dbd9302575e5e9e757d05740d896953

Observation 0edc626f-8814-43c8-b43b-70c8fc73aec6 · outbound

This paper cites Towards trustworthy llms: a review on debiasing and dehallucinating in large language models,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards trustworthy llms: a review on debiasing and dehallucinating in large language models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.062327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.916756Z digest=sha256:d73ae057387a715ee86f7724858d5d12b296fb85ed315ff8468fa8f3092c4efb

Observation c7ad77f5-909c-45ae-9d6b-93fce6f8cad5 · outbound

This paper cites Sometimes painful but certainly promising: Feasibility and trade-offs of language model inference at the edge,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Sometimes painful but certainly promising: Feasibility and trade-offs of language model inference at the edge,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.920457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.920457Z digest=sha256:5f9ef23a84a60e66ac2251135837a6400066f9c0c4c662a649449b64cd792e64

Observation 5b153ab8-f3a6-4034-be29-cb7af28c0ff2 · outbound

This paper cites Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.924187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.924187Z digest=sha256:236e7da22afc94b605d50a16bb2f93698b43e42653267c2647e77c719a01af47

Observation 2763b83c-da63-41cd-9755-02619710b949 · outbound

This paper cites Fly-swat or cannon? cost- effective language model choice via meta-modeling,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Fly-swat or cannon? cost- effective language model choice via meta-modeling,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.050512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.928196Z digest=sha256:8af2b157273276ddbd87633e28badfb42bb8c31e2f733f0813dcac8f62a8b0c6

Observation 3eb1df40-ec5d-4c25-9ffe-fb80b44b4a74 · outbound

This paper cites Routoo: Learning to Route to Large Language Models Effectively.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Routoo: Learning to Route to Large Language Models Effectively

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.932292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.932292Z digest=sha256:54fe02ef6923cecb5b1bc84d9dd598ef578aa31639a10d7fcd6c58fd51a90931

Observation 369e619b-111c-4d37-9083-b7527010f245 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.936401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.936401Z digest=sha256:d6a488da62236725ef0e5d58817d87b87b8ce71c7b20f2e6d2dbef22db6d7bba

Observation df27739d-3704-40bb-8686-deff835a0c54 · outbound

This paper cites OptLLM: Optimal Assignment of Queries to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques OptLLM: Optimal Assignment of Queries to Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.940328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.940328Z digest=sha256:d43a2d93180cc3d75c381cb1def428c4fab96e86290c92dc323e7ef7489edb20

Observation 0c9c37a1-40eb-474b-8550-5c81a28095c2 · outbound

This paper cites Chatgpt,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Chatgpt,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.039550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.944254Z digest=sha256:0160af8fef428e770ec5ce1c0265a74bcc3e0d99969545b93a00897bea2afb96

Observation 42d3dc4a-2a64-41c9-99a2-6b0499547c30 · outbound

This paper cites Amazon titan in amazon bedrock,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Amazon titan in amazon bedrock,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.028097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.947839Z digest=sha256:72801044bb38b4f14d0d9dd2d8211b181159436273d78cdda15de287a06f5ce7

Observation f221f8e4-7d12-4957-84a7-c3924a8fdfe8 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:58.017464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.951394Z digest=sha256:8a76a9baa31883c124bbd22ec4e833517cef6fc912bc11f84206679d876c9a2f

Observation b1854396-ee05-40c7-bc01-4155a4feb7b4 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:58.007120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.955137Z digest=sha256:eb3c1ee222f3094e2bdd80650d540aab7b05dab3d1bdf32bfe5c1f71f2dc54ee

Observation 03951d37-536d-408e-baeb-00a54b6df901 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouteLLM: Learning to Route LLMs with Preference Data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.958761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.958761Z digest=sha256:76f5b338278b36a3e7894207860049cbadc9472d56ff468cc1f5e5a466a24b7a

Observation ddbcf273-44eb-4338-99b8-695ce6e04f59 · outbound

This paper cites Cache & Distil: Optimising API Calls to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Cache & Distil: Optimising API Calls to Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.962993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.962993Z digest=sha256:173054b62603b655b0f4b99260107d7198a5cf9a5433e66f40b82472aadc3394

Observation fd6524a2-e721-4cb8-8606-52316e8eb167 · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques AutoMix: Automatically Mixing Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.966738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.966738Z digest=sha256:9e474efeea54fa502cf6a7508acaa728999683b8c8993f5ff1254dba488c2151

Observation 6547ef9b-67bb-4c49-a87f-454fe5b60ac1 · outbound

This paper cites Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.970762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.970762Z digest=sha256:a523cbf3dfcd38fd9ca664ab2945cf2eb60c3acf6d10c031d8e3a213bbb4a6a0

Observation 41f0d6a2-dacd-4685-9ca5-421c3cbd558c · outbound

This paper cites Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.974688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.974688Z digest=sha256:b51dd1a4287b1038a1211ff6a300d7418f148a29861462bccf2217c0c4d259ee

Observation 39d584d9-4d79-42ad-a834-f6714631ce83 · outbound

This paper cites Query by committee,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Query by committee,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.978614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.978614Z digest=sha256:7c6c787750e6c4c9eabe644777b8902874f47f23414e438dc70aca838b78d825

Observation bc603551-9d84-4c2e-8cea-a0e8cf2a4c24 · outbound

This paper cites A review on edge large language models: Design, execution, and applications,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A review on edge large language models: Design, execution, and applications,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.990897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.982111Z digest=sha256:97168d7ce0a489efd33418078834310e2763ef0a8b7ef4e137301125f9827e02

Observation 92755cfd-9dbb-452b-be1b-aa346eb0f90f · outbound

This paper cites On-Device Language Models: A Comprehensive Review.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques On-Device Language Models: A Comprehensive Review

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.985768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.985768Z digest=sha256:e9b14a9568019cf5e59ec3c09d31622c0621326f1e94926415e3d10f575a5b43

Observation b897ca20-8400-4e57-94b6-ba9e268f6016 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.989615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.989615Z digest=sha256:de89e3bbfd0d99e77138406a88f4e2482abc8b430ff00264d30161845cdae743

Observation 48890071-a8b5-4073-b2c3-b0e37f358712 · outbound

This paper cites Research.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Research

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.980441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.993763Z digest=sha256:3858abc4d4cf0a29bf317dba8c6dc0c74d0972be8eb6a7087c8c35a03c38a2a8

Observation 6d0c202c-d3d9-4b8f-aace-aa73d8d9d901 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:57.970312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:56.997453Z digest=sha256:d6fc8dfde52370a48cc1e9e2b73d88ae371bd27a1395ff24ce78d89a358bac2d

Observation b853f9ce-101b-432a-bc3d-cfbee91bce19 · outbound

This paper cites Edge-first language model inference: Mod- els, metrics, and tradeoffs,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Edge-first language model inference: Mod- els, metrics, and tradeoffs,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.959881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.001152Z digest=sha256:bded0e85ac7586d6cd3bc9c885e76911eba3f1582f40182a9dfe5a9bc714638e

Observation 47546da5-a959-4dd8-ae59-e3497242d1d4 · outbound

This paper cites UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.004691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.004691Z digest=sha256:72207ad4d5cc8a7ed8f5ffb97bc8074aac1d8415adbdd47c7334254ff95c6c5d

Observation f2707b7c-9c27-42ad-9a28-db59fb148ffb · outbound

This paper cites Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.008858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.008858Z digest=sha256:883c35a7d74a047ce0d9ab20b1e6408d07f7ffcd54064319f2fd24bda5261edf

Observation f8fee592-837e-4d70-81fa-7623c76188d6 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.013028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.013028Z digest=sha256:f6c8a60f5065fc2202b7186b9f17107af976077b87179db11468152c77efa59e

Observation 0a64c309-4bed-4ea1-9647-9e54130324de · outbound

This paper cites MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:56:57.208416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.016882Z digest=sha256:122d65094aedae6f345383c017edf5ad91c392d2e1ee4fe9b8ec1e6a3235a121

Observation 0d5ee2e9-8d16-4ba5-9f93-528ee93467ee · outbound

This paper cites Multimodal large language models in health care: applications, challenges, and future outlook,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal large language models in health care: applications, challenges, and future outlook,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.949512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.020774Z digest=sha256:cdf2fd8021e989908930453fcd75645dbc8d79aa4fba25833fae892bb41e84d6

Observation 37db8c79-d54a-49f6-9a13-f52a9e9359da · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.939194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.024727Z digest=sha256:9e3db5f727f0cbd542617faebffc5784aa0bf00896e4b4c3322af0af9c5f7980

Observation 8356649b-3d40-4c5e-8c39-f0310af58151 · outbound

This paper cites Lexical complexity predic- tion: An overview,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Lexical complexity predic- tion: An overview,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.927904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.028587Z digest=sha256:fa355e8f1245a4bc59acdfc257ea575a998da11e26a1351fb09fe4698a94d446

Observation 45265c53-9cc2-4c58-ad15-7189185fc78c · outbound

This paper cites Collaborative cross- modal fusion with large language model for recommendation,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Collaborative cross- modal fusion with large language model for recommendation,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.916891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.032260Z digest=sha256:e8d372e5e276fd2d67a4f9bb1fa11ad11f42995bf63594a588713073cb21048b

Observation 639313a1-8624-4d86-bd63-0c21eb2fcc4b · outbound

This paper cites Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.036167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.036167Z digest=sha256:717c841111efc3d7d86b86abfee55b5e746a4111878b8126b9d566153eca73f6

Observation f9cf4371-264d-46b4-9bf7-326b519ac0a1 · outbound

This paper cites Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.040284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.040284Z digest=sha256:7748233d54c5642b0460e280e26ff7ffc3f6c46a45726010f8713a37cdfb02ff

Observation 5fa027e9-05a7-4355-bc66-3c02c9e9534c · outbound

This paper cites Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.044447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.044447Z digest=sha256:4444e90c54bb2b59d47fe11f17c9114fbdcde11bdafcad7569fb22802cfdce74

Observation 19bbad48-bb75-497c-ae97-ff936e166272 · outbound

This paper cites Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.048415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.048415Z digest=sha256:3c6d5cbfb2d33e5e1ca72b5a5463ea28de147ce377c40ad8924ea4467e72f22f

Observation ccf21cc0-c770-4c37-83a8-322785e006a9 · outbound

This paper cites XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.052089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.052089Z digest=sha256:e0f9b1559e3237cf3057ae939e412771e193f77e91e40f3b736d387a0e8bf2e2

Observation 65439416-3169-4241-99f4-6904b9f83c1c · outbound

This paper cites PickLLM: Context-Aware RL-Assisted Large Language Model Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques PickLLM: Context-Aware RL-Assisted Large Language Model Routing

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.056470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.056470Z digest=sha256:1f8f8c65e37a285eab3bbd1954adc233636b5eb39d380cb1b682ce61cb66f368

Observation 2c26cbc3-633d-40a4-8602-45546a00ee96 · outbound

This paper cites Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:56:57.138941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.062055Z digest=sha256:3c7f6409e19611f74d3f1c564b330e9620ea44a942f07ce9b8c70ef148610a0c

Observation f6fcb862-cb94-4226-a586-8dd499f3a109 · outbound

This paper cites Trusted llm inference on the edge with smart contracts,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Trusted llm inference on the edge with smart contracts,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.906148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:56:57.066381Z digest=sha256:d3fa15e6cf1284b0bbee7cc8e302bda10b9a86f4235b366a4b4469bf75f19fac

Observation 055c6e2d-3fc3-4bae-b046-10296140e601 · outbound

This paper cites Encryption-Friendly LLM Architecture.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Encryption-Friendly LLM Architecture

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.070193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.070193Z digest=sha256:afc2c5fa49cb2dc85654c120e0fb223408610c44a4e48a799fe6651de76fa7da

Pith citing papers

Observation 8bd8a001-206d-4097-95ca-4f1690881d27 · inbound

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems cites this paper.

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:10:53.372941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:10:53.372941Z digest=sha256:c9255a3f76547dc5c435de20e59b7768cf2bcdbff43d8c4f174d371624fb1e3e

Observation 51c91ec7-75e7-407b-bacb-1972153792d3 · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:25.790205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:25.790205Z digest=sha256:2237b963e2bb9c6affd3a10d565b4f820e0b9cd63b04b9af3d366edd08832a92

Observation 744745d6-6957-4b28-bb15-51ab1b51eaec · inbound

When Less is Enough: Efficient Inference via Collaborative Reasoning cites this paper.

When Less is Enough: Efficient Inference via Collaborative Reasoning Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:42.531986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T19:27:04.267404Z digest=sha256:240b16fb6324611716a5a5bcadfa8bc494732b2fc86ef2b274d1dd4bd3b05eec

Observation 34674ce5-e34d-40c4-97ff-f856eab0a7c5 · inbound

SOMA: Efficient Multi-turn LLM Serving via Small Language Model cites this paper.

SOMA: Efficient Multi-turn LLM Serving via Small Language Model Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:04.221906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T01:37:27.361504Z digest=sha256:b3613abb6c48d12e6b1460f81edd2ee64c2f62be891608a2e6d817339abef952

Observation 4f495410-efb5-4926-87f1-8dcf3557bdb5 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.605739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:e5c6e37539aceb37c260a5dde0bcc74c0801a5ca300599c8065907342163e4d1

Observation 4fbda2fd-a0ef-45e3-b4ae-6be656850ee8 · inbound

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing cites this paper.

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.806454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:54:28.452796Z digest=sha256:585e50f23d7c581e156828ee6287ee5e4eedfc3757424f1872230102b241a8b0