Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

As of 7 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 5 inbound Pith citation observations for arXiv:2506.06579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06579 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:56:57.070193Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:34:25.790205Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 18d0f5d8-50b3-4731-9b87-671596bc7b41 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.676293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.676293Z digest=sha256:fac9ed3aad9209ab3a9b84ba072ed23fa5ffc28fc28e65a9c78e941b7951a05e

Observation 81b4d2ff-f2a3-422a-9949-b79583aae3c4 · outbound

This paper cites Gpt-3: What is it good for,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Gpt-3: What is it good for,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.681421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.681421Z digest=sha256:cc4a23f401af7282903427e8b3f319e70d636482dc616f895019d14c4eb4e148

Observation 05ee312a-aeb1-4c58-afe7-62044f8759ee · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.685256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.685256Z digest=sha256:e0ec10005b8008d720243bd337ee62403bf68f728e790d9d8ba52a0e2fa043ed

Observation b3796b1a-f0bc-4e92-a908-bdbe0898d0e7 · outbound

This paper cites Mobile Edge Intelligence for Large Language Models: A Contemporary Survey.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Mobile Edge Intelligence for Large Language Models: A Contemporary Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.689467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.689467Z digest=sha256:64ded98bd351dc45aeab07d72096ec6b5a9a7b24c661acbfe8549447582c42a9

Observation 46ff31e9-dd1b-46b2-96a3-52481f9c4c92 · outbound

This paper cites Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.693866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.693866Z digest=sha256:2a9c2b925f9942024859e74249bed6f03f8929c86c9a9a71c897a8ab79ff9c83

Observation 33ce2435-327e-49b8-953f-09acee7540f3 · outbound

This paper cites A Survey of Resource-efficient LLM and Multimodal Foundation Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey of Resource-efficient LLM and Multimodal Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.698136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.698136Z digest=sha256:be673079ca17172d90312cfcc6d03dbb34bd960b15c8def5179e2184a065459d

Observation 9fab6bad-1562-4300-b2ef-c1200ed5900d · outbound

This paper cites Efficient compressing and tuning methods for large language models: A systematic literature review,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Efficient compressing and tuning methods for large language models: A systematic literature review,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.703061Z digest=sha256:6eb6a5ccff0fa3a133ce37f3d3244f453f895a4de8afa22c67b86ad0dbfac2f0

Observation 61647dc1-bbbe-496c-bfb1-6c23b4989b34 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.706799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.706799Z digest=sha256:c600b42c3e148180b459d07afb393241f06fa30c7f4a8157906184584f50ec1f

Observation 3e1f6fc4-ce73-4bd6-988b-01bc1c177bbd · outbound

This paper cites Evaluation of the phi-3-mini slm for identification of texts related to medicine, health, and sports injuries,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Evaluation of the phi-3-mini slm for identification of texts related to medicine, health, and sports injuries,

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:56:57.788517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.710904Z digest=sha256:8fba869f29a94cc81a8dc76bbc8cbaeb6d57417e4f9a3e383d710cbd5686f039

Observation e5184609-c34b-4f89-b9d8-454290dbc810 · outbound

This paper cites Harnessing moderate- sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing moderate- sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.714951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.714951Z digest=sha256:bb2c2b37349ea081d80179a9a297aea80ac402a36344ce71bfe7f3285cc7ce7e

Observation ef61aed4-6660-40d4-a88f-3ed439004692 · outbound

This paper cites Overview of small language models in practice,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Overview of small language models in practice,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.718751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.718751Z digest=sha256:a9943f6f5aac5d49d423e3ac6cf134e0efcade3634f72e4578faa92bb086741e

Observation 25ca84e2-fe07-4ddb-9284-b90423a8e8de · outbound

This paper cites MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.722420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.722420Z digest=sha256:20cc4b59560856e99913000a1cd4dad13b5f7206b4d2e9eb393d69dc117bccbd

Observation 5cf3d3c2-0771-478c-a869-9735cf428956 · outbound

This paper cites EcoAssistant: Using LLM Assistant More Affordably and Accurately.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques EcoAssistant: Using LLM Assistant More Affordably and Accurately

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.726464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.726464Z digest=sha256:639c889c30e3dc1618438a4b7c3bab7090d5068d65715b8e7a5ce6afa2c737cc

Observation ebb02bfe-3add-451a-adfe-b3185ef308e7 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.730282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.730282Z digest=sha256:34ff2f7e6a985723b5fe02e2205b9063ed411806372df06b9acee02ee86b93f3

Observation 1adff952-db0e-4c63-9f30-5560021ae50e · outbound

This paper cites Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.734663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.734663Z digest=sha256:c8ebed5e317a7c8d4c2c053a045e7c461dbcd419f38624a6ddd6544c8f091a4b

Observation c53fac9d-a16e-49b1-9c65-099db0fc7b9b · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.738323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.738323Z digest=sha256:49b69fe1c9217d54895990b32f4f7b7b70ff4680018c7961f40f7ddd33b2b6af

Observation e13a777d-61db-46c8-bd3e-3388bac034aa · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.742261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.742261Z digest=sha256:f33976b504c93c67360ae7a2208a23d03ccd3df0f6192b527ba2bdcb73e18a1e

Observation 109d4bb6-2963-4afc-a98d-8f293c457236 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouterBench: A Benchmark for Multi-LLM Routing System

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.746355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.746355Z digest=sha256:9d1b7482c42d514d4125b092404dd4ad4ba75413031ab0456fe0046c88d0aa5a

Observation e14414f2-d12c-4f9d-85c1-466be764662a · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Efficient Inference for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.750392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.750392Z digest=sha256:3a8365c2d4a5e3ae5235ab336e7218dd4df8988634b88e686be7ad646228a8ce

Observation 5e8ed8fa-1544-40d7-b95d-3d65bd8a07dd · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.754256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.754256Z digest=sha256:8696fc2e18cb268c6fe04a9fba6a4a4300d722b193f655dac3004d667550fc45

Observation 0c823151-99af-4d27-9d87-0b6c60cb9593 · outbound

This paper cites Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.758394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.758394Z digest=sha256:42721ca81ffea1605d06ebebad8c8b96a28f8a6c30e9f2ab795f335ce4d0428b

Observation 23b5ca1b-7de2-48db-b85d-84b7783cced2 · outbound

This paper cites LLM Inference Serving: Survey of Recent Advances and Opportunities.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Serving: Survey of Recent Advances and Opportunities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.762293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.762293Z digest=sha256:e1b9409875ea391720f86b88eaeb0a8d4ab8f7bdc139aedf1cb76e72ea65d57a

Observation 8a93df01-9df9-49f3-a6fd-8babc8be4a46 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.766216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.766216Z digest=sha256:bf56eb565b045484000883db5b620d1a87073f976937af52f15ff5ef465afdc2

Observation 1d8ad172-a5ce-46a3-ad55-399e14c4ea5e · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.770100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.770100Z digest=sha256:8a5592e4636882000f354c7683283488755a5ff9b6c896d6dae70f4f1f4ad524

Observation 5a9e363e-7f67-4f63-9e44-a51864a34603 · outbound

This paper cites The Efficiency Spectrum of Large Language Models: An Algorithmic Survey.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques The Efficiency Spectrum of Large Language Models: An Algorithmic Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.774041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.774041Z digest=sha256:ed8d10507b101d7523cb3c11434f7f340ac547a845276af4a28052f06325ed43

Observation eb037d63-9b43-4127-a192-d9fc29db83ea · outbound

This paper cites Harnessing Multiple Large Language Models: A Survey on LLM Ensemble.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.778198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.778198Z digest=sha256:3ffe405dde03377d4950b9fe4843dd879435fb9a1d54763c9ee92c87d773a28d

Observation 6039e81e-d05b-4f0a-ae3f-a52ab2991376 · outbound

This paper cites A Survey on Collaborative Mechanisms Between Large and Small Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Collaborative Mechanisms Between Large and Small Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.786194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.786194Z digest=sha256:7090314dd398198cc158121e9f17328e2a1da4c07715c425cf1c8e6240eb1e87

Observation 5de78550-fd1b-4748-9cb0-1c7e64efeb56 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Mixture of Experts in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.789923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.789923Z digest=sha256:fadf09a5e65409b2ba894c1c7103f83474958307adfad96b5bd04f9b6d018642

Observation 3577678f-29a4-4934-b91a-f0345d91bd84 · outbound

This paper cites Attention is all you need,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Attention is all you need,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.793800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.793800Z digest=sha256:63b3f2edc43046c13fdd526ab739d80f6d5874f6f1e8f302f1fd7f8cd5c5e290

Observation 0bbbec51-4dd0-4c8b-a035-89a8f7eb3720 · outbound

This paper cites Improving text classification with transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Improving text classification with transformer,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.797365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.797365Z digest=sha256:0bddd35c85df29f3490b2df9fdb4f658a5ca00e337bd795df00b0835001b0f47

Observation 978ff6f8-5e7e-437d-be5b-5011412a4460 · outbound

This paper cites Scaling Neural Machine Translation.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Scaling Neural Machine Translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.800726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.800726Z digest=sha256:ec474b3efd5433c087a202df4b387c9b653372f5bfdc1078ec9fe7c203144319

Observation be4655b7-aea5-4364-b2f2-dd5ae28b5a34 · outbound

This paper cites Block-skim: Efficient question answering for transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Block-skim: Efficient question answering for transformer,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.804693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.804693Z digest=sha256:9dc9d15da1e334a255c54b4df8e05df2779b0a42100da795ce8d6af55db2ed06

Observation b2660a75-897a-4d33-bb3f-35ccf3443b7f · outbound

This paper cites An attentive survey of attention models,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques An attentive survey of attention models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.808452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.808452Z digest=sha256:6d0bd97ef8ef891818cfd9b42c73a958dcdebee0481ae9db29c2ea8e7460679c

Observation c6c1e688-a417-42a8-b552-c22a9386b9f3 · outbound

This paper cites Attention, please! a survey of neural attention models in deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Attention, please! a survey of neural attention models in deep learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.811824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.811824Z digest=sha256:54a45f5a5a776799b0650b495a3db9af9f645f78aca2cd37d91cb5b8d89dee61

Observation 6710cb2b-7a60-47f6-bbc8-cc316fb1a6b0 · outbound

This paper cites Show and tell: A neural image caption generator,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Show and tell: A neural image caption generator,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.815532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.815532Z digest=sha256:3f5ba6e3d6bc928af954992b5c361f5acb2dc081144aedb96375546b55c8ea79

Observation 3ae31ecd-3f5e-4c04-b5b3-662fbbc0d25d · outbound

This paper cites Deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Deep learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.819461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.819461Z digest=sha256:587aa88e4994f716e5cba953be5232d69258ec3f453cf6e1cb0b6e4cc5df1782

Observation 19dd7d3c-cf54-4f03-9153-8cece3bdd3da · outbound

This paper cites Deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Deep learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.823248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.823248Z digest=sha256:bca6a8ab35e182582edb4345a0a1d967cdb0d8e1da56d689bf82f4fb1ea5a914

Observation ed2ed15a-de27-4b25-a446-5ba96e48fe8b · outbound

This paper cites Long short-term memory,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Long short-term memory,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.827223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.827223Z digest=sha256:07ecbe9bf0b1275e8fab7044124163e0813172a60d2aea9835762aa2512481c6

Observation 91a31ed1-48ed-4645-aca7-f84d249d3ae8 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Expert-Level Medical Question Answering with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.830717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.830717Z digest=sha256:c8ca7941f1d6aa23a9d7cac2bd0828e0c09621a7ce9eeb1f8fbb8b6d3e994e1a

Observation 61ad6437-42e6-421f-a37d-6989b2f3066d · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Training data-efficient image transformers & distillation through attention,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.834768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.834768Z digest=sha256:8e2a824019f597a9a6c2e9f71940c795a4c3817a97b29fe474b9024196eeb601

Observation 026b1338-c47b-486d-8f6f-1b19e7c41a99 · outbound

This paper cites SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.838942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.838942Z digest=sha256:e66cea60f30a1da93f41a7f97d0f4dbc8eceaf491b0bb48051ae122b1e009711

Observation 17d69191-bd0e-4a14-bc3c-c4a18dc13238 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.843455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.843455Z digest=sha256:ed66124e02891a7be62bf22f6ab2b465f8ff3e3361df9ab27b45250c8891d6bf

Observation 022ef8eb-2881-47dd-8cf3-a1b946ccdd12 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.847663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.847663Z digest=sha256:55bf84bb3b76e15b7a373baf6236f427f51512d97186bff12705d78afe076b22

Observation f37dcbe8-5666-47ea-9135-c094306895a8 · outbound

This paper cites Gpt-4 is here: what scientists think,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Gpt-4 is here: what scientists think,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.851308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.851308Z digest=sha256:c9a937590df9240ea043aaaccf6019f72509b694f993525da98cce36a5789a72

Observation 7d9c7891-7d51-4166-a825-b38212f263f3 · outbound

This paper cites Language mod- els are few-shot learners,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Language mod- els are few-shot learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.854998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.854998Z digest=sha256:06a4a74ea28d18747d0387d991bd23593ef3957b80abdaeaea50256377f5e1ff

Observation 5296042b-7465-4dcb-acdb-81f0fe42a09e · outbound

This paper cites Multimodal deep learning.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal deep learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.858659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.858659Z digest=sha256:295ca36878e7ddab6e85c24df3860d51b831aa6ba4046985ae1d65dc9c3fc685

Observation 22078948-1812-47f7-aa95-707bf650affd · outbound

This paper cites Multimodal machine learning: A survey and taxonomy,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal machine learning: A survey and taxonomy,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.862350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.862350Z digest=sha256:e12e597bfad964108300e791c963174d439b5d8dec249b0578457ab83e004a7f

Observation 81b85cd3-5be3-4683-a409-c9a9637d77f2 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal transformer for unaligned multimodal language sequences,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.866015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.866015Z digest=sha256:384cb5971e5994c24f5386b3367f372ef4380169aa743138041942fa73ddfc6e

Observation c9922b25-8d4c-44a1-a3b3-5ced6bfeb68f · outbound

This paper cites GPT-4 Technical Report.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.869757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.869757Z digest=sha256:ead646559f0b2f63943e207ec25a168d49aedd0e11161d043c988635d73680ea

Observation afed7e32-5d07-4dbe-837e-0300bfc2d845 · outbound

This paper cites Crossner: Evaluating cross-domain named entity recognition,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Crossner: Evaluating cross-domain named entity recognition,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.873536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.873536Z digest=sha256:64ad9bbd997485183294d13ff81f23e55df751f15666b70ea40dfe6c19c57699

Observation 770a5b30-5cab-4377-97c3-9a1bac24c3d7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Learning transferable visual models from natural language supervision,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.877137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.877137Z digest=sha256:b9b479891d84ed331674f3f8b93ccc4482ed98223d3c6f486042cc864925ead9

Observation 087f0a8a-8a25-44d6-bc85-4db173ac2277 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques On the Opportunities and Risks of Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.880936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.880936Z digest=sha256:e9227ea6de681e0d4defef47f14321b936183f20b0311f90d20d0c8c8c82551a

Observation bade9839-b308-4cfe-bf10-2a8dc2d847d5 · outbound

This paper cites Xtext language engineering for everyone,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Xtext language engineering for everyone,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.885110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.885110Z digest=sha256:8d91a499825a1fff52c132c0432d69729613e07ffc365db20bb7265dccfe65c9

Observation ec150ed6-967c-40c8-a0fa-2959ee49851e · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Align before fuse: Vision and language representation learning with momentum distillation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.888633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.888633Z digest=sha256:61e7b916f14c244d721cb0c985adcfa19cf3f24a70675b8de13bf3713bdbe0b4

Observation 1a5c4180-e441-4f2a-9a48-1367d5f4289d · outbound

This paper cites Training language models to follow instructions with human feedback,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Training language models to follow instructions with human feedback,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.892388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.892388Z digest=sha256:0e6998e5edea86b7876d03240f49f141a0941fb66054ea1ddf8baf8099610fa4

Observation 3b22685c-af14-40c0-8907-133b90e7a683 · outbound

This paper cites Hymba: A Hybrid-head Architecture for Small Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hymba: A Hybrid-head Architecture for Small Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.896718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.896718Z digest=sha256:94249408070e145f2a8d0dddcae44885c70d2c32bf49a89a768268f3d73899bd

Observation b3a4de7e-c2fd-4b95-a2d4-19d0257026f4 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.900912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.900912Z digest=sha256:eb3e36617c711d59a0016e9bd3f341aa73dcb22ed812c492a0c53205b2071050

Observation b66c68af-5e7d-4b6a-b49d-f98e1538b848 · outbound

This paper cites Small Language Models: Survey, Measurements, and Insights.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Small Language Models: Survey, Measurements, and Insights

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.904953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.904953Z digest=sha256:d57d6531261592e2cb4084b14615e6ef60095bb6a91ec28e432e1424c0aa2992

Observation ec4eb572-b1a4-4129-b4ec-aee6e84b0c0c · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.908957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.908957Z digest=sha256:a4297bc5c01642c521263a0361c04235d4aa408921d83814cec8319fc362916d

Observation d1c0e700-7c5e-4bb3-ad41-ee82a0f844d6 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.074869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.912852Z digest=sha256:a8d4050020af140acc71181de7afe695b74a9d035e10eb93a5fd7270bba5b783

Observation 0edc626f-8814-43c8-b43b-70c8fc73aec6 · outbound

This paper cites Towards trustworthy llms: a review on debiasing and dehallucinating in large language models,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards trustworthy llms: a review on debiasing and dehallucinating in large language models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.062327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.916756Z digest=sha256:52eee3207d64f3ab4d30f7b16f95fab8ce092fd0f0c620047543a13d78a46fa6

Observation c7ad77f5-909c-45ae-9d6b-93fce6f8cad5 · outbound

This paper cites Sometimes painful but certainly promising: Feasibility and trade-offs of language model inference at the edge,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Sometimes painful but certainly promising: Feasibility and trade-offs of language model inference at the edge,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.920457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.920457Z digest=sha256:45e09757c2fde743b62c106be862b0999d3a13504c23e60dd2f858f509309799

Observation 5b153ab8-f3a6-4034-be29-cb7af28c0ff2 · outbound

This paper cites Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.924187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.924187Z digest=sha256:9bc785e9c207eaa036a3227e735b86cda7a3aaee7e009110aaf109313aa10ec1

Observation 2763b83c-da63-41cd-9755-02619710b949 · outbound

This paper cites Fly-swat or cannon? cost- effective language model choice via meta-modeling,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Fly-swat or cannon? cost- effective language model choice via meta-modeling,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.050512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.928196Z digest=sha256:d516acf0d1b6efc9ee37bc5c999cf8435d3d97688087a572de475ec00868bb57

Observation 3eb1df40-ec5d-4c25-9ffe-fb80b44b4a74 · outbound

This paper cites Routoo: Learning to Route to Large Language Models Effectively.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Routoo: Learning to Route to Large Language Models Effectively

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.932292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.932292Z digest=sha256:57803b0405b55a37f37fa108a5242e356ea7efc524a70250f7f41801aec9b09d

Observation 369e619b-111c-4d37-9083-b7527010f245 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.936401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.936401Z digest=sha256:e86161875f43696091547ed5e8e8dc6a79ed88a1b2440035d8828036a15c734b

Observation df27739d-3704-40bb-8686-deff835a0c54 · outbound

This paper cites OptLLM: Optimal Assignment of Queries to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques OptLLM: Optimal Assignment of Queries to Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.940328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.940328Z digest=sha256:d20a3dfaef7455a2726d0398fe5a6b4273724b8917062cc093c10d5ba7008004

Observation 0c9c37a1-40eb-474b-8550-5c81a28095c2 · outbound

This paper cites Chatgpt,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Chatgpt,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.039550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.944254Z digest=sha256:cff675683bf1b1b7186e8eab1f403a68df30d9399865b343b045b833ccfd616a

Observation 42d3dc4a-2a64-41c9-99a2-6b0499547c30 · outbound

This paper cites Amazon titan in amazon bedrock,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Amazon titan in amazon bedrock,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.028097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.947839Z digest=sha256:45a935f3fbc7e9d1f782075aa069d9613bb2f3d5704ce161ee341886f8dbcba7

Observation f221f8e4-7d12-4957-84a7-c3924a8fdfe8 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:58.017464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.951394Z digest=sha256:7154df5cd7c6389d7ab19128776c2be517211026eba79877940b415c90646f75

Observation b1854396-ee05-40c7-bc01-4155a4feb7b4 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:58.007120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.955137Z digest=sha256:3e609e235dda5b325a4ddeb353865f08520ff7ffb7cb5fdb1e64d51a69f0951d

Observation 03951d37-536d-408e-baeb-00a54b6df901 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouteLLM: Learning to Route LLMs with Preference Data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.958761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.958761Z digest=sha256:d2fd1a7654a1913343637a3c73b87679d51a468d78e52cbc8f478332975c74d0

Observation ddbcf273-44eb-4338-99b8-695ce6e04f59 · outbound

This paper cites Cache & Distil: Optimising API Calls to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Cache & Distil: Optimising API Calls to Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.962993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.962993Z digest=sha256:5c08f5212622d814189421617223b0bf4e12ac0838799cd3693b71b5cb35dafa

Observation fd6524a2-e721-4cb8-8606-52316e8eb167 · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques AutoMix: Automatically Mixing Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.966738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.966738Z digest=sha256:39e245d078361bbcac008065f0ec7fff5c7d4c85c0445b188a3c6489978e0c19

Observation 6547ef9b-67bb-4c49-a87f-454fe5b60ac1 · outbound

This paper cites Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.970762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.970762Z digest=sha256:2c11385e2bde615f191e7b0b10435507af10cd2bfcf79508b10fb265f53cb02b

Observation 41f0d6a2-dacd-4685-9ca5-421c3cbd558c · outbound

This paper cites Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.974688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.974688Z digest=sha256:290f974cab7483b0fc2fc5ec80f5f8d10a56fbf948d6b2ef24f64a17b8894863

Observation 39d584d9-4d79-42ad-a834-f6714631ce83 · outbound

This paper cites Query by committee,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Query by committee,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.978614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.978614Z digest=sha256:430787d846e0ab724f14535227d0f2691c2d09ea42a3ece79b9ba91e7e9a96aa

Observation bc603551-9d84-4c2e-8cea-a0e8cf2a4c24 · outbound

This paper cites A review on edge large language models: Design, execution, and applications,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A review on edge large language models: Design, execution, and applications,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.990897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.982111Z digest=sha256:27b4379556000f1642e67119358980f79658f21bc213b9b411abf6e7192b6ac4

Observation 92755cfd-9dbb-452b-be1b-aa346eb0f90f · outbound

This paper cites On-Device Language Models: A Comprehensive Review.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques On-Device Language Models: A Comprehensive Review

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.985768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.985768Z digest=sha256:87a0838faa657718bb5ceadf7e9cb200afc56798aac3dc5a8922b39988a9ee0f

Observation b897ca20-8400-4e57-94b6-ba9e268f6016 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.989615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.989615Z digest=sha256:8f0cd14c842eed2fce4274f14930ee8483e1a251667278db0d176c461fb7e68a

Observation 48890071-a8b5-4073-b2c3-b0e37f358712 · outbound

This paper cites Research.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Research

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.980441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.993763Z digest=sha256:d8108deecee904d5a6b7a62b9b834af23bdee5b3ad4b5835b4a94beb377ef99a

Observation 6d0c202c-d3d9-4b8f-aace-aa73d8d9d901 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:57.970312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:56.997453Z digest=sha256:2b204ba856b26600511af4584340b47552732458a60c62002339bf766a84670b

Observation b853f9ce-101b-432a-bc3d-cfbee91bce19 · outbound

This paper cites Edge-first language model inference: Mod- els, metrics, and tradeoffs,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Edge-first language model inference: Mod- els, metrics, and tradeoffs,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.959881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.001152Z digest=sha256:5c082c013e6ba5995958b087fe79002e14f67c1eb14a279de2f9a416392d9d14

Observation 47546da5-a959-4dd8-ae59-e3497242d1d4 · outbound

This paper cites UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.004691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.004691Z digest=sha256:539c9a344d75b8bd370f6cb26222d101fade13ec7df68d83868cf0d197597840

Observation f2707b7c-9c27-42ad-9a28-db59fb148ffb · outbound

This paper cites Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.008858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.008858Z digest=sha256:24f662a55137afcbaaddc2d1b9756d2fdd55fa86cab377b6480d3f285b468f65

Observation f8fee592-837e-4d70-81fa-7623c76188d6 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.013028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.013028Z digest=sha256:8deecd491d1ef71d3477729cda20724c0efb431a398b156eea9ee9239578ce27

Observation 0a64c309-4bed-4ea1-9647-9e54130324de · outbound

This paper cites MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:56:57.208416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.016882Z digest=sha256:a4c16708af272fb05f9f4cba3a8c6589fc6f44091b542aff73d70eaecd176783

Observation 0d5ee2e9-8d16-4ba5-9f93-528ee93467ee · outbound

This paper cites Multimodal large language models in health care: applications, challenges, and future outlook,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal large language models in health care: applications, challenges, and future outlook,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.949512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.020774Z digest=sha256:5b872c81264d533ad84ed7129b69c93c9c78cb2d03298ff6dd2528b186e78326

Observation 37db8c79-d54a-49f6-9a13-f52a9e9359da · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.939194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.024727Z digest=sha256:d00c11e9449aaf449352ed7d05b92b159a929d0005be4114430393c74ce440be

Observation 8356649b-3d40-4c5e-8c39-f0310af58151 · outbound

This paper cites Lexical complexity predic- tion: An overview,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Lexical complexity predic- tion: An overview,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.927904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.028587Z digest=sha256:574502813f3ad4ce724f45c06ad07ea8daab6ea1f4e21cdbe699103330898fc3

Observation 45265c53-9cc2-4c58-ad15-7189185fc78c · outbound

This paper cites Collaborative cross- modal fusion with large language model for recommendation,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Collaborative cross- modal fusion with large language model for recommendation,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.916891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.032260Z digest=sha256:64e5f5e1608447b46a9f289d923d50adf83219594a538360dc1755de77d0f5c4

Observation 639313a1-8624-4d86-bd63-0c21eb2fcc4b · outbound

This paper cites Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.036167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.036167Z digest=sha256:58bd0be335a58294a3d4d637271211b7880e512ecfad3b0183eed69a1e7d680f

Observation f9cf4371-264d-46b4-9bf7-326b519ac0a1 · outbound

This paper cites Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.040284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.040284Z digest=sha256:ff38c55c6dca100895c2f6eff1cc3757d0cae4178f02a1309819acc70ed51ad0

Observation 5fa027e9-05a7-4355-bc66-3c02c9e9534c · outbound

This paper cites Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.044447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.044447Z digest=sha256:e869481f7170cc97e1b4eba86cb0df1dc01cfc120296ab51b750b7b5cb1911e6

Observation 19bbad48-bb75-497c-ae97-ff936e166272 · outbound

This paper cites Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.048415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.048415Z digest=sha256:37027e7f5c087ba66251a12d2e8b4c1f60def99a80ab5f91d68815e355e5dc2d

Observation ccf21cc0-c770-4c37-83a8-322785e006a9 · outbound

This paper cites XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.052089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.052089Z digest=sha256:598221065b576a6414ab0e3081d1ba608ca7b505e4cf51276c869ffc286472e0

Observation 65439416-3169-4241-99f4-6904b9f83c1c · outbound

This paper cites PickLLM: Context-Aware RL-Assisted Large Language Model Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques PickLLM: Context-Aware RL-Assisted Large Language Model Routing

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.056470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.056470Z digest=sha256:52449ac468ff6a52cb17ec8a2ac71de8f345d1a4ae6c36e6a3f94bd3a0c643c8

Observation 2c26cbc3-633d-40a4-8602-45546a00ee96 · outbound

This paper cites Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:56:57.138941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.062055Z digest=sha256:d8540eafa2ff262c23106e92fee0cbf9b28d42a0bfb9550b52caf62485bd195f

Observation f6fcb862-cb94-4226-a586-8dd499f3a109 · outbound

This paper cites Trusted llm inference on the edge with smart contracts,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Trusted llm inference on the edge with smart contracts,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.906148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:56:57.066381Z digest=sha256:e0dbeefecee9b81712fe550ca22caade912dd5e4147918f0e31f56e27720ca32

Observation 055c6e2d-3fc3-4bae-b046-10296140e601 · outbound

This paper cites Encryption-Friendly LLM Architecture.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Encryption-Friendly LLM Architecture

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.070193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.070193Z digest=sha256:c64822072fa83a3a61b3180ccf09ad3b5c66cc9a6edf875c0fdc9c7aa0d87194

Pith citing papers

Observation 51c91ec7-75e7-407b-bacb-1972153792d3 · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:25.790205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:25.790205Z digest=sha256:5b23505091eded3961aca2349f75743fc26e7c409ae97b81115300aa39dbe483

Observation 744745d6-6957-4b28-bb15-51ab1b51eaec · inbound

When Less is Enough: Efficient Inference via Collaborative Reasoning cites this paper.

When Less is Enough: Efficient Inference via Collaborative Reasoning Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:42.531986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T19:27:04.267404Z digest=sha256:d6032450572044eb3369da8ad6abd5cad80a00656fc539534f7462ed5f5cb728

Observation 34674ce5-e34d-40c4-97ff-f856eab0a7c5 · inbound

SOMA: Efficient Multi-turn LLM Serving via Small Language Model cites this paper.

SOMA: Efficient Multi-turn LLM Serving via Small Language Model Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:04.221906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:37:27.361504Z digest=sha256:0efd498ffe44c669f56c4f503c493ce0b4b04a89689447516973f6d605ece233

Observation 4f495410-efb5-4926-87f1-8dcf3557bdb5 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.605739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:caba6f932cb8490735a1a20918d6514bccd03e8a0c8e11913303540e4f823296

Observation 4fbda2fd-a0ef-45e3-b4ae-6be656850ee8 · inbound

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing cites this paper.

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.806454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T22:54:28.452796Z digest=sha256:b009ec6c8eb7c283f407ead1c5c7ec8ce92211e98a1e252d153889180f557b18