Pith. sign in

Paper Citation Record · LEDGER

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.08239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08239 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:18:54.165282Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d36c36fa-dbd6-42eb-8e77-ef705d33ba8e · outbound

This paper cites Switchcraft: AI Model Router for Agentic Tool Calling.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Switchcraft: AI Model Router for Agentic Tool Calling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.674282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:18:54.095105Z digest=sha256:b46a49ceb69ca8cecba9c843bddba8bf36ec81bd702687635eb286c616ee2511

Observation e5690e71-0c1a-4b56-83af-97dd95198134 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouterBench: A Benchmark for Multi-LLM Routing System

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.111313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.111313Z digest=sha256:0db68fa973bec76df8778cc256e5984c589f4617845d09b5ca63ffadf479a102

Observation acd73c12-76f2-472b-96cb-804c6265671b · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.116775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.116775Z digest=sha256:559c94fa57a32e187683d08abbe1529d2b66de152a2c8e4f14b7dc4674521167

Observation e8f3a029-e178-4d69-88eb-09b24bd3b99e · outbound

This paper cites Universal Model Routing for Efficient LLM Inference.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Universal Model Routing for Efficient LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.121405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.121405Z digest=sha256:ab09b9d896f6452b60e6309d4d0f4c3b6c4a04c9a291eb4645caac0621bdadd8

Observation 90abd128-3dec-4217-9494-d7aa174c556d · outbound

This paper cites Routerarena: An open platform for comprehensive comparison of llm routers.arXiv preprint arXiv:2510.00202,.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Routerarena: An open platform for comprehensive comparison of llm routers.arXiv preprint arXiv:2510.00202,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.135081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.135081Z digest=sha256:1281cc67a8f257bdc141424e2f6e7b965d99def525941fcf59f4406d338cd1e4

Observation 9c2ced0c-2fde-4a61-b8ec-754313d4dd25 · outbound

This paper cites Odar: Principled adaptive routing for llm reasoning via active inference.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Odar: Principled adaptive routing for llm reasoning via active inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.139092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.139092Z digest=sha256:16a46fe66ddecaf672f69c1a2fe0f2926e3d5f54ad40aaa97fdc15383ab7dcbd

Observation b5cfd66b-ea8c-41a7-9d32-a05164ad4746 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouteLLM: Learning to Route LLMs with Preference Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.143606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.143606Z digest=sha256:367a1eb58d59f19360663361356dc32594a6f4c8ed7a880e9742e6aff12abbe4

Observation 5b8145ba-adce-4712-9048-a90929237e80 · outbound

This paper cites Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.149064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.149064Z digest=sha256:8a4fc8a72b9674ce13a2ce481ff350605e7db4e340b00a4a110c259969e91102

Observation 13092041-523f-4119-9c43-b263d655017b · outbound

This paper cites Qwen3 Technical Report.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.159441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.159441Z digest=sha256:34e8f5db960534eff004ae3f7cd5048328cdfc5237c0e4ce63b66b1c0c07f54d

Observation 80955fb7-5ad3-486f-902c-838badcf76d9 · outbound

This paper cites TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.165282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.165282Z digest=sha256:0dbb58f1ab61b90f31646badf67e10feae9d5ab86149786b4b9825f9cb75d022

Observation 025e749e-fa89-470a-84e5-2907cecea79d · outbound

This paper cites Step-level Optimization for Efficient Computer-use Agents.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Step-level Optimization for Efficient Computer-use Agents

Reference 2016

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.240259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:18:54.154484Z digest=sha256:0313f6a21507bb1695d0dbb0ea3a525b03b6331426955c369a1123c302a3c8ad

Observation 042ff5c3-d06e-4ba3-928b-45ff557256b5 · outbound

This paper cites RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.573778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:18:54.126620Z digest=sha256:52860303508861c82f0fd36eecc5462360eb4c114c738470776a32e4c0280a1d

Observation bd504762-ced5-400c-9089-1abeb2024d29 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.106043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.106043Z digest=sha256:4c07b360d45aef661898bdab5fcb5cacb68a0753cf2a035644dbe6d1a7161664

Observation c823f036-d0cf-4726-b02c-791109e2997b · outbound

This paper cites Session-aware agentic routing: Continuity- aware model selection for long-horizon llm agents.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Session-aware agentic routing: Continuity- aware model selection for long-horizon llm agents

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:18:54.690524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:18:54.131034Z digest=sha256:08662ce6bc982df5b18ff8da5fc728c55dfdd499967a899fb959999142229524

Observation 4c00d2ec-c5f4-474b-b92c-311bbc0ed06e · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World AutoMix: Automatically Mixing Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.100605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.100605Z digest=sha256:eb276a24b07047b32205b3015910ffc5c0d9f5e4bce837f9a79e1b6a0a2b95c9

Pith citing papers

No inbound Pith citation observations are available.