Pith. sign in

Paper Citation Record · LEDGER

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

As of 14 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.08239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08239 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:18:54.165282Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d36c36fa-dbd6-42eb-8e77-ef705d33ba8e · outbound

This paper cites Switchcraft: AI Model Router for Agentic Tool Calling.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Switchcraft: AI Model Router for Agentic Tool Calling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.674282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:18:54.095105Z digest=sha256:5c6804f83365fca25c33543f09a72680513e483f7113a80012e6c49e01dce5d9

Observation e5690e71-0c1a-4b56-83af-97dd95198134 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouterBench: A Benchmark for Multi-LLM Routing System

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.111313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.111313Z digest=sha256:e3d090cafc990200b8652f5387ef633383d3c3c3d2a5fd7b62f8aadf05fbd070

Observation acd73c12-76f2-472b-96cb-804c6265671b · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.116775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.116775Z digest=sha256:db004f8c5c9d1de03d8e0ba607b4346d2a40efb92ebc47557b6510156cd3234d

Observation e8f3a029-e178-4d69-88eb-09b24bd3b99e · outbound

This paper cites Universal Model Routing for Efficient LLM Inference.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Universal Model Routing for Efficient LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.121405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.121405Z digest=sha256:42f89547a147951fbfe147b8a6af43af58b7d58bb38d80a99bf7a7fbc26da752

Observation 90abd128-3dec-4217-9494-d7aa174c556d · outbound

This paper cites Routerarena: An open platform for comprehensive comparison of llm routers.arXiv preprint arXiv:2510.00202,.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Routerarena: An open platform for comprehensive comparison of llm routers.arXiv preprint arXiv:2510.00202,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.135081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.135081Z digest=sha256:05061e5d5cebf75729e064eeeb9cc53215a4771b3f2d2645dd6277cf2436a406

Observation 9c2ced0c-2fde-4a61-b8ec-754313d4dd25 · outbound

This paper cites Odar: Principled adaptive routing for llm reasoning via active inference.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Odar: Principled adaptive routing for llm reasoning via active inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.139092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.139092Z digest=sha256:eaeb21f194b6029f41be91b254120885f9e4968238c126f2cc4b412a65b898c1

Observation b5cfd66b-ea8c-41a7-9d32-a05164ad4746 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouteLLM: Learning to Route LLMs with Preference Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.143606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.143606Z digest=sha256:b77d2b22688687b4461c300d199cf2e3540ce7bd4bf5c1002db0f65e97420d21

Observation 5b8145ba-adce-4712-9048-a90929237e80 · outbound

This paper cites Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.149064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.149064Z digest=sha256:e6d05a9fa2ae245a2b273649ac2f9776c71b2a1c80700de46cc6f92df6261e7e

Observation 13092041-523f-4119-9c43-b263d655017b · outbound

This paper cites Qwen3 Technical Report.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.159441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.159441Z digest=sha256:fcf2860ed587c5c029f7fc8a902a7eb5470911526234eb71eb2a0243938286e4

Observation 80955fb7-5ad3-486f-902c-838badcf76d9 · outbound

This paper cites TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.165282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.165282Z digest=sha256:bcabd97c3b9c5e95ed7994b3ced6248148f3d6858c7a66ed74bf45b711e6a631

Observation 025e749e-fa89-470a-84e5-2907cecea79d · outbound

This paper cites Step-level Optimization for Efficient Computer-use Agents.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Step-level Optimization for Efficient Computer-use Agents

Reference 2016

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.240259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:18:54.154484Z digest=sha256:d595baa9df2f9ff4ee2aa9bfa59b0cbf8fa60b0c89daa6105e4cca8e72817902

Observation 042ff5c3-d06e-4ba3-928b-45ff557256b5 · outbound

This paper cites RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.573778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:18:54.126620Z digest=sha256:4055154e4db2529901541c8de6e851982a937d468bae27387eda63459fadb644

Observation bd504762-ced5-400c-9089-1abeb2024d29 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.106043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.106043Z digest=sha256:9e22dd0e541cd54acd70bec5e670fe8c75c512e2fb06d35bb38ec04850bf25b5

Observation c823f036-d0cf-4726-b02c-791109e2997b · outbound

This paper cites Session-aware agentic routing: Continuity- aware model selection for long-horizon llm agents.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Session-aware agentic routing: Continuity- aware model selection for long-horizon llm agents

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:18:54.690524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:18:54.131034Z digest=sha256:b373a7e7ecf19d4d3467b6db3d871a0b044e59d09a4076a8d92d9e67527adc81

Observation 4c00d2ec-c5f4-474b-b92c-311bbc0ed06e · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World AutoMix: Automatically Mixing Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.100605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.100605Z digest=sha256:6f6953d4db00d1de80e0af208bb5931d48e8975c5acb8cf8fa0ac7dbf53d0997

Pith citing papers

No inbound Pith citation observations are available.