Pith. sign in

Paper Citation Record · LEDGER

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.08160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08160 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:24:03.892023Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c978dad-8c8e-4668-9f37-f53f8112b1b7 · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Towards a Human-like Open-Domain Chatbot

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.785496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.785496Z digest=sha256:8896944b1edcd91fa0edbb7a39f209ab9f3e640f65eeebd29d881d15db4707e4

Observation c6a18712-fc71-414b-ad6f-898742112cba · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Prometheus: Inducing fine-grained evaluation capability in language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.357075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.802957Z digest=sha256:1782efede8d728d48e1b2b971c526eb4cbd38c7dca9d66f9ec678c830beec481

Observation 36f5f8fc-770d-46e7-afd7-39ad98dcd032 · outbound

This paper cites Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.341954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.808280Z digest=sha256:8b0f59f04a6c41ccb0c25db58525e5ea51f69b3870acc7ffefab2d9159462f49

Observation 68b75013-c94b-4c02-ac03-8690e3070cb1 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.813314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.813314Z digest=sha256:c262675cea7122552043558b219c923961a65463f629541cf2bfbec49b517734

Observation 1607fd6b-28bb-41ab-a7b5-e57cda87559e · outbound

This paper cites Player-driven emergence in llm-driven game narrative.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Player-driven emergence in llm-driven game narrative

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.308832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.823848Z digest=sha256:7d62775bd5586d27340dc904bb8d8dc796b8e26c1d48cbf00db8849feec11ca5

Observation 35975542-e343-4935-81f1-f10bde27460e · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives LaMDA: Language Models for Dialog Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.860945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.860945Z digest=sha256:88be70316d9369a211e39333da41d78da3f1b6d8e412224f60ba0512d4dc8628

Observation b551a7ed-98ce-4ea7-8349-3f2efffd131a · outbound

This paper cites Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.199992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.872129Z digest=sha256:7a36fbcd3b696fd08dd53f832a61dcb3be2d6b1e32e5d0ada04790bd16f893b4

Observation 0440e59e-1a60-4e6d-ab2d-0535c010162d · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Agentless: Demystifying LLM-based Software Engineering Agents

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:24:03.876817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.876817Z digest=sha256:de9f89972226dfbbd1028413b9fe367340db2d0d4ff81702c1174a78b1a4b763

Observation 4b20e77d-dc52-46a8-a0c7-f6ffda46cd4c · outbound

This paper cites Qwen3 Technical Report.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.881671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.881671Z digest=sha256:8576acc7faf0452967d4959ca1f5bedab83efe4ad8ac46305a650172c2e40cc0

Observation bec70c5d-a753-4fa3-96f0-94225235b708 · outbound

This paper cites Score: Story coherence and retrieval enhancement for ai narratives.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Score: Story coherence and retrieval enhancement for ai narratives

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.887012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.887012Z digest=sha256:4c1003923ed2476589f5f424aeebcc2bff581f0cb9a2690fccefe86bc5ed1011

Observation 134ab024-c815-485d-b717-19c6318be0a9 · outbound

This paper cites Interesting.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Interesting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.182847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.892023Z digest=sha256:ce7e93f5956afc33b6a23ab06ac221fd7209a5e1b63e81a9f0fa7dcc7f7103e2

Observation ba62213a-19f4-414a-bb3f-09d14ba10607 · outbound

This paper cites Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.325603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.819040Z digest=sha256:89b086eac1875aa2ce06b3f88f975ce02018dbc50baf167cc709bd12eecab3ee

Observation d4cb5e96-c8f0-437a-9a00-e67905585a6a · outbound

This paper cites What makes a good conversation? how controllable attributes affect hu- man judgments.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives What makes a good conversation? how controllable attributes affect hu- man judgments

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.257169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.839707Z digest=sha256:a21e0c0a27fe19d5c5a419d6c57f9bedc04330ac881a21058d324427aee850d6

Observation 54b0f562-912b-44ee-b09f-f9bd56f901b5 · outbound

This paper cites Plot- machines: Outline-conditioned generation with dynamic plot state tracking.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Plot- machines: Outline-conditioned generation with dynamic plot state tracking

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.275910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.833990Z digest=sha256:fc3b0fae7bc590181db4e4559bce4ff11249780b0266ca41adb4bc63047abab8

Observation c7051bb3-133e-4697-b26d-bb735d28ebf8 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Kimi K2.5: Visual Agentic Intelligence

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.850635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.850635Z digest=sha256:3d0238d52a90dafddc838797a41b969e72ddc39493d78ecf65e113ae3073b917

Observation 8425bebe-4687-4561-8b98-44732b6b51ad · outbound

This paper cites Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.844992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.844992Z digest=sha256:82bf589af8a6f0d54f715e27b85e3ced5fa4fbb15dbccaffa75b51bfccacef45

Observation b50cad13-f1d4-4bbc-ae47-7c4e9bea53ef · outbound

This paper cites Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.222720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.867337Z digest=sha256:c6e684d212e48abaa4c377e0c5edd13c41d538ed916f81385b3811331057de84

Observation a74d2641-a08f-4694-a266-9d1200787481 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.791845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.791845Z digest=sha256:2e7978e90e120d7a55c151df164fc1928a4c1ca21224605ce7f2d69f35ac7593

Observation 8743f3e3-b051-4735-9fe1-29456f6b55b6 · outbound

This paper cites Red teaming language models with language models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Red teaming language models with language models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.292768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.828593Z digest=sha256:ce20907179e66e8d096e0bae0599ac79c91a7b5976415a0f1d7c01ee155914c9

Observation 5464d4e5-f611-466d-8cb5-7484c48e3df1 · outbound

This paper cites GPT-4o System Card.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives GPT-4o System Card

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.797463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.797463Z digest=sha256:021bd7c542ee5b46418024e9e8e0333aa94f3b6d2a3a429d2405a76df02bd264

Observation b3ec1cb2-5686-4179-a1b6-ab08ad2558fa · outbound

This paper cites T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.239975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.856261Z digest=sha256:920ee11d72d27929a2e95b84d309fc3ecd4a73d3c685a1f221acb547915bcf42

Pith citing papers

No inbound Pith citation observations are available.