Pith. sign in

Paper Citation Record · LEDGER

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.08160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08160 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:24:03.892023Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c978dad-8c8e-4668-9f37-f53f8112b1b7 · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Towards a Human-like Open-Domain Chatbot

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.785496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.785496Z digest=sha256:f563a054e0d8ca5301a079ad51d20e26e236484682989f5ad541d25c91598f55

Observation c6a18712-fc71-414b-ad6f-898742112cba · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Prometheus: Inducing fine-grained evaluation capability in language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.357075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.802957Z digest=sha256:a916e0995d03576de7574d750c4251809364da7056ec05d3d597fd4f0a45c1f9

Observation 36f5f8fc-770d-46e7-afd7-39ad98dcd032 · outbound

This paper cites Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.341954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.808280Z digest=sha256:d7559b624d21b8e61b48657b687d34c3a8648a57f1e39329e5c77be1b2b7d3e3

Observation 68b75013-c94b-4c02-ac03-8690e3070cb1 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.813314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.813314Z digest=sha256:fc855a34563fb63dadeef947de3c7531467e5c1bfbf287433f3154ed438156d7

Observation 1607fd6b-28bb-41ab-a7b5-e57cda87559e · outbound

This paper cites Player-driven emergence in llm-driven game narrative.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Player-driven emergence in llm-driven game narrative

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.308832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.823848Z digest=sha256:ffb6bb97d3d1a362f160e0ab34e03969f6f89d280bb4f4f682a3dbfc646cf455

Observation 35975542-e343-4935-81f1-f10bde27460e · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives LaMDA: Language Models for Dialog Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.860945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.860945Z digest=sha256:b460772b6f856f67e97553e77fe40a596a38df3edd7838f16c8ea86bbc12d430

Observation b551a7ed-98ce-4ea7-8349-3f2efffd131a · outbound

This paper cites Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.199992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.872129Z digest=sha256:c3e7f3a85d433bbcac321e7b80bb500c1d3d442654aa0cc2f90bf90fad853c1c

Observation 0440e59e-1a60-4e6d-ab2d-0535c010162d · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Agentless: Demystifying LLM-based Software Engineering Agents

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:24:03.876817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.876817Z digest=sha256:c8875a8a772ce937ea55cbf98d6752e99551f575f5d0e7e59b90b78609ecb02c

Observation 4b20e77d-dc52-46a8-a0c7-f6ffda46cd4c · outbound

This paper cites Qwen3 Technical Report.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.881671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.881671Z digest=sha256:05cbbd2d9e3c4dfcf066e0102626e95099959a51a0c74671d05e2502f036e094

Observation bec70c5d-a753-4fa3-96f0-94225235b708 · outbound

This paper cites Score: Story coherence and retrieval enhancement for ai narratives.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Score: Story coherence and retrieval enhancement for ai narratives

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.887012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.887012Z digest=sha256:630cec0c8d9f4d4d5619de1c2a00cf703ec9ccbfbd9edfc310907a171395d9ae

Observation 134ab024-c815-485d-b717-19c6318be0a9 · outbound

This paper cites Interesting.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Interesting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.182847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.892023Z digest=sha256:7e0115f40e524d67a7cafcd7b5bcc9b50cb5482ebfdf7c63f88eab5215718449

Observation ba62213a-19f4-414a-bb3f-09d14ba10607 · outbound

This paper cites Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.325603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.819040Z digest=sha256:dce17ce655bea3daf04ee5c2ee0232a47dcf1f0c5584d9dfb802694d708b33da

Observation d4cb5e96-c8f0-437a-9a00-e67905585a6a · outbound

This paper cites What makes a good conversation? how controllable attributes affect hu- man judgments.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives What makes a good conversation? how controllable attributes affect hu- man judgments

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.257169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.839707Z digest=sha256:6dc5fd127f7f9bd39793d642cdbc254e8ec653d9513e64523981550ff1dbb80e

Observation 54b0f562-912b-44ee-b09f-f9bd56f901b5 · outbound

This paper cites Plot- machines: Outline-conditioned generation with dynamic plot state tracking.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Plot- machines: Outline-conditioned generation with dynamic plot state tracking

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.275910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.833990Z digest=sha256:9405c462c3bf7537d3da47f050914900b509e23a5d50242eb68eea9de8cb3669

Observation c7051bb3-133e-4697-b26d-bb735d28ebf8 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Kimi K2.5: Visual Agentic Intelligence

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.850635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.850635Z digest=sha256:299077af1aab9bce05280ac966e58f4280fcc458f4acacefdb727b457e6e9238

Observation 8425bebe-4687-4561-8b98-44732b6b51ad · outbound

This paper cites Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.844992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.844992Z digest=sha256:2d6378f1ef78970f919b476408fa4991f128686ce1a2d65f9cf19a094f37333e

Observation b50cad13-f1d4-4bbc-ae47-7c4e9bea53ef · outbound

This paper cites Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.222720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.867337Z digest=sha256:1039ccff2a98933005d3e5158312d0e6da6c0639bc9de57c0245898d53db572f

Observation a74d2641-a08f-4694-a266-9d1200787481 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.791845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.791845Z digest=sha256:c9ccf715b22932544e6395356532f0e4046b0e7b04017d10948cafa84affc840

Observation 8743f3e3-b051-4735-9fe1-29456f6b55b6 · outbound

This paper cites Red teaming language models with language models.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Red teaming language models with language models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.292768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.828593Z digest=sha256:9eb1a65b82b8910d59cf086cf15cb3915a4ac3388d786fcf82848f963ccfc3d0

Observation 5464d4e5-f611-466d-8cb5-7484c48e3df1 · outbound

This paper cites GPT-4o System Card.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives GPT-4o System Card

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:24:03.797463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:24:03.797463Z digest=sha256:29849816a7c1f644c38afd4ec627629307c012b3b3d8c0cbe9860ad17e35a67e

Observation b3ec1cb2-5686-4179-a1b6-ab08ad2558fa · outbound

This paper cites T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al.

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:24:04.239975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:24:03.856261Z digest=sha256:4162afd1fd0e3eead751e713d72ceb3a7d827482d139526edf6ff9b3d5dbde50

Pith citing papers

No inbound Pith citation observations are available.