Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:24:03.892023Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.08160.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:24:03.892023Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6c978dad-8c8e-4668-9f37-f53f8112b1b7 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Towards a Human-like Open-Domain Chatbot
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a18712-fc71-414b-ad6f-898742112cba · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Prometheus: Inducing fine-grained evaluation capability in language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 36f5f8fc-770d-46e7-afd7-39ad98dcd032 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 68b75013-c94b-4c02-ac03-8690e3070cb1 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1607fd6b-28bb-41ab-a7b5-e57cda87559e · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Player-driven emergence in llm-driven game narrative
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35975542-e343-4935-81f1-f10bde27460e · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives LaMDA: Language Models for Dialog Applications
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b551a7ed-98ce-4ea7-8349-3f2efffd131a · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Wang, L., Lian, J., Huang, Y ., Dai, Y ., Li, H., Chen, X., Xie, X., and Wen, J.-R
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0440e59e-1a60-4e6d-ab2d-0535c010162d · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Agentless: Demystifying LLM-based Software Engineering Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b20e77d-dc52-46a8-a0c7-f6ffda46cd4c · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Qwen3 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec70c5d-a753-4fa3-96f0-94225235b708 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Score: Story coherence and retrieval enhancement for ai narratives
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 134ab024-c815-485d-b717-19c6318be0a9 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Interesting
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba62213a-19f4-414a-bb3f-09d14ba10607 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Self- contradictory hallucinations of large language models: Evaluation, detection and mitigation
Reference 2003
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4cb5e96-c8f0-437a-9a00-e67905585a6a · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives What makes a good conversation? how controllable attributes affect hu- man judgments
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 54b0f562-912b-44ee-b09f-f9bd56f901b5 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Plot- machines: Outline-conditioned generation with dynamic plot state tracking
Reference 2011
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7051bb3-133e-4697-b26d-bb735d28ebf8 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Kimi K2.5: Visual Agentic Intelligence
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8425bebe-4687-4561-8b98-44732b6b51ad · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50cad13-f1d4-4bbc-ae47-7c4e9bea53ef · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Are large language models capable of generating human-level narratives? InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a74d2641-a08f-4694-a266-9d1200787481 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8743f3e3-b051-4735-9fe1-29456f6b55b6 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Red teaming language models with language models
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5464d4e5-f611-466d-8cb5-7484c48e3df1 · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives GPT-4o System Card
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ec1cb2-5686-4179-a1b6-ab08ad2558fa · outbound
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives T., Liu, H., Liu, T., Wang, C., Liu, T., Zhang, Y ., Shipman, F., et al
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.