Pith. sign in

Paper Citation Record · LEDGER

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 9 inbound Pith citation observations for arXiv:2506.13356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13356 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:08:12.087697Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:36:05.751022Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T02:29:25.233098Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact3
  • verified fuzzy30
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69f57222-94a4-4327-b27b-0aed18bfc6a9 · outbound

This paper cites L-eval: Instituting standardized evaluation for long context language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns L-eval: Instituting standardized evaluation for long context language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.769044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.915064Z digest=sha256:d17a49a34f8af9964a178768eddd2575515b81372d67e2b0b4ca96808f05f3c3

Observation 7ed84ab1-713b-4448-b30b-14f2e79ef3aa · outbound

This paper cites Claude 3.5 sonnet model card addendum, 2024.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Claude 3.5 sonnet model card addendum, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.756334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.920323Z digest=sha256:b8cfe2d088792ccde7d3941ed845f105a9b702ef158f766bc0e1b985eb8dcfde

Observation 1abb18d2-cf36-4687-989d-71b09d92895e · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Longbench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.744342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.923913Z digest=sha256:419cbae5cc5c3d3ce79ae39684f2da3454410cba5b968c94e156651b27aded2a

Observation e4301444-94b0-4798-8966-9f82e87aea12 · outbound

This paper cites Opportunities and challenges for ai-based analysis of rwd in pharmaceutical r&d: A practical perspective.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Opportunities and challenges for ai-based analysis of rwd in pharmaceutical r&d: A practical perspective

Reference 4

Resolution
verified exact
doi, observed 2026-08-15T20:08:12.255655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.928097Z digest=sha256:9a921e17dfa71ad3167d6c2f01c9f19e48fe19d910fc189402493674ac65f9cb

Observation 52f1139f-f545-4249-be2d-6553fd5ce5d3 · outbound

This paper cites Longformer: The Long-Document Transformer.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Longformer: The Long-Document Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.932113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.932113Z digest=sha256:0848ea01e4007ee0777d61d521f95911c07128f5093ade3efe9a84674f231752

Observation 64c3555e-fdfc-4410-b5b2-7320115c2b16 · outbound

This paper cites Doubao-1.5-pro, 2025.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Doubao-1.5-pro, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.732354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.935671Z digest=sha256:15098bc413d70862afdb9a179c6ee3d5844988fef229c3c9bcaeca4c9f6bd6f8

Observation 76bb980a-89b2-4566-ace6-fe53e7a76961 · outbound

This paper cites Beyond prompts: Dynamic conversational benchmarking of large language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Beyond prompts: Dynamic conversational benchmarking of large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.720802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.939133Z digest=sha256:798efe2849653604d9d2f349b590f4b57bb11bfbd5bf7668beff08834bc39d0f

Observation 75025580-5c20-40bd-a85a-598cf664da1f · outbound

This paper cites A survey on evaluation of large language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns A survey on evaluation of large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.942410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.942410Z digest=sha256:9e8058bd9f8d60e55b03787fca7bc854caaca02a54fe84327150a8c72eb0cd51

Observation 30825cda-7ffa-477b-9936-6531273d6041 · outbound

This paper cites HSR-Enhanced Sparse Attention Acceleration.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns HSR-Enhanced Sparse Attention Acceleration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.945104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.945104Z digest=sha256:09fa9db75be47c5f8c81e5126ed9321cbf507283c080d570216623020581bee4

Observation c8f64c8d-fabf-4ddc-8e22-cc4ad1adbe3c · outbound

This paper cites Llf-bench: Benchmark for interactive learning from language feedback.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Llf-bench: Benchmark for interactive learning from language feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.702509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.948517Z digest=sha256:a45ecce89f5a0807f6492d75e090d293d3eac7c15a4a4f13f6f56e190a25f46b

Observation 1d8dc7ee-df25-400d-871d-6d960ceaa94b · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.951699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.951699Z digest=sha256:122ff36c1b10792604bcff7d3a380059d3ba39782495481068eb52cf5be1caac

Observation b7030dd1-ff99-4c8f-8cc3-f9f7b72fb9a3 · outbound

This paper cites Generating long sequences with sparse transformers.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Generating long sequences with sparse transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.691205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.955423Z digest=sha256:44ad917a6e548ed6f46a8cf692e07d8f7e97075502181855803bada284e50592

Observation fbb4c577-d707-491e-9575-d4ee2ddda1cb · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.958869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.958869Z digest=sha256:d09a99b4eb9d7066befa4185744c4387184a1c8ab6b77f5ccb8d49d4b9588f9c

Observation 20c3f5dd-f178-4685-9bb3-9d23f8578e9d · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.680009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.962474Z digest=sha256:aedb7a345d1d3e7325e4d372a463bdf0a8eca102f2f348e53f9c216a55a126b8

Observation b2c98317-7572-46b2-aef7-3745b43a8fa7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.965879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.965879Z digest=sha256:1fe42c5e3bf28801e4dbc24d4d64989ec272fc437484ca4473dd8db7340a9415

Observation 9f1e6352-ea59-40ed-9f65-c7d6e524c7f0 · outbound

This paper cites Cervical cancer segmentation based on full-scale feature fusion with cascading-attention and dilated convolution.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Cervical cancer segmentation based on full-scale feature fusion with cascading-attention and dilated convolution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.667483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.969800Z digest=sha256:4bced5883e136c49c710b47bcffbfa51055445210aa9608514b03686a5d36a2a

Observation 038ec467-1243-4d5d-a0f7-844cf694a475 · outbound

This paper cites BAMBOO : A comprehensive benchmark for evaluating long text modeling capacities of large language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns BAMBOO : A comprehensive benchmark for evaluating long text modeling capacities of large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.655732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.973152Z digest=sha256:6feaa86e6e200687bd0e1b2795e1725b4bff296c06139ed58228fcdbeab6bd8e

Observation 33be54da-4aa4-4852-a1f0-5f1157324d84 · outbound

This paper cites Knowledge retention: 8 main strategies to improve it, February 2024.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Knowledge retention: 8 main strategies to improve it, February 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.644375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.976644Z digest=sha256:6076f2ecf560023456ea56d632996e019d8be1e659824b4fd63872793ddfaa26

Observation 622a8289-afee-4a07-a1d1-4f732e5a7fdf · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Mamba: Linear-time sequence modeling with selective state spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.979699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.979699Z digest=sha256:c8fd30492f36da59b91f212a8a3d6675e739fd569389a16d35db47cbb2776888

Observation 0549d748-8eb2-48c3-8c70-8efd52c650c9 · outbound

This paper cites MAPLE: A Mobile Agent with Persistent Finite State Machines for Structured Task Reasoning.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns MAPLE: A Mobile Agent with Persistent Finite State Machines for Structured Task Reasoning

Reference 20

Resolution
verified exact
doi, observed 2026-08-15T20:08:12.233969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.982835Z digest=sha256:bb52a686f0bba984dc7cdc9134cf2e27b1b8c3cb1e0f103eb554a3c3c464fd78

Observation f645c181-9bad-4323-9966-509ca3f0f567 · outbound

This paper cites Hipporag: Neurobiologically inspired long-term memory for large language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Hipporag: Neurobiologically inspired long-term memory for large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.626692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.985766Z digest=sha256:a916f0fa58462b166c8280cb3ea8fcb4fca869963520799e7e7e6c6c888489c8

Observation 4811842e-b86c-4b43-ab18-0619868e0c43 · outbound

This paper cites Layer-adaptive low-rank adaptation of large asr model for low-resource multilingual scenarios.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Layer-adaptive low-rank adaptation of large asr model for low-resource multilingual scenarios

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.615869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.990084Z digest=sha256:abfa95c0220a5d9c870b6469e62008bd579fa753752100b9f23e69041b165ad5

Observation 517d0326-d456-4670-902e-84f9fccaddff · outbound

This paper cites Synthetic Data in AI: Challenges, Applications, and Ethical Implications.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Synthetic Data in AI: Challenges, Applications, and Ethical Implications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:11.993263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:11.993263Z digest=sha256:7e3c4e72c70f04beebfd3901167d457978dc59c3817426f619f62f146bfebd5e

Observation 5b6804b3-1f29-4533-9e90-c7150b8a4141 · outbound

This paper cites Ruler: What's the real context size of your long-context language models? CoRR, 2024.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Ruler: What's the real context size of your long-context language models? CoRR, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.604910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:11.996846Z digest=sha256:bc63cb669daa7cf3190447e31f6b87ccb357aac9d96dfe060a8fbe80940a14e4

Observation 53b4a1ad-deb7-4bdb-9acd-c98d0990871b · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Needle in a haystack - pressure testing llms, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.000185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.000185Z digest=sha256:0f07ba2a62861172b28847980d52f97cdd2a738d5539b0c1efce8d99403ca26c

Observation 12c40333-58d1-4706-a49c-a65dd8ae1b6e · outbound

This paper cites Reformer: The efficient transformer.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Reformer: The efficient transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.003410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.003410Z digest=sha256:e0fc348a2a81f5082b8abb4d686ebe55ed6c67a6a95b7cfa2edfbfefc440ae6e

Observation 3f79d0d3-3e9a-48ed-b760-31b958a054c8 · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Babilong: Testing the limits of llms with long context reasoning-in-a-haystack

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.580116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.006634Z digest=sha256:7cef4282a296bdc8b725266d88d9e9d1b4d57b87d12bbdda63c7e4ae8999a443

Observation 6d22d922-771d-4387-a45c-df86270cba5c · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Efficient memory management for large language model serving with pagedattention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.010674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.010674Z digest=sha256:5c3c59b032e672875fbf6f510bb38f41e8eb1e6d134e31bbb5e8858ba4062a1b

Observation 880da2a9-324a-4cd8-9e0a-d6ce8c8139e8 · outbound

This paper cites A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.013881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.013881Z digest=sha256:6c8eb0f5a0326dbfef3038739cb1cc3f64fd7c1190a26735e597ed194160c032

Observation a33f96fb-ac34-4107-8291-a5f15609a4fb · outbound

This paper cites Loogle: Can long-context language models understand long contexts? arXiv e-prints, pages arXiv--2311, 2023.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Loogle: Can long-context language models understand long contexts? arXiv e-prints, pages arXiv--2311, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.561384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.017573Z digest=sha256:085a3123280ce56e467cdd86fb4375e5efa0ce38cad3df7866172ea729a376c8

Observation 3867c032-6a6e-4bd8-82b1-7d5e96f31564 · outbound

This paper cites Ringattention with blockwise transformers for near-infinite context.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Ringattention with blockwise transformers for near-infinite context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.551788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.021339Z digest=sha256:dcf7a243966bb95bd0ce503c61643e57e185a43c08b0f556188b814b5c55f843

Observation 942a7a22-6870-4fd7-8cb8-6a825ee71d37 · outbound

This paper cites Best Practices and Lessons Learned on Synthetic Data.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Best Practices and Lessons Learned on Synthetic Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.025044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.025044Z digest=sha256:ea1631864af626e1030410735d49dbcc85f7dd8c711151642dd367c8f02a0248

Observation fa3a05e5-cda4-461e-9e73-4deb49414c07 · outbound

This paper cites Agentbench: Evaluating llms as agents.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Agentbench: Evaluating llms as agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.541685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.029113Z digest=sha256:0322f45d53af476a91c3fb7f920ae65f6de0dbe795d4e98d16a1e7029cb402dc

Observation e87b24f4-fa0f-449a-aec4-298bc9af453a · outbound

This paper cites Hello gpt-4o, 2024.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Hello gpt-4o, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.033418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.033418Z digest=sha256:f5026153974ae7ace060f3120c37014db188f25afbe2fb82c1f10add07d2acb4

Observation df2a3246-c07b-4e3d-925f-43b818c3d8ab · outbound

This paper cites Memgpt: Towards llms as operating systems.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Memgpt: Towards llms as operating systems

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.524801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.037587Z digest=sha256:6f23dacdd2e0afa1daffad13828eac7edc45a7dcf389188ae924664fc3f31e00

Observation 08278005-64cb-42c2-88bc-c5ecf46618e7 · outbound

This paper cites Faster causal attention over large sequences through sparse flash attention.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Faster causal attention over large sequences through sparse flash attention

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.514056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.042114Z digest=sha256:1802e4493b2241195d5aba774232aaab132352d009433f4af780b0774a2d3e11

Observation ce6c93a0-232e-4884-a3ae-0ee440bd6132 · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Rwkv: Reinventing rnns for the transformer era

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.503812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.045250Z digest=sha256:48a85f5ff421925276d7eaea547dfbfb85c0b08d6df9a7dab3a6087073e710c8

Observation 3f795acf-3c7c-48f9-a558-e40fa83addef · outbound

This paper cites Do transformers need deep long-range memory? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7524--7529.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Do transformers need deep long-range memory? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7524--7529

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.396431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.048708Z digest=sha256:1e01c4a44867ada02f6ed9d346b6f28adf11166e0351505f6f5fd94cef0a55a9

Observation ce696ee4-53fe-4da3-82a4-6c0ad90444a4 · outbound

This paper cites Zeroscrolls: A zero-shot benchmark for long text understanding.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Zeroscrolls: A zero-shot benchmark for long text understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.385792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.052059Z digest=sha256:a7da1be940e39116e748fcbb598caf59317b568d18dce98d093eee3a353891d5

Observation 3f7e7546-8de6-4ca4-8df8-851fdb57839f · outbound

This paper cites Cognitive Memory in Large Language Models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Cognitive Memory in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.055184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.055184Z digest=sha256:4827b1374411f9aa15979a25e34b1f65f372a2f340e169738b3cb5ae6f2e2fdc

Observation ba3fd16b-0a37-4b3e-a28b-d8457f571f12 · outbound

This paper cites Chapterbreak: A challenge dataset for long-range language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Chapterbreak: A challenge dataset for long-range language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.374814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.058341Z digest=sha256:a83f6b6efc2a48dbeff3ee7a87682c45b677facb56db31de1b71e8d1e1fe4cf5

Observation 1235262d-f1f6-4c76-a6ec-553d096a8e14 · outbound

This paper cites From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:12.062760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:08:12.062760Z digest=sha256:8dd7fceb54cf3ae8ed938b7e1a6cba31a2a6af91edb55af51e8d735e113bcf74

Observation e8988212-e8e1-433f-80ae-323fbf6efeb9 · outbound

This paper cites Data and System Perspectives of Sustainable Artificial Intelligence.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Data and System Perspectives of Sustainable Artificial Intelligence

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:08:12.123172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.066826Z digest=sha256:03e52f6ca37cef85663845097e7add6618e77028952558fec0ba2733464e46f0

Observation 198e470d-761b-4afa-b4b1-1a2c2b384e8e · outbound

This paper cites Long time no see! open-domain conversation with long-term persona memory.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Long time no see! open-domain conversation with long-term persona memory

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.364239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.071189Z digest=sha256:79ebd3effae6ac6cdc12848890cc25500a23dfcfe063b64852fcbf423620b61d

Observation 8a92f438-4eb9-4409-9cf8-14a5e02d4c28 · outbound

This paper cites Memoryscope, 09 2024.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Memoryscope, 09 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.352260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.074534Z digest=sha256:9dcf4908976d469449eabe33091ab2412ad5c8f4b6ed07e1b751966d7b2ecd27

Observation 8e3d5258-fa73-4fea-a4b5-b8d4143315d5 · outbound

This paper cites A survey on multi-turn interaction capabilities of large language models.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns A survey on multi-turn interaction capabilities of large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.342056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.077663Z digest=sha256:5525c308dc794cf58bed110f4933eb25592f36988a5fcabf65ea50ccd8cb6152

Observation efe79906-4df8-4582-8fc7-89ca9fefdd50 · outbound

This paper cites bench: Extending long context evaluation beyond 100k tokens.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns bench: Extending long context evaluation beyond 100k tokens

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.329481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.080639Z digest=sha256:529bab73f5c9f0c0256c6e1455e5a03775c987c21c4834ab71cb3fdd37766555

Observation 7b4c4dbd-2ba5-4df7-81b8-e95e555afd8c · outbound

This paper cites Memorybank: Enhancing large language models with long-term memory.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Memorybank: Enhancing large language models with long-term memory

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.319299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.083371Z digest=sha256:70df1d2abd4120b3b96afb2065a2f2104fc89bb132d9ac7d1a8d63acd3fce9b9

Observation e16ef2a4-d38e-4a14-a8d8-322a289bf299 · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents.

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Webarena: A realistic web environment for building autonomous agents

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:08:12.308084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:08:12.087697Z digest=sha256:37771335a00fb1a7f52c9b5af9c88aa6bb8d160282d7ca8c1b53ef9d9bd13f9a

Pith citing papers

Observation fc514041-df33-48a7-8dc8-e135856eddf9 · inbound

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions cites this paper.

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:20:22.260588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T21:20:22.146417Z digest=sha256:e2da39ea0a92dc2a14c2ffe474203ffba3667136bcb681d36af7b3344f000df2

Observation 99af9aae-2f5e-47a5-a31e-b252d9768f55 · inbound

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions cites this paper.

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:36:05.751022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:36:05.751022Z digest=sha256:ffa53275df2d29be3efb0d81113af445416379348d17c1f3aa1dd766dc3ca87d

Observation f28ed759-8378-4802-89c6-4dda3ee488f0 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.770267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:8111eca5b67943cf625749b41510719dc5c7362d112c15708fc1f539f713c3ac

Observation 9f37cf52-1f65-4a4c-b070-763b30f745ef · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:16.118967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:dc5cf2a7506b858a82ddfcc37c5f7c8417d8aa6074f78587621a66338fe74dad

Observation b1c0d3c9-3af2-4d47-92cc-23664cac1401 · inbound

Cost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large Language Models cites this paper.

Cost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large Language Models StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T11:01:43.577854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:01:43.577854Z digest=sha256:6ab342711b24fcb7230f7565160f19ab5502f45b84e0ce38ee7b64374715d5e1

Observation 7448e09e-0062-44f7-86b9-ba38d68b59c2 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:42.468772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:42.468772Z digest=sha256:b987a2b686a7d81a929260905ef041f46da00026294c3be479d2f580920ee677

Observation eb03873c-da9d-4f31-93b8-ae173f8e4696 · inbound

From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms cites this paper.

From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:40:58.950764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:10:12.281373Z digest=sha256:65e0502b0686a19bf9f566d0f3512d0f467eebdd373704174b9a42a974b0fc9b

Observation 532bf676-1e9a-4949-b3dd-e4ac8cf8a4b3 · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.465188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:22b5448a6140d7c3aac1ce9d1c51f3060b97ace74741c1de3080f0b55c500518

Observation daa066a6-8d23-49d6-95ea-e2a59ac9b31d · inbound

MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts cites this paper.

MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:29:25.236761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T02:27:03.246763Z digest=sha256:59ad63f3bad134e1552548469b5c3589f935d8ca45285c9b973d4d2069fde99e