Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:26.736743Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2505.19293.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:26.736743Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:40:58.250915Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T05:00:21.895979Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7765fd1d-24eb-47ff-a883-0d7f005b270e · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7785cc7a-5b1f-475b-956d-a36804d325e9 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a986590-0b74-4e28-a975-fd73578628f2 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b688bf22-8169-4d5f-9dc2-cfb763875804 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9213b459-1029-4b1c-ba61-9c03bf6eb69b · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Many-Shot In-Context Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c728fe26-7aa8-4634-9e18-24b52d96d5b0 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? L-Eval: Instituting Standardized Evaluation for Long Context Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb476a2-5b16-4523-94d7-f7fd05e20bb2 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0daae9fc-557d-49ca-b9c5-826db961a409 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a821fbc7-3ddb-4320-aa9c-93c253bf5643 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Extending Context Window of Large Language Models via Positional Interpolation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46d8258-ff30-447f-9756-d2eb5aa8260e · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbac3db7-9e4c-4bc9-b544-60ac81a6badc · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23951161-ae45-4f9c-8cb6-b09ebe855f14 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b193c1-b023-4dfe-bef3-985d11c5bdc5 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? MedOdyssey: A Medical Domain Benchmark for Long Context Evaluation Up to 200K Tokens
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9a4335e-851d-4fc3-848f-3c771c57f95a · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff83b31-5369-4e9f-9c35-d957cb6dac6b · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b2b4f5d-42fe-479d-bbb2-971369faef7f · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa4703b7-b1d2-4a88-94c5-99b999c0c675 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8e0671-926e-4089-9ee3-c86c824530ca · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Liger Kernel: Efficient Triton Kernels for LLM Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 845c94eb-5dac-41e9-a23e-553281c6f0d1 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb754909-6e43-484e-a40d-a28a508b4009 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00993ed0-2ba8-433d-b8ba-24bacf7d5cc5 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LooGLE: Can Long-Context Language Models Understand Long Contexts?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b19a446-827a-4ff6-87aa-f8fe4cbc08b9 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Sequence Parallelism: Long Sequence Training from System Perspective
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c65d74-f6c6-43ae-b1db-8229c215db9e · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Compressing Context to Enhance Inference Efficiency of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f459c01-634f-4b2c-9482-3ea5afab542c · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63264196-1f47-4755-9156-30de23677639 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 230c6493-93ac-4731-b6aa-60f2fcce3d7e · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50486a72-61a9-48ce-ac92-c1cbedf19cef · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? A Controlled Study on Long Context Extension and Generalization in LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4476c01e-0511-439f-a552-e9eca5dcd166 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Landmark Attention: Random-Access Infinite Context Length for Transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f5cc80b-5417-43cd-82b2-9031e1968f8f · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209228c9-fd6c-4ec9-b7f3-fb5c3f6e08ff · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? YaRN: Efficient Context Window Extension of Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b931d1-cca0-4832-86c7-03529d525cb7 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f5f048-b73d-4ef5-8bf5-d2d54a49344a · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8bbd1dc-631f-4ba0-bba4-95b283e3cd07 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Massive Activations in Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1501f1a1-72aa-4039-a278-e0920a9a9b5b · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd10410-c72e-44a1-9d17-4fcdf0c3360d · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66cf8d94-71d7-4122-bc93-fd08332f0159 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9bd265d-1a31-4ee3-84ce-cd1fab6893aa · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72e3d66b-1e85-4b29-81d7-24fdb87c2829 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Efficient Streaming Language Models with Attention Sinks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4197dba-3d4d-4a2f-9023-2461b166a976 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Effective Long-Context Scaling of Foundation Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6080f2-1949-4bc6-9f7b-ade0e9f6c6fa · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Retrieval meets Long Context Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712a7281-3054-48fc-ae0a-4479ad7bbbed · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Qwen2 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e77dc3-cefc-4599-b02b-e6c49b622e11 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8835e627-336e-4a43-b395-415ed7333fa4 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Yi: Open Foundation Models by 01.AI
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e67bf589-4100-49f3-a77d-8ef444fae1a8 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb15062-fd8c-4a26-9137-5e4771313c26 · outbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0acb92be-2151-487a-a413-a395146d8257 · inbound
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6121b6e2-da21-4c2b-9b6e-dc4fb3b358bd · inbound
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b118c4f5-32ae-4d1a-aa9d-fbb97ea9b043 · inbound
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.