Pith. sign in

Paper Citation Record · LEDGER

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2505.19293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19293 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:26.736743Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:40:58.250915Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:00:21.895979Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7765fd1d-24eb-47ff-a883-0d7f005b270e · outbound

This paper cites online" 'onlinestring :=.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.126509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.126509Z digest=sha256:70c6bbbcadf3f5de4aed40137f86f86ff4ba4f4a6c7f64ec8f2d4cab256fd550

Observation 7785cc7a-5b1f-475b-956d-a36804d325e9 · outbound

This paper cites write newline.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.180478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.180478Z digest=sha256:3c89c2a20f646de2bd83be5e7d0d833024d102b65d214a6c8538ff77ea8e193d

Observation 7a986590-0b74-4e28-a975-fd73578628f2 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.244756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.244756Z digest=sha256:dbb22e39a7036289e1ba529b45f0c143a3051af154050e16d8a7bb1afd9eedbc

Observation b688bf22-8169-4d5f-9dc2-cfb763875804 · outbound

This paper cites GPT-4 Technical Report.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.261232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.261232Z digest=sha256:1c10ac3af2d816fa9ceb2160a8139d0aadcc096cdbcdfab8f1274a63100d4e6c

Observation 9213b459-1029-4b1c-ba61-9c03bf6eb69b · outbound

This paper cites Many-Shot In-Context Learning.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Many-Shot In-Context Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.272210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.272210Z digest=sha256:64a0cbb3208c9942b312a73c42e0c7a0a8256a7a5e717dd03f3e4b392f82bba8

Observation c728fe26-7aa8-4634-9e18-24b52d96d5b0 · outbound

This paper cites L-Eval: Instituting Standardized Evaluation for Long Context Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.321920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.321920Z digest=sha256:c505a4eebf9e19690fdd2aa83c5cb3e04809d3d1815bfbb423df1118c0c01fd1

Observation ffb476a2-5b16-4523-94d7-f7fd05e20bb2 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.418649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.418649Z digest=sha256:4ab211577500493a220695c060a890705e428ea8798259cd2789977804b0df13

Observation 0daae9fc-557d-49ca-b9c5-826db961a409 · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:28.458074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:20:24.483016Z digest=sha256:fd87066135968f2f3c929fa0e0cc27c4d7e19f0623725cd67c982222ca26f0b3

Observation a821fbc7-3ddb-4320-aa9c-93c253bf5643 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Extending Context Window of Large Language Models via Positional Interpolation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.550556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.550556Z digest=sha256:0946db457164acaec89c60e7054891db3792a42b17c3c93c654dbc1f085dad15

Observation a46d8258-ff30-447f-9756-d2eb5aa8260e · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.621120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.621120Z digest=sha256:5af72d75eb5ce40411d8f83380088ec752d989907e4f213593445a57ca2266cf

Observation dbac3db7-9e4c-4bc9-b544-60ac81a6badc · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.696735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.696735Z digest=sha256:943550e7708da66c13a7058197c293c0c0b2567cddec0c7259eef1a03a362b86

Observation 23951161-ae45-4f9c-8cb6-b09ebe855f14 · outbound

This paper cites The Llama 3 Herd of Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.781338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.781338Z digest=sha256:f9106e1fd53ff3c2ff48144f8b1d822dbf8fd4c9b3e1e181fcd7db9572c63a35

Observation 93b193c1-b023-4dfe-bef3-985d11c5bdc5 · outbound

This paper cites MedOdyssey: A Medical Domain Benchmark for Long Context Evaluation Up to 200K Tokens.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? MedOdyssey: A Medical Domain Benchmark for Long Context Evaluation Up to 200K Tokens

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:27.805331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:20:24.793978Z digest=sha256:56fa23e8f1b839613ec72afe5eb3c55e4794ad8faff63fa7ac6eff727ed91a41

Observation c9a4335e-851d-4fc3-848f-3c771c57f95a · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.804087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.804087Z digest=sha256:028b784ee521b1b277296eb3956c1b87e5a72233a6812493ef850ec36cab8b5f

Observation 5ff83b31-5369-4e9f-9c35-d957cb6dac6b · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.810361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.810361Z digest=sha256:d5d28d7b872f6a07f9d4b758d707e848271469925855b2cd9ee33c39a0411025

Observation 7b2b4f5d-42fe-479d-bbb2-971369faef7f · outbound

This paper cites CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:27.578318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:20:24.835670Z digest=sha256:329ec29c0ac20dee7a8585c988fb5a3e7644f74674ad9f536af9fc615f84ef27

Observation aa4703b7-b1d2-4a88-94c5-99b999c0c675 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.924084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.924084Z digest=sha256:caa362521cf218a69fa6450f4c917f27d7b8eb1e7921494cd65ddec4220d045f

Observation 5f8e0671-926e-4089-9ee3-c86c824530ca · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.008242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.008242Z digest=sha256:8b2262068f195bb43b237eaecf199262cf06025a1bb12d71b9e3920ce6318633

Observation 845c94eb-5dac-41e9-a23e-553281c6f0d1 · outbound

This paper cites LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.095761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.095761Z digest=sha256:06270b8e5666a6a6906db38b53513a7776eb3c76fd544d1b97d4f14a127e6d96

Observation fb754909-6e43-484e-a40d-a28a508b4009 · outbound

This paper cites LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.146848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.146848Z digest=sha256:36b7dfd4712f53a06e18c574f2619682fe1bf1446280d3ba4761532ef5e86103

Observation 00993ed0-2ba8-433d-b8ba-24bacf7d5cc5 · outbound

This paper cites LooGLE: Can Long-Context Language Models Understand Long Contexts?.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? LooGLE: Can Long-Context Language Models Understand Long Contexts?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.205683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.205683Z digest=sha256:c534b43dfb4261af6807a1c6eb36d869a196f8c3ba67acb2d71f51fe7cf60d96

Observation 7b19a446-827a-4ff6-87aa-f8fe4cbc08b9 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Sequence Parallelism: Long Sequence Training from System Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.288961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.288961Z digest=sha256:25c510175dcdafdffd34f2f141bb59be760d2b627cd16dce51ce87a304189858

Observation 57c65d74-f6c6-43ae-b1db-8229c215db9e · outbound

This paper cites Compressing Context to Enhance Inference Efficiency of Large Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Compressing Context to Enhance Inference Efficiency of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.308833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.308833Z digest=sha256:6c5cb3a615eea85699c311af2c063a9f140df84feff80da8af0bf4431956d89f

Observation 0f459c01-634f-4b2c-9482-3ea5afab542c · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.419215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.419215Z digest=sha256:ce084c353d091de2e68bf1a081bd1f8a818b797c3b3210a064901c086b3f8b08

Observation 63264196-1f47-4755-9156-30de23677639 · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.442595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.442595Z digest=sha256:5b58f0c492bc8b662ea2f688dcc58848226b40df171b3d6e46b9e8f4edeec95a

Observation 230c6493-93ac-4731-b6aa-60f2fcce3d7e · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.469693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.469693Z digest=sha256:7c7a3b6843f6d7ce44738154aafcfac76f05ed8d09a0fd5718375fbeb935d65f

Observation 50486a72-61a9-48ce-ac92-c1cbedf19cef · outbound

This paper cites A Controlled Study on Long Context Extension and Generalization in LLMs.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? A Controlled Study on Long Context Extension and Generalization in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.515017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.515017Z digest=sha256:078982201f220e63f980eab3d67e2f8240d3ceebf77f498722c9d39343173df7

Observation 4476c01e-0511-439f-a552-e9eca5dcd166 · outbound

This paper cites Landmark Attention: Random-Access Infinite Context Length for Transformers.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Landmark Attention: Random-Access Infinite Context Length for Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.577961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.577961Z digest=sha256:eaa9586d33731928411580cf3d786ae29abc22eea8391c2e9c18367c5a0b936a

Observation 8f5cc80b-5417-43cd-82b2-9031e1968f8f · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.626842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.626842Z digest=sha256:bfadafbc5f68616a3e2fa95fb0a3bd19a31a89f226f5033e8ea608f126b8e476

Observation 209228c9-fd6c-4ec9-b7f3-fb5c3f6e08ff · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? YaRN: Efficient Context Window Extension of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.669954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.669954Z digest=sha256:d0e154efab0fd17330b78e6428e8ee2661b709e41b416721e5bb5ad7bc70b440

Observation 16b931d1-cca0-4832-86c7-03529d525cb7 · outbound

This paper cites Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.752009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.752009Z digest=sha256:3acadd41f8184dc501ab969f5b07fe7516c5b8f3f29302c74a397b065c8313e9

Observation a1f5f048-b73d-4ef5-8bf5-d2d54a49344a · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.875196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.875196Z digest=sha256:75c625a701e6aa9b66f716449eeaa0f23b6900ce0ee390d4e6251190115762cf

Observation e8bbd1dc-631f-4ba0-bba4-95b283e3cd07 · outbound

This paper cites Massive Activations in Large Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Massive Activations in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.968312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.968312Z digest=sha256:379c9718388c14705869d09bb9ee3f50364ce05f3660217f6d6c99959959ae87

Observation 1501f1a1-72aa-4039-a278-e0920a9a9b5b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.013754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.013754Z digest=sha256:0350d284b0d59ebe6ce81f943f296474f96848ed408f543d841b801deb43eb69

Observation ebd10410-c72e-44a1-9d17-4fcdf0c3360d · outbound

This paper cites Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.146395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.146395Z digest=sha256:14e62c472d0afda9bf1c449baae73f4047acad0d99ca5e37ff505fdddf3d15ee

Observation 66cf8d94-71d7-4122-bc93-fd08332f0159 · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.228218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.228218Z digest=sha256:4fb563ee9f7843007f3650e418cdd8e744d33fda8085a5642304137bc0c74c5d

Observation d9bd265d-1a31-4ee3-84ce-cd1fab6893aa · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:28.217434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:20:26.277548Z digest=sha256:8527ea6a6345458fa8878578d5bcd58ff4d3c5af4cc76ea2cdcc4f5a3248d578

Observation 72e3d66b-1e85-4b29-81d7-24fdb87c2829 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Efficient Streaming Language Models with Attention Sinks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.335648Z digest=sha256:8a5009a6ab39bb49d23c8e25a6d5f8d815ad54269b4df29b90a4826a4f77f851

Observation b4197dba-3d4d-4a2f-9023-2461b166a976 · outbound

This paper cites Effective Long-Context Scaling of Foundation Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Effective Long-Context Scaling of Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.452682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.452682Z digest=sha256:43978e3ea822d1bc654d031eed5d746d70b42aa0190179cbd9a1aaa3ed6654b1

Observation 9d6080f2-1949-4bc6-9f7b-ade0e9f6c6fa · outbound

This paper cites Retrieval meets Long Context Large Language Models.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Retrieval meets Long Context Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.529399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.529399Z digest=sha256:1f2a5baf3406cb9f343eacd38989e8b65e3ea1cd8e1fbe18bd74c58546484a89

Observation 712a7281-3054-48fc-ae0a-4479ad7bbbed · outbound

This paper cites Qwen2 Technical Report.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Qwen2 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.597274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.597274Z digest=sha256:e1994db710a38157bf68971f823824db5f3f54f79ed9233e27b2c7663fdd5b25

Observation 87e77dc3-cefc-4599-b02b-e6c49b622e11 · outbound

This paper cites HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.649185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.649185Z digest=sha256:e66d34ce8b18cdd38da538ed0b2cc1b37f2144bd7320964cbc9b974c902f1a01

Observation 8835e627-336e-4a43-b395-415ed7333fa4 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Yi: Open Foundation Models by 01.AI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.695690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.695690Z digest=sha256:5e26174dab6b865c0c6fa6dad3e4b74df0794d1a506e782495b273906380be05

Observation e67bf589-4100-49f3-a77d-8ef444fae1a8 · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:26.729592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:26.729592Z digest=sha256:1daea0f2c7012c30d50e709db8d41f24a1f5482eb85635773fae3f6eaeafa9cc

Observation ffb15062-fd8c-4a26-9137-5e4771313c26 · outbound

This paper cites an unresolved cited work.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:28.135544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:20:26.736743Z digest=sha256:4a6ccbcb220b0d0edece1b953fe9df0d3f94cb290d79db344e6f67a8d822a808

Pith citing papers

Observation 0acb92be-2151-487a-a413-a395146d8257 · inbound

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks cites this paper.

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:03.614635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:58:15.184063Z digest=sha256:dd39bb82fb695d147bb24591ec0ff1cd184dfe1636acf83de171004db4b2c006

Observation 6121b6e2-da21-4c2b-9b6e-dc4fb3b358bd · inbound

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning cites this paper.

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T03:52:24.872919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:52:24.872919Z digest=sha256:332a48392c39c7cae859427a9d08b2781c1cc768f509b45582bbd8e4b813cc73

Observation b118c4f5-32ae-4d1a-aa9d-fbb97ea9b043 · inbound

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning cites this paper.

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T07:40:58.250915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:40:58.250915Z digest=sha256:dc8d0d22557a260d2347a105c779f8691234acb3ed3db8ec6d79ca46021e714e