Pith. sign in

Paper Citation Record · LEDGER

BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2406.10149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10149 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:31:43.210471Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dcbe3a3d-3b5a-46fa-b0fa-600ee6d9fdf8 · inbound

RULER: What's the Real Context Size of Your Long-Context Language Models? cites this paper.

RULER: What's the Real Context Size of Your Long-Context Language Models? BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:20.655690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T03:55:20.355345Z digest=sha256:9406871cb04d03eead43c38a9ad704ee63faf158e8320b7f0228c97addf084d3

Observation 54cc1b36-120f-4524-a2db-2901e5781a05 · inbound

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models cites this paper.

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.869247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:09.869247Z digest=sha256:b5b7c40f5dd9939f3566809c3184c9ed7a1d64ebf487df588241d6fb4d080ebc

Observation 37534588-8791-4d18-a9c2-29caca12fa5e · inbound

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts cites this paper.

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:43.210471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:43.210471Z digest=sha256:f429759e8da473557dd156803daa6dba80dce0d98be4d0ec3b94c803cad49103

Observation 2e03c3cd-abd0-49d9-a16a-3c4563e4d3a2 · inbound

IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments cites this paper.

IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.221816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:03.221816Z digest=sha256:9a75d20313c0b6f1851b396a98be5d84d1bd85929592b0793360cc93071f6be3

Observation 39fbe1fd-1b36-448d-8992-9e86b99e4ae2 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.769302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.769302Z digest=sha256:546493ecf3caa05e30209d8726e2bb0d406f0e0cc038969e7f9cca8511f97654

Observation e97b852f-4873-4ae1-adbe-de8d3a8c5f21 · inbound

LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework cites this paper.

LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:12.466041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:45:12.466041Z digest=sha256:61c8a6c7984c5be747d49a8856e3a509c2fd05ece6490aa3c80612dd4b83540a

Observation 902bc66f-9254-4443-bada-97d472af23bf · inbound

Positional Biases Shift as Inputs Approach Context Window Limits cites this paper.

Positional Biases Shift as Inputs Approach Context Window Limits BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:12:56.219629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:12:56.219629Z digest=sha256:e14fcec0ca859a6ce7b79e073ea5ddbb7436053261dee55fb2b02ad9ecacae04

Observation 7e42fbbc-7002-4020-9ab2-847602c1be3e · inbound

BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections cites this paper.

BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:42:07.490897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:40:51.426489Z digest=sha256:833f30c573467e4f1e39acf3d5ae523702d35011f6a26f5ee76ae80a23963818

Observation e56728d6-6922-42a7-8896-6aef5a8a41dd · inbound

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs cites this paper.

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:42:27.870418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:42:27.870418Z digest=sha256:b83e45a71d32a65ad3cfe6d046c36cd5c50364c67a847fd801d6947c5aeccacb

Observation cc35383c-c636-4fbf-954f-9f50c968b197 · inbound

Retrieval and Multi-Hop Reasoning in 1M-Token Context Windows: Evaluating LLMs on Classical Chinese Text cites this paper.

Retrieval and Multi-Hop Reasoning in 1M-Token Context Windows: Evaluating LLMs on Classical Chinese Text BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:26:09.838239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T16:46:02.443761Z digest=sha256:398c875fa2c960ce1a3da56d108ec2e6c39602ded93a596e1b1acbdc94e0e365

Observation b95822b3-b6f5-48a9-811f-eb0545578443 · inbound

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks cites this paper.

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:21.910436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:58:15.184063Z digest=sha256:ce753e082edca99209035123e7400a2196ffc919accdb80c056547c0dc5efba0

Observation 9fdcd53c-50f1-43f1-a1ff-93b82fc0110a · inbound

Diagnosing Evidence Utilization in Long-Context and Retrieval-Augmented Language Models under Matched Evidence Conditions cites this paper.

Diagnosing Evidence Utilization in Long-Context and Retrieval-Augmented Language Models under Matched Evidence Conditions BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:46:59.863455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T00:56:58.759665Z digest=sha256:16f2d9dd19f2d02f3b28aed2229941a63ee200e0fcfbd8638bd35b07d69a23f5

Observation 6e02aa06-8a04-4e02-9488-c52188e65884 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T22:31:21.491643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:48f91a0777448c740a3fc0a0c1415c9f7a501888a70a72142fb13d19f66d2f88

Observation 5d286266-e25c-4358-8884-435a42ecc16e · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 123

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.661388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:4588c5089139225eea616375774e2186fbc9428652f754d925db987977b0eaf8

Observation 445558d0-253e-443a-afe8-19916fc7558a · inbound

Randomized YaRN Improves Length Generalization for Long-Context Reasoning cites this paper.

Randomized YaRN Improves Length Generalization for Long-Context Reasoning BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:59:46.621214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:14:33.226848Z digest=sha256:576b35f552ccd63302c31b37a6a9e83b7205c6595eb0f0956f0022dab2169615

Observation 133c98d7-fa10-43da-8001-1d13b617990c · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.091424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:28967aa96572017effd4ec1a54d38cfcb4f246e597a2237ca5e7a8af1b4e47dd

Observation 550637b4-7d3e-4350-a0bd-dd5be3140afc · inbound

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning cites this paper.

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T03:52:24.872919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:52:24.872919Z digest=sha256:70ce4d5446cfe68a9e5f3b8e7f29dd1e0dff6c81beb17db86d484cac5a55b0f6

Observation 5b0fcf00-ff01-4511-b418-633c4506c10f · inbound

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning cites this paper.

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T07:40:56.540931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:40:56.540931Z digest=sha256:b8e7e09c940b3b8b0f28f03416684cc2d50bd4e00c406c91bb375febfb77b90f

Observation 1498b5e4-839d-4439-9266-7ebf3bdaad0a · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:74d4d5d7536ef312fb76b47afb5495632f836addade44644176c4b185ac0a147

Observation 3174dc67-2214-4bee-ac8f-24deb82b378a · inbound

Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention cites this paper.

Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T14:02:40.746028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:02:40.746028Z digest=sha256:106c5f8de8b9a11f996fdbad1bf395b680eeb31cd567acd206b724a30db758e4