Pith. sign in

Paper Citation Record · LEDGER

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 14 inbound Pith citation observations for arXiv:2502.05252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05252 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:25:16.984092Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:41.438347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:19:23.876486Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4bb915a-8c43-4bb3-93a5-c825f3db962c · outbound

This paper cites @esa (Ref.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? @esa (Ref

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.740385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.740385Z digest=sha256:b0fce20546c3aea4ffac64df8bc8e1623bf4c1a260bf3f7ca79b0480b2a679f9

Observation 8ee7e233-7850-4e89-815c-7d80b9ef8b8d · outbound

This paper cites an unresolved cited work.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.747628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.747628Z digest=sha256:67b0760821dd1f212a1d09f7910e9e811527e6d669d11d41a486a795b63fdd1a

Observation 23794d80-f9ca-4366-8a56-22813ae5adce · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.753303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.753303Z digest=sha256:4ec9fb0de5ffc5190e7efafb4dd99614766a2f33aaa5fe4047eb5de75372b6f2

Observation 767b9b18-8209-4b7e-b845-7ff8bba35e4a · outbound

This paper cites write newline.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? write newline

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.761081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.761081Z digest=sha256:1c716af9e95d616a44fbbcc90c0b67b6744a547a5ce1a15fb1f9353f872b5883

Observation 70001e99-033b-4608-9428-9cf998ad5983 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.766880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.766880Z digest=sha256:88b85fc33b94b04df4d9d0dffccd256ce2d55f143c97ab2fdcf3fbed01842708

Observation e456d5dd-cf02-4266-848d-c477b5ed4276 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.772448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.772448Z digest=sha256:3afaaf21d47b7e8b591e92afb4c2574b3f71180fb2b6d391f512432dc99d5ed3

Observation 37095ade-0bf5-4673-941a-15fc4536eb9f · outbound

This paper cites Longformer: The Long-Document Transformer.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.777759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.777759Z digest=sha256:91602e0ac9714562b6ee9315f24ee3667a56de17e153ffbd58635b5032426a6f

Observation af607390-b8c6-49c2-94ef-527c11dcf5ae · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.788920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.788920Z digest=sha256:9ef86c29970dfacb7e54e65857dea8de3d1462fcacd9bddb5727ba378616800f

Observation 2816ea9e-b8a9-4e7c-a8bb-2b4aed73cd54 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.800528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.800528Z digest=sha256:aae53d2a316169b46258bd449eea22a7feb39220c2e7011bdffde639a7b3ab96

Observation deb7c6a1-35a5-4fd0-98f0-f325b302d22c · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.805549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.805549Z digest=sha256:89103cb7aa2e3ff54d13d12cd64fce55444009dffb93a7188a5c19eb3d840d31

Observation f22df550-9487-4c3f-96d5-4cc61017e205 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.682833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.810585Z digest=sha256:4a895329d4e39b656ccf2ce3f9203bafb15630466f7476ab835227cbe3a33c2f

Observation fd4ba3bb-527f-403a-9c2e-188f9529c300 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.815476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.815476Z digest=sha256:ed97b2523dccc8284768e8ef1578717f36dce770346442cd60c01a8cccd139be

Observation cc7a17b9-e075-4884-95d5-0a71615f0d35 · outbound

This paper cites Neural Networks and the Chomsky Hierarchy.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Neural Networks and the Chomsky Hierarchy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.820859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.820859Z digest=sha256:c9ed7151eda425c137f65bf28d1d751ff51a72e4c6b59ea800131219670c3a9c

Observation c2fa464d-b6be-449f-96e3-ecf758f6a627 · outbound

This paper cites The Llama 3 Herd of Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.826047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.826047Z digest=sha256:154db7ad6bec77b2b30d3b7a73ce820c9ad1a2a4ac7b3c103588a68bf96fedd4

Observation 911d184e-580f-41f9-9fee-f9ec652b0794 · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Needle in a haystack - pressure testing llms, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.664012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.832004Z digest=sha256:987725e2d2b08a0faeeab4eb73a7736fc6f0f2b2da6b64ac2e73c38216c928e7

Observation f57ef75e-9fff-47cb-91c9-5ab88eb228c1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.837781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.837781Z digest=sha256:19867c5d1c291bb6152ed07f8cd7359ae4aaca42115ae1b76cb6b44665f1cffe

Observation 2a8d70f9-c571-4ef6-9661-5c72e8cb49e8 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.848492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.848492Z digest=sha256:f1748285da2c47cd0f81e4fc1aec2bcf8e6f874b6a8bab2631c88aae605b7005

Observation 8b667d32-78fd-4e74-990b-52721a025d7e · outbound

This paper cites Mistral 7B.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.853620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.853620Z digest=sha256:74c0e78106df1315489eed173cdc79ba547b3d9f73611edd669f35dd9eac9ac2

Observation 5cb60de0-520f-42c2-91d2-faa9b7e76904 · outbound

This paper cites Active Retrieval Augmented Generation.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Active Retrieval Augmented Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.859507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.859507Z digest=sha256:635f771cabc71723b87eae8becd8e5df3ebf7076bae2e2a5b2522c685dcdf23c

Observation fa549915-4c4e-4630-b053-f062f4cad42d · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Needle in a haystack - pressure testing llms, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.645657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.864472Z digest=sha256:0193ff1e1ca225f21e25445e2e81f1bf9cbc0b98c7848e69607badc6e74a54c3

Observation 190ef306-8a5d-4f36-ba38-65563c3d4242 · outbound

This paper cites BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.869442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.869442Z digest=sha256:ffc808238fb27247ed64a8467361d9c4c95d439851f83cef933a2b7c14619ece

Observation eb849715-65bf-4fe6-9e3c-4df0bb4bb99e · outbound

This paper cites Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.874664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.874664Z digest=sha256:7bfde1146a21c2fa5083eeb21637d214301cf5968be275c5b5b84960e780c270

Observation d5b394f4-c6a7-4a30-8bc6-ee4cdde6c0cb · outbound

This paper cites Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.880139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.880139Z digest=sha256:3f6007a6af908337250177026b3879374769eb16f0f3e3e45fcde3988d0e7d71

Observation 2a7a3af3-51a2-46c5-bbf0-96150d6fcedc · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.885207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.885207Z digest=sha256:73420062855937d41f312472062d91b01d0434e9d4e280b08a515943ba28e270

Observation 937d0ff0-3e26-4cde-aab8-2b32467d4ba9 · outbound

This paper cites Long Context vs. RAG for LLMs: An Evaluation and Revisits.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Long Context vs. RAG for LLMs: An Evaluation and Revisits

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.890221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.890221Z digest=sha256:1d758225fd79dac12584d5cbe5deaf08cb8f8ce8a5e3b6232a2cacefcb64cac8

Observation 2f5691aa-e510-4aca-a989-4e9e6de23538 · outbound

This paper cites Retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.627462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.895307Z digest=sha256:eef0dc4efc56c8be2851f9fa95263519486ff8b6230273c723771402e5237c06

Observation 4520cab5-8190-40ec-9c3b-db670fd2140e · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.900015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.900015Z digest=sha256:973eedeac5b28f769bddd375ed07d52038cf514a8dcb5d5624e138019a3de35d

Observation 1292359e-02cf-47e3-826f-2e24f6be073d · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Lost in the Middle: How Language Models Use Long Contexts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.905007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.905007Z digest=sha256:1a14275863b76607f9e675b13c46e4c9c8d456abea5be2258800ac7a85c1ab77

Observation da20d73c-593a-4b08-8a0b-38a07b2b10af · outbound

This paper cites DafnyBench: A Benchmark for Formal Software Verification.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? DafnyBench: A Benchmark for Formal Software Verification

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.910497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.910497Z digest=sha256:40313c3c0422b5f9709fcb33465a4787e1cf43e5c01a4e69219ebddef755a409

Observation e4214ef1-4b51-46d3-9c0c-f24b46781a31 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.916090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.916090Z digest=sha256:3e1cea2ca1d6a4a0561e55d7a0ef9a3a596fa283f6b03f97da34bb336dbfe7d1

Observation 34d51ff1-5307-4ae5-b21c-113fca06d93d · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.922469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.922469Z digest=sha256:cbbf2f0aff1dc064cbaa4f9fab6b6928ae4e3b56ce7acec6bdb113a0aabc7230

Observation 9947327d-b5e3-4c06-8b34-13b21beb5bf3 · outbound

This paper cites Qwen2.5 Technical Report.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.928032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.928032Z digest=sha256:27e040d3f7f17bd06a4ec540ae9467da3b09909c6d36ec7b1eba56d5afe26f8e

Observation 516599cc-a18b-4080-ae52-665b23f30ef4 · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.933498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.933498Z digest=sha256:f25de7d8fa58fba553cf0178b7c23b9a48c1ec07a1a98b51a6bea3cdc8104c9d

Observation 208e2cbb-bcb8-44c1-bc73-2e1eead2fe66 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.939267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.939267Z digest=sha256:dfe6d5c2441a16903839416ac0d1c8bcc0cb5d1f8c7fd08b02fef7dde8d359ba

Observation 2f2b01ab-d196-479a-b1a9-a493a486b12d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.944984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.944984Z digest=sha256:150ea15887bede974d71d6b3ee6ecf7594627add51abec9db2d9bd17c2f87f71

Observation 82a51e3a-6c9a-418c-8cad-62795881d5b4 · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.951317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.951317Z digest=sha256:c55cebfdf43c5089802af3ffefa7c14f3672360f253bb9a88b5f47465ea82f7f

Observation ee703469-e33e-4977-bfe5-8e082c6167ae · outbound

This paper cites Modular elliptic curves and fermat's last theorem.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Modular elliptic curves and fermat's last theorem

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:17.607210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T20:25:16.957513Z digest=sha256:b5118ab967d7979a266fc6894e834b076d907049476d6cb8131758b2d62c8f40

Observation f8675a3e-e296-438c-86b1-91e4ee6e5ac3 · outbound

This paper cites Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.962705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.962705Z digest=sha256:fdcacaf9babef1f90d0f025ff17fa3dbdcba453ad2f6b0482af31a9a7f52cbdb

Observation 542152b1-4ed4-4c88-b19c-2b68971428c2 · outbound

This paper cites Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.968703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.968703Z digest=sha256:00044a3aa6c6bf03c8d600166a304342b4c0ba6d5c71e18a10865e44a16fdf89

Observation 87b9c447-6dd1-4c54-ac8d-e1a7496cf3d4 · outbound

This paper cites In Defense of RAG in the Era of Long-Context Language Models.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? In Defense of RAG in the Era of Long-Context Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.979011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.979011Z digest=sha256:d261b2853810c892574b92449d95b7bb5b6bb9efa9da33f98ad48e4aa523df56

Observation c4e73476-9160-44b4-97d4-1b5ad729b0da · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.984092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.984092Z digest=sha256:d4d81ad22c9ffb94043705fc970afd2108eddd3b2881a778988cccbb57a38ebb

Pith citing papers

Observation bfdbaa7f-bb57-45e9-a4bb-78d238f20d31 · inbound

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings cites this paper.

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:41.438347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:41.438347Z digest=sha256:91963d19f6b4843ff3999b694d94ff81a79ce4d9b03abb6df1b492d439c3739d

Observation 2448123f-d8e5-4f46-ada2-a791e302adcd · inbound

LongReasonArena: A Long Reasoning Benchmark for Large Language Models cites this paper.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.176410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.176410Z digest=sha256:597ddc6545cdeb3fd2c4dc270ee5e1a813a0e6d7da6a57e01082da9d05cf3b8f

Observation 9c9a0ee7-7f47-4758-bf05-0b9669aa0dc6 · inbound

vAttention: Verified Sparse Attention cites this paper.

vAttention: Verified Sparse Attention GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T11:21:10.093674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:21:10.093674Z digest=sha256:ef6800980d0757e7d6658ebf8f7fd92794ce842638afbe7e35e5ed6e582803e0

Observation 35219edd-3f3e-48a5-8495-413cc13ba9d3 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:33:32.819008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:dfb21872bc6ff511323a65ff348bcb365423268a7d37bfbe7d52f1f31dc83d54

Observation 513603ef-f1ea-4074-be2c-07e15eadb755 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.602502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:13929d710aebe6b259bf0d7b55cced857462def57a1d06568591344866b3e960

Observation 0ab7533e-c155-4bc4-91d7-dcefc186051d · inbound

From Local Corrections to Generalized Skills: Improving Neuro-Symbolic Policies with MEMO cites this paper.

From Local Corrections to Generalized Skills: Improving Neuro-Symbolic Policies with MEMO GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T15:10:27.013215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T15:10:27.013215Z digest=sha256:2ed023ba4361d0606a1e6c48854c2b5dd4c072b219a88ac85a1e315527003c42

Observation c81149b1-e343-401f-9449-19308cc205c0 · inbound

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning cites this paper.

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.740864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:34:14.413131Z digest=sha256:d33d0f2941f500047b77c92773b94412e05a3b711f19052f0fbc416fd4c57171

Observation 74e4b96d-0ca5-4a0e-acd1-7008e2a33d6b · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.190972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:8872be25e56f29725b890ac6a7285e743b033de4909cdad67fddc9b683ebb1bc

Observation 1d00e09b-b71d-4e6e-8d9b-49d4d7b10ae9 · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:2794a686bb5f08ad50daa00efb9974c05c00a3e6131f76b07229b9d40a31cdd2

Observation bb9b8bf6-909d-405b-aa7d-9a7e2cb7a8e6 · inbound

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding cites this paper.

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.770407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:25:20.511933Z digest=sha256:28c4394c0353f870ff4c86b3fdca60940c34786818f775ba11c8407b593c684b

Observation 6a15e411-70c4-41e4-aa92-dc30f06307ed · inbound

ATLAS: All-round Testing of Long-context Abilities across Scales cites this paper.

ATLAS: All-round Testing of Long-context Abilities across Scales GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:27.979086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T13:30:09.778806Z digest=sha256:d4c891010fd5559f325ab6bc09ae58c2aba681989ea8a1db8065f9956cc2fde3

Observation ff892fdb-d0b3-4522-8990-5e2477da3a48 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.878014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:d8d5bc6fbd8607e3b6d377ed8257c3be9e62f5ff9690ef1cf8e6629974942d57

Observation 09806160-cf7e-4d87-899c-67888cbaa6d6 · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:04.371774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:04.371774Z digest=sha256:dc1a97fea04ea1080b8229eb6022f1b5d6c5ead11bcfff4e85b28886d3e0de2e

Observation 3a8dba01-3ae5-4436-b557-10af259cd432 · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:53.900522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:53.900522Z digest=sha256:8a30815ec96881baf2a1478217447f7efbad010a2e90407747e713e63cac1e99