Pith. sign in

Paper Citation Record · LEDGER

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 10 inbound Pith citation observations for arXiv:2506.18880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18880 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:58.003217Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:11:50.688938Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4c49ba96-fb44-4476-b86a-3ccd9dca7be0 · outbound

This paper cites Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.825201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.825201Z digest=sha256:c18b0a29209fb6a32727da6c3dba6f8f6bd25ce217404457ce9bfff70ebade43

Observation 413c4957-2b95-4a80-936e-0c5632568c77 · outbound

This paper cites The mathematics of deepmind models.The Mathematics of DeepMind Models (November 01, 2024), 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization The mathematics of deepmind models.The Mathematics of DeepMind Models (November 01, 2024), 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.596042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.829880Z digest=sha256:d3ef7d4bc731f5feb5629bb14cba2fdf7e060b4c239c05555d82bc58bff3c5aa

Observation b251ea2a-1ed4-4599-aa5b-15aa881e54f3 · outbound

This paper cites 2024 aime ii problems/problem 1, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization 2024 aime ii problems/problem 1, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.578352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.834002Z digest=sha256:a0f923b79f9360574e45283552b0c034a3dd3ed16c56177cc842f00de663ff16

Observation 04bc275b-8698-489d-8c42-1bfc7c8d76a0 · outbound

This paper cites Creativity and artificial intelligence.Artificial intelligence, 103(1-2):347–356, 1998.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Creativity and artificial intelligence.Artificial intelligence, 103(1-2):347–356, 1998

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.566520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.838151Z digest=sha256:041dbff0badb57ff06769a2e2217bb4414f8b2e9a667358a4e19e578b8066643

Observation af0a9c24-1f09-490d-8dd0-04f6a5b3b475 · outbound

This paper cites Compositionality and Generalization in Emergent Languages.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Compositionality and Generalization in Emergent Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.842320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.842320Z digest=sha256:f6bf48bf33b5abfaaabdf9c9420df797266d323bddc97f71ba9a95f2363c009e

Observation fb8c3a10-d888-46e0-96a2-deacf7e66cf3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.846603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.846603Z digest=sha256:22d2bed3f21b26cce64cd9bbe8d46846cd7ea9354d7177711676cef0a55aed2e

Observation 39156a0e-efd7-4038-943a-f46fa4cc07d2 · outbound

This paper cites Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.850744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.850744Z digest=sha256:2713cdc217760dffc6703e0363e2c385a9a4201e308e05b0e2c644f4e0a3d270

Observation 9c8ce179-d975-4c0b-a68a-51f3f6867c53 · outbound

This paper cites Metamathqa, 2023.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Metamathqa, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.547650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.854767Z digest=sha256:85d693255cc1039dff7cec9329a26a88e4d8bfd8b5dcbad0f403ddbd8b7a5956

Observation a6ec6f48-764b-4ec5-9780-ecf65165d822 · outbound

This paper cites Improving Text-to-SQL Evaluation Methodology.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Improving Text-to-SQL Evaluation Methodology

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.858624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.858624Z digest=sha256:7782135ada71f37851ff362c19117526c11628e7c3bdad5ddf90e44af65ca8e8

Observation 5bc87e77-95a3-48a1-831c-14d125ca03bc · outbound

This paper cites Deep learning with long short-term memory networks for financial market predictions.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Deep learning with long short-term memory networks for financial market predictions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.535725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.862712Z digest=sha256:32688334e399002569e37dc2bc3fe4f70a985d7fbe4d4faa35540b0898cbe18d

Observation d708f446-0c28-4d1a-8405-9eb09090be86 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.866666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.866666Z digest=sha256:735b9ad273f8bf5938580c5010bd026c478698df7ad7e335a0611810ac1bc5b8

Observation 2c278d11-9ba9-4d26-84dc-0ff445939953 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.870835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.870835Z digest=sha256:397ea41f1201f4b95484e4fbbccb5dc46851ec7e5fb5f0478eaafa6a62e8818c

Observation f8eebb7a-5381-4e4a-b163-18f14c06a519 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.874797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.874797Z digest=sha256:8d391b820b4b5360db1fb03a1e405c20c5fe2dfc3c4c3392439bcd8b77a32f3f

Observation 287b1408-3754-45d5-9029-e360cf92e0ce · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.878375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.878375Z digest=sha256:950cd0f51c1af7279a1bd85326d01738c6248fc99d8b2d69b82e16d69dd275f9

Observation ab8612df-863e-455f-80d5-90142378ea31 · outbound

This paper cites Compositionality decomposed: How do neural networks generalise?Journal of Artificial Intelligence Research, 67:757–795, 2020.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Compositionality decomposed: How do neural networks generalise?Journal of Artificial Intelligence Research, 67:757–795, 2020

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.515058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.882450Z digest=sha256:6940d7223407084b6497277d8496d61204ce1dbb7539f11910a6978040472b8d

Observation 9c8a0537-59f6-4073-b5ab-248a746aa363 · outbound

This paper cites Measuring Compositional Generalization: A Comprehensive Method on Realistic Data.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Measuring Compositional Generalization: A Comprehensive Method on Realistic Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.886073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.886073Z digest=sha256:3df11a5949c58fa8a76ba59e2bef01cb201e925c6153f5f2ca4b448835f246fc

Observation 50b1205a-f1e2-43e9-b5fe-f030ddfb6270 · outbound

This paper cites Measuring compositional generalization: A comprehensive method on realistic data, 2020.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Measuring compositional generalization: A comprehensive method on realistic data, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.502787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.889892Z digest=sha256:31609f9f740695ad67298a3643e717ad3d239046ba7fc571cf89399aee86b4bf

Observation c01db1ad-85ae-4b29-bb97-f8cf6e5d835b · outbound

This paper cites Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.893487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.893487Z digest=sha256:7e49635ef129a28ba4c363cd6a9525a3f1b5e3adde93afff89fa849927f50a9e

Observation 6675cde1-1fe2-4ae1-8ae2-568ce4f706fc · outbound

This paper cites Solving quantitative reasoning problems with language models, 2022.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Solving quantitative reasoning problems with language models, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.897821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.897821Z digest=sha256:3df3272dc555259d013f7cd2001b96c2367aad06c3b342e025ce9685768db524

Observation 4be8921c-b3a4-4e82-bb62-34aad40b7ee1 · outbound

This paper cites Numinamath.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Numinamath

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.474997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.902199Z digest=sha256:69a2ecf02ef92d11bde4f4c9f1fe05a4d9cbb766d6c7423fa069a5c841af60cc

Observation 7eb29bb5-7887-44c3-99d8-256f31919b46 · outbound

This paper cites Probing out-of-distribution generalization in machine learning for materials, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Probing out-of-distribution generalization in machine learning for materials, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.462392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.905837Z digest=sha256:36474442eefcd57ea3348b61ac1c29a3c40fe30151309f8261ea428d163bbe56

Observation 0bd7f082-a1e4-450c-89fb-6ab849c9fd9a · outbound

This paper cites Gsm-plus: A comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Gsm-plus: A comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.448652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.909496Z digest=sha256:f085b32c019fa8c97077c84e0cc77db47089a7e9c5984e89a229ae940e3068e8

Observation 7fcdc50d-c832-496e-ba54-3ea34ce3aef7 · outbound

This paper cites Towards out-of-distribution generalization: A survey, 2023.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Towards out-of-distribution generalization: A survey, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.913037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.913037Z digest=sha256:d4c190e6755121ab129d9d19f1473c2c6e2cd1aae6de7016aafe8ab1672a5f4f

Observation fb35d26b-a10e-49c6-a693-7e217cb88bdb · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.916405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.916405Z digest=sha256:a434653c4d7456f44462ef071f2c92f45838222b78cd1ac3a02f4d82b54fea00

Observation 21a9a9b5-b234-4026-87ce-c265df39b69d · outbound

This paper cites Compositional generalization by learning analytical expressions.Advances in Neural Information Processing Systems, 33:11416–11427, 2020.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Compositional generalization by learning analytical expressions.Advances in Neural Information Processing Systems, 33:11416–11427, 2020

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.429281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.920220Z digest=sha256:f23fc207317edc073d2d1cbdfcf511b0947a9f9d5048d7449542160abaa22f34

Observation 841e4e97-a348-48fe-aecb-6d8bea74d693 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.924842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.924842Z digest=sha256:dd010e9a22c6b8b79dea9e359778e0b2fa95ccf03ab0498f58b5cf6e100fcd45

Observation 72cb1a5c-fd3f-428f-9223-edf8b72ec7fc · outbound

This paper cites Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.929508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.929508Z digest=sha256:179cfa4742cda658151d01920f6f2295460ac5c5589f8910f7ae39a210374f59

Observation 4b06513b-cc66-4bb4-b67a-31241f400e33 · outbound

This paper cites Gsm8k, April 2022.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Gsm8k, April 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.416982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.933504Z digest=sha256:34ec7cb9b27b174a4e62d4e70895d8d62f176b1b6c0e95a870a3f48c5ea70124

Observation dbafe4e1-dd2b-40f7-b86f-8ae33e5b2df9 · outbound

This paper cites Learning to reason with llms, September 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Learning to reason with llms, September 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.937329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.937329Z digest=sha256:1fe6d1c619208d71fd2c34ad829c5cfe655812dde75eb4b5ab947d77327ae75d

Observation 85be693b-36fb-4318-8dda-bf18b6ead1cd · outbound

This paper cites Math 500, November 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Math 500, November 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.398067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.941089Z digest=sha256:66d13ec8103ae9419acb406d909d9f237fcb716634f0970e47e693db25805eb0

Observation 2d6a2901-5fb8-428e-8cc7-bad3e1d337dc · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.946217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.946217Z digest=sha256:7a51271b77a7cbaf5b5f4f5b92b3d27f9836d1c458fec7f8492f83203e01573e

Observation 7e2fc55c-16e6-4ef7-9a82-cb309150cae1 · outbound

This paper cites Climbing the ladder of reasoning: What llms can-and still can’t-solve after sft?arXiv preprint arXiv:2504.11741, 2025.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Climbing the ladder of reasoning: What llms can-and still can’t-solve after sft?arXiv preprint arXiv:2504.11741, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.950314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.950314Z digest=sha256:15214e865c8701ec51ac0a9a9b383797117c779bf84c723249954f5ffbf2f162

Observation 82d3b542-5ad7-4c43-9036-ce8a3ed41083 · outbound

This paper cites Mathscale: Scaling instruction tuning for mathematical reasoning, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Mathscale: Scaling instruction tuning for mathematical reasoning, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.953862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.953862Z digest=sha256:6caabc468c35125abe8e4c557b8f14ca3a306176b1aebca2645762654a52c58c

Observation c4cb4824-27e4-44ad-8037-8f70e4f00768 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization The claude 3 model family: Opus, sonnet, haiku

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.375061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.957314Z digest=sha256:c9a501fa4b1ef455560f2fb49e4f73f684bc0de1776b0609a68d1969aa237f1c

Observation 3657a068-ad2b-497b-8d6e-d34aa6463fc0 · outbound

This paper cites Openmathinstruct-1: A 1.8 million math instruction tuning dataset, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Openmathinstruct-1: A 1.8 million math instruction tuning dataset, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.361658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.960941Z digest=sha256:519b60ffe519abaa1f1ad42462669306c368c42229174b031e93182c1efbbd3f

Observation 632248f7-a2ae-41d6-90b4-cc393bba50f5 · outbound

This paper cites Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.964503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.964503Z digest=sha256:1c1d69b4383dc1f6f57ea0672d2df492019ff09786781eb819cd11d5abfc698b

Observation bd17a92b-19b7-4ee2-98c0-eef0e8ffdb1a · outbound

This paper cites Towards a theoretical framework of out-of-distribution generalization, 2021.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Towards a theoretical framework of out-of-distribution generalization, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.343032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.968152Z digest=sha256:83312df6b7305dd1942502469e8e2951e5a1b8cce22d8b539120530977f718f8

Observation 147c34e6-d020-4e39-8219-ca0b8bd18de8 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.971551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.971551Z digest=sha256:e25750ecc744b8664aada089410f7f7fabc4505c3b05ec41983d939517ea49f5

Observation 073e41ac-1ee1-4177-9364-f3e7c34b5d71 · outbound

This paper cites Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.975316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.975316Z digest=sha256:a11c60cc9fb80ebb219b97eb7fdc5bd4cc48bea12ff555ae1a222cc61b103e6d

Observation c44c3d13-2f48-4716-9def-b3198ef423e5 · outbound

This paper cites Evaluating the performance of large language models on gaokao benchmark, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Evaluating the performance of large language models on gaokao benchmark, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.330336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.979227Z digest=sha256:77ac3b90214cc220477ad2961ae141fe210cb2ae127b48caf03b2ea27b98495a

Observation 8b282595-51b0-4b30-a444-9e848b9a909a · outbound

This paper cites Can models learn skill composition from examples?Advances in Neural Information Processing Systems, 37:102393–102427, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Can models learn skill composition from examples?Advances in Neural Information Processing Systems, 37:102393–102427, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.314082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.982617Z digest=sha256:b7619f0bfbc5732031160d0aa039f0307870c5194c0f8756a27808a2bcd215dd

Observation 7527bc67-e77d-4288-a537-56d461fca4fd · outbound

This paper cites GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.986232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.986232Z digest=sha256:76dec559c145cec9efe4fe81c76cbae91531c505aa5b8c0f1f34036206696ea9

Observation ed18e313-d1ea-4dd5-86b5-a08ca1df6ff2 · outbound

This paper cites visualizing.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization visualizing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.299553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.991015Z digest=sha256:8f92d068e71daa17312211fef9c523930607dc48cc845095207132c18666c89d

Observation d56b2f23-c352-47a0-b072-177155a845eb · outbound

This paper cites conjecture.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization conjecture

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.287433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.995535Z digest=sha256:01f552348d32fac60d15e8765e03ee2a9ba15b1e5312a8360fb11b40d74e07ab

Observation 3b4428ef-c97d-4243-947f-abcbfcf47d56 · outbound

This paper cites computation.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization computation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.274551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:57.999597Z digest=sha256:14b478031883d1509b3d854b5aac3b3768c66de77cda03132a59c6f7b6cb78cb

Observation fed93d16-0519-4544-8441-bb4121582713 · outbound

This paper cites computation.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization computation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.262364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:46:58.003217Z digest=sha256:45f0a6537bd7fa7c227751fea8f475729e12b3b9af2fc6bdc753fe0d52f11c75

Pith citing papers

Observation a4955607-c80d-4a2a-8e1d-ba00be34a93d · inbound

Rethinking the Illusion of Thinking cites this paper.

Rethinking the Illusion of Thinking OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:39.097178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:03:39.097178Z digest=sha256:818b443ae2198fe96c6fe0dd117f978c56a8c6655d9da632df3677e38c8d4ad9

Observation 32cb581c-3213-4fb9-b9df-af707d223f04 · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:16.918871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:16.918871Z digest=sha256:e2fe78a12de6f9ad1bc2249e6506c0ba94704e2eb40dbc73a4c4c6145e33b052

Observation ff53744e-1a34-46c0-a3cf-94671faf3c08 · inbound

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments cites this paper.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.137256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.137256Z digest=sha256:62b0afd7010d6556d75dde087990906df2a4b5bd0b36376aa94d9ee5561c7e66

Observation 1356ad6f-4a21-4a9c-bcc5-88494310d421 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:56:11.366338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:0f805df632d47dec1f9ccd60443763b1d68f547d33d741dfe38b625e5ee0c1e3

Observation 2c7fef75-70e6-49ce-a412-f59312a560a7 · inbound

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers cites this paper.

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:50.818472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T19:11:07.615149Z digest=sha256:6de68f4f7121d87bd0ac82c331fa898b63d9d77f415fb852d2b2abf7d056b633

Observation 4afce806-873a-4d04-be21-9a925ef2d5de · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.904520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:e407a09a5cda2f6cee038ff9fe3c55308b3c734f246f80e8c3a2b1ec83bba902

Observation 3ab79f2f-46a2-43a7-aa3d-7b632af45827 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.449215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:9e2f00994156231a40aabc122d46f19b6340ed38bfb9b4a5226423c7542a3089

Observation 52890016-2d81-4841-8b4c-b804953a11da · inbound

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages cites this paper.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:15:30.210152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-08T19:08:06.748619Z digest=sha256:074f45c7a8dd5bb57cbdab2e7163ca78219090201829ba81c2efef1a805fd3a4

Observation bb9418fb-fd26-484e-91b6-02a8fea9c261 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.759751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.759751Z digest=sha256:6215a933abe26b6076298db5e79fcd181d2162e502e191c7325527e61bd933f4

Observation e3115606-8d7d-4b7a-9c8c-e296f1eaddaa · inbound

Ask-E: An Environment for Calibrated Question Generation cites this paper.

Ask-E: An Environment for Calibrated Question Generation OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:11:50.688938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:11:50.688938Z digest=sha256:d23ff23942b2946e2a9767b27881d8d38faf35db46e822def316b43075d87f8b