Pith. sign in

Paper Citation Record · LEDGER

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

As of 23 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 10 inbound Pith citation observations for arXiv:2506.18880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18880 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:58.003217Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:11:50.688938Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4c49ba96-fb44-4476-b86a-3ccd9dca7be0 · outbound

This paper cites Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.825201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.825201Z digest=sha256:eb5fa2d4fc9f6978174e7c3799b642fa76be01004a0b7a68aaa1cceb286ab42a

Observation 413c4957-2b95-4a80-936e-0c5632568c77 · outbound

This paper cites The mathematics of deepmind models.The Mathematics of DeepMind Models (November 01, 2024), 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization The mathematics of deepmind models.The Mathematics of DeepMind Models (November 01, 2024), 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.596042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.829880Z digest=sha256:af38bd1faf13e20bb8acfef67546d9d6fc4cb3cc0c84d1ce83b3131d92d767b8

Observation b251ea2a-1ed4-4599-aa5b-15aa881e54f3 · outbound

This paper cites 2024 aime ii problems/problem 1, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization 2024 aime ii problems/problem 1, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.578352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.834002Z digest=sha256:9f15d32e41e559424390272f17d9829f57096dc98de9aab1f8229f92587def7e

Observation 04bc275b-8698-489d-8c42-1bfc7c8d76a0 · outbound

This paper cites Creativity and artificial intelligence.Artificial intelligence, 103(1-2):347–356, 1998.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Creativity and artificial intelligence.Artificial intelligence, 103(1-2):347–356, 1998

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.566520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.838151Z digest=sha256:d9c0b58efde9ca46a7418f8837a05c0d1934dd21190a8eeceac600fc10a20ff8

Observation af0a9c24-1f09-490d-8dd0-04f6a5b3b475 · outbound

This paper cites Compositionality and Generalization in Emergent Languages.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Compositionality and Generalization in Emergent Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.842320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.842320Z digest=sha256:3178f5a8e76c8917f3a7d25bfc721e2d4875257fe018f4ce21ecd4f000eb3080

Observation fb8c3a10-d888-46e0-96a2-deacf7e66cf3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.846603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.846603Z digest=sha256:c5198ea7f02fe1688b835784dfa80cbffef726ed4e89f2c820ddec9bfbcab8ce

Observation 39156a0e-efd7-4038-943a-f46fa4cc07d2 · outbound

This paper cites Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.850744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.850744Z digest=sha256:6fc0ebed9ee6667f74bd456958e7422c108cf6c347e457d6428f0346c44d2e80

Observation 9c8ce179-d975-4c0b-a68a-51f3f6867c53 · outbound

This paper cites Metamathqa, 2023.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Metamathqa, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.547650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.854767Z digest=sha256:46b5cdd64b606d3dde302a16b8b63bbfe9d93e4d1b1d06eb08a2a88391c4f022

Observation a6ec6f48-764b-4ec5-9780-ecf65165d822 · outbound

This paper cites Improving Text-to-SQL Evaluation Methodology.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Improving Text-to-SQL Evaluation Methodology

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.858624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.858624Z digest=sha256:346743e564a64cfca1f700b5dd841c1c64fcbe821f62dd9eeafdf686feec726b

Observation 5bc87e77-95a3-48a1-831c-14d125ca03bc · outbound

This paper cites Deep learning with long short-term memory networks for financial market predictions.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Deep learning with long short-term memory networks for financial market predictions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.535725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.862712Z digest=sha256:239d1233f11020a344caee297500c7e3e9675d127a39bf99831aa4121f29a597

Observation d708f446-0c28-4d1a-8405-9eb09090be86 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.866666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.866666Z digest=sha256:681f14768bbf420bf7c993225e017d869a7fa5364890909d714cd5d7ac1319b4

Observation 2c278d11-9ba9-4d26-84dc-0ff445939953 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.870835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.870835Z digest=sha256:a617ebb4abf3c6682bc92312d2a77fd00cef1c4093f3438c34b779593c74aac4

Observation f8eebb7a-5381-4e4a-b163-18f14c06a519 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.874797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.874797Z digest=sha256:73b6775303ee3d33da21162591b0404779dd93058a7c6e3d8a4fabeb03680566

Observation 287b1408-3754-45d5-9029-e360cf92e0ce · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.878375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.878375Z digest=sha256:e4bad1f212a7c8b52983c70b65cd9b4b34882ce533d0cc5be4af06384d00f0d5

Observation ab8612df-863e-455f-80d5-90142378ea31 · outbound

This paper cites Compositionality decomposed: How do neural networks generalise?Journal of Artificial Intelligence Research, 67:757–795, 2020.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Compositionality decomposed: How do neural networks generalise?Journal of Artificial Intelligence Research, 67:757–795, 2020

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.515058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.882450Z digest=sha256:4140afae9d51d1295938816c6bec0381dcb5b7eafafb55a71948c86b07228acc

Observation 9c8a0537-59f6-4073-b5ab-248a746aa363 · outbound

This paper cites Measuring Compositional Generalization: A Comprehensive Method on Realistic Data.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Measuring Compositional Generalization: A Comprehensive Method on Realistic Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.886073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.886073Z digest=sha256:deef3f53ea2dcd53d077ac2f71834e0cb64ccfa139c3bca6776b84eb8d7986b8

Observation 50b1205a-f1e2-43e9-b5fe-f030ddfb6270 · outbound

This paper cites Measuring compositional generalization: A comprehensive method on realistic data, 2020.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Measuring compositional generalization: A comprehensive method on realistic data, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.502787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.889892Z digest=sha256:b3f9207a4796ab97653f1a4ebd8af9af38ad32e0d74380cc27344624c73af504

Observation c01db1ad-85ae-4b29-bb97-f8cf6e5d835b · outbound

This paper cites Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.893487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.893487Z digest=sha256:7366d6f974b217c95a09bede67f70cf3dc20336da27e525afbf035cc9a9dfc9c

Observation 6675cde1-1fe2-4ae1-8ae2-568ce4f706fc · outbound

This paper cites Solving quantitative reasoning problems with language models, 2022.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Solving quantitative reasoning problems with language models, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.897821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.897821Z digest=sha256:9778a4c26511e9d958717c8ae514ee592520cda778a0d7bd7741ae534a4a0a6f

Observation 4be8921c-b3a4-4e82-bb62-34aad40b7ee1 · outbound

This paper cites Numinamath.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Numinamath

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.474997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.902199Z digest=sha256:b3cfbf966a74d7546c2c2e049dbf2a356a8eb854856a7de86f0e4c339d8c34b2

Observation 7eb29bb5-7887-44c3-99d8-256f31919b46 · outbound

This paper cites Probing out-of-distribution generalization in machine learning for materials, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Probing out-of-distribution generalization in machine learning for materials, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.462392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.905837Z digest=sha256:b8e67a3faeab1038639f11c888a057db4c96f9a8b48b36bb54aea62d60947ad2

Observation 0bd7f082-a1e4-450c-89fb-6ab849c9fd9a · outbound

This paper cites Gsm-plus: A comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Gsm-plus: A comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.448652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.909496Z digest=sha256:c35431ca42a43d3016e6884d05022b4d611be668485d3974f2147bbe6e9c4109

Observation 7fcdc50d-c832-496e-ba54-3ea34ce3aef7 · outbound

This paper cites Towards out-of-distribution generalization: A survey, 2023.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Towards out-of-distribution generalization: A survey, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.913037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.913037Z digest=sha256:23893f481b8b8931768170ec019e70edacaeff15e78f540f2c19ed01347b5c5e

Observation fb35d26b-a10e-49c6-a693-7e217cb88bdb · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.916405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.916405Z digest=sha256:c76e3053da1fde2da66395fdebd0539df594b7e4aac03aab1fff88ee096d5008

Observation 21a9a9b5-b234-4026-87ce-c265df39b69d · outbound

This paper cites Compositional generalization by learning analytical expressions.Advances in Neural Information Processing Systems, 33:11416–11427, 2020.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Compositional generalization by learning analytical expressions.Advances in Neural Information Processing Systems, 33:11416–11427, 2020

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.429281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.920220Z digest=sha256:24ea04b3d734c454259da60d5baff56e6e1bfb5f9aa3ca6e18f604d951f51b43

Observation 841e4e97-a348-48fe-aecb-6d8bea74d693 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.924842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.924842Z digest=sha256:d762784ef125bfa682ee05abeb3b9ad318ed8a87b3168ea4623a5d4910571d8c

Observation 72cb1a5c-fd3f-428f-9223-edf8b72ec7fc · outbound

This paper cites Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.929508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.929508Z digest=sha256:df1b96f98fc909bd6c822f34319f2722a5dfc25f584b6ec183a94ae16d52adb2

Observation 4b06513b-cc66-4bb4-b67a-31241f400e33 · outbound

This paper cites Gsm8k, April 2022.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Gsm8k, April 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.416982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.933504Z digest=sha256:f8639c944ed50c06c88d63d76330b88d55aa41ed944940698186381c9aa4e497

Observation dbafe4e1-dd2b-40f7-b86f-8ae33e5b2df9 · outbound

This paper cites Learning to reason with llms, September 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Learning to reason with llms, September 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.937329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.937329Z digest=sha256:ed9911d9ce990fd7a825201000b3f54b50f7abee17aa26e3209ecfe8cc7bcc32

Observation 85be693b-36fb-4318-8dda-bf18b6ead1cd · outbound

This paper cites Math 500, November 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Math 500, November 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.398067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.941089Z digest=sha256:96789ff8797e21dcbbe0c3926bbb1f0446c122d41d03ac31b6d9b4caad992094

Observation 2d6a2901-5fb8-428e-8cc7-bad3e1d337dc · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.946217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.946217Z digest=sha256:a3c7cbdace40e1569bc9e62ecdcae94e2a8045820ff5f3788f393c69b0bf5a60

Observation 7e2fc55c-16e6-4ef7-9a82-cb309150cae1 · outbound

This paper cites Climbing the ladder of reasoning: What llms can-and still can’t-solve after sft?arXiv preprint arXiv:2504.11741, 2025.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Climbing the ladder of reasoning: What llms can-and still can’t-solve after sft?arXiv preprint arXiv:2504.11741, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.950314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.950314Z digest=sha256:48181b2595b63cbb636e03583d67cc7cfb0db7b9639f04fa12c3ab2fcb30745b

Observation 82d3b542-5ad7-4c43-9036-ce8a3ed41083 · outbound

This paper cites Mathscale: Scaling instruction tuning for mathematical reasoning, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Mathscale: Scaling instruction tuning for mathematical reasoning, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.953862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.953862Z digest=sha256:451a525630358d20e2699bff2778aece4bfdf77e17e7e866931eb33e6d4d7221

Observation c4cb4824-27e4-44ad-8037-8f70e4f00768 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization The claude 3 model family: Opus, sonnet, haiku

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.375061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.957314Z digest=sha256:5affc41aa613bc3e2b64eb74c953b1fa994abd3d61d9cb9d506325e6ae2e084f

Observation 3657a068-ad2b-497b-8d6e-d34aa6463fc0 · outbound

This paper cites Openmathinstruct-1: A 1.8 million math instruction tuning dataset, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Openmathinstruct-1: A 1.8 million math instruction tuning dataset, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.361658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.960941Z digest=sha256:38c3b6f66ba61d0f2d3894366dba2f8afbcf97c6c6a662bd12c199ae8528f3a4

Observation 632248f7-a2ae-41d6-90b4-cc393bba50f5 · outbound

This paper cites Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.964503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.964503Z digest=sha256:2c690b5c2ddb60df77e197b9a5ca51d769124eda166faac8ab3cdf265c6689f6

Observation bd17a92b-19b7-4ee2-98c0-eef0e8ffdb1a · outbound

This paper cites Towards a theoretical framework of out-of-distribution generalization, 2021.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Towards a theoretical framework of out-of-distribution generalization, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.343032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.968152Z digest=sha256:c4c68b352ebbdfa02148a2076fa75a95b4ea0702319090211eac4280b46305a2

Observation 147c34e6-d020-4e39-8219-ca0b8bd18de8 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.971551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.971551Z digest=sha256:691b5010d6af225f2d81d8cf72dd1cab10d553640d4d1d7c2bea9135077dca80

Observation 073e41ac-1ee1-4177-9364-f3e7c34b5d71 · outbound

This paper cites Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.975316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.975316Z digest=sha256:f3d290ffa7c755db541568effea58251a71310670006fb6224716b43973ac03e

Observation c44c3d13-2f48-4716-9def-b3198ef423e5 · outbound

This paper cites Evaluating the performance of large language models on gaokao benchmark, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Evaluating the performance of large language models on gaokao benchmark, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.330336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.979227Z digest=sha256:c4821a691a1b85b78398039080a21b47ee0ddc0be4657b903b660f9c7ce7a113

Observation 8b282595-51b0-4b30-a444-9e848b9a909a · outbound

This paper cites Can models learn skill composition from examples?Advances in Neural Information Processing Systems, 37:102393–102427, 2024.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization Can models learn skill composition from examples?Advances in Neural Information Processing Systems, 37:102393–102427, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.314082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.982617Z digest=sha256:2f5b8c1a9f94dcf90d1b8ca67f61741be76400998a7275eb41d874b5b53c5efc

Observation 7527bc67-e77d-4288-a537-56d461fca4fd · outbound

This paper cites GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:57.986232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:57.986232Z digest=sha256:947cea76fbcfe07e698e2b53d1340fd16053a0c1aa68a47a48f0a561fc9ffe95

Observation ed18e313-d1ea-4dd5-86b5-a08ca1df6ff2 · outbound

This paper cites visualizing.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization visualizing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.299553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.991015Z digest=sha256:e4bb0428ff9e5f9dd185273f31c9c35d6395b2e67ea24e551a3f956acb85aeb5

Observation d56b2f23-c352-47a0-b072-177155a845eb · outbound

This paper cites conjecture.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization conjecture

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.287433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.995535Z digest=sha256:41e8858b5e57a64f1bda758c7b5b834ed3283eeff7c7bb658d0f4d15457f2b70

Observation 3b4428ef-c97d-4243-947f-abcbfcf47d56 · outbound

This paper cites computation.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization computation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.274551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:57.999597Z digest=sha256:6c26fb2ab35856df0d8cbfaddf6171bf7d17af30231167251b8821fe88985f21

Observation fed93d16-0519-4544-8441-bb4121582713 · outbound

This paper cites computation.

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization computation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:58.262364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T18:46:58.003217Z digest=sha256:74b5316f21add1ad1a88b72398d639369a513b968cb636af963283053fa01eef

Pith citing papers

Observation a4955607-c80d-4a2a-8e1d-ba00be34a93d · inbound

Rethinking the Illusion of Thinking cites this paper.

Rethinking the Illusion of Thinking OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:39.097178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:03:39.097178Z digest=sha256:88eb64d0ce2b3732134691ee8e599e5001ab64e0a844038a05d65b6f898a035b

Observation 32cb581c-3213-4fb9-b9df-af707d223f04 · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:16.918871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:16.918871Z digest=sha256:4c040afb6a5fdaa6bb6c499be01ab6f00062e502d03c17744ab719778106fdb6

Observation ff53744e-1a34-46c0-a3cf-94671faf3c08 · inbound

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments cites this paper.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.137256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.137256Z digest=sha256:5c9e06ed8628915e78f26ef65a72e48a565fca85bf54e3682293841f3931b661

Observation 1356ad6f-4a21-4a9c-bcc5-88494310d421 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:56:11.366338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:0852e44078ef1b706ecd7ae5ce39134174b03a156741dffcab1985aa3f8464ef

Observation 2c7fef75-70e6-49ce-a412-f59312a560a7 · inbound

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers cites this paper.

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:50.818472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:11:07.615149Z digest=sha256:29ef0e59bf6faf2a68b57ede2945675017bb424329239644fd5f5223092a7acb

Observation 4afce806-873a-4d04-be21-9a925ef2d5de · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.904520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:7b4f379d8b8ce7b5541d4430d35befe41de1ce55e2d2ff8880dd634e29e938af

Observation 3ab79f2f-46a2-43a7-aa3d-7b632af45827 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.449215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:cf8f974cc9c1f009f5be478da73e1ecb29443f717860cc521f5f20372f35fcb0

Observation 52890016-2d81-4841-8b4c-b804953a11da · inbound

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages cites this paper.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:15:30.210152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-08T19:08:06.748619Z digest=sha256:00da6ebb596ec7714075a0a15a51cfa3ad1ca14cab8baf1eb569b3d9ca4a639f

Observation bb9418fb-fd26-484e-91b6-02a8fea9c261 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.759751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.759751Z digest=sha256:d85d383f3720c36adff8c0387b483b9cdbd51294b28bb49f1422a06da9fcd6ee

Observation e3115606-8d7d-4b7a-9c8c-e296f1eaddaa · inbound

Ask-E: An Environment for Calibrated Question Generation cites this paper.

Ask-E: An Environment for Calibrated Question Generation OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:11:50.688938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:11:50.688938Z digest=sha256:60f4f24c88c6e4f891637a72c8739687f5dcdbfed9ff4990cb7007011b659452