Pith. sign in

Paper Citation Record · LEDGER

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

As of 18 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2505.08744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08744 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:51:21.341349Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T09:45:18.385925Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 07b94738-330a-460b-8a3c-abe63e78e2e6 · outbound

This paper cites Counterexamples in Real Analysis.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Counterexamples in Real Analysis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.527925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.291082Z digest=sha256:9d5ff925799412a6363a6aabd044760b44e28c6f3ad9745d1965432a0062e2ee

Observation 03b1ffa7-01f3-41ff-9ea6-d6ac25abcada · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.296102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.296102Z digest=sha256:d7921026a3102e82aae5d156472e5d27ff2546bb225847b005785fdb55cb35dd

Observation 979e5bd1-3c6a-4f60-9573-ba7d0b2a1a11 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.300560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.300560Z digest=sha256:b41e20daac7081dd3f35139ff8338cf18e5cf8060fa506e08e27f06516069895

Observation e076a5ee-e88e-4890-be4f-2145f48bf88b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.304758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.304758Z digest=sha256:5b2b3260c28bf31046a150a20a939bab4e65a8f7d5efad1973f70a3a0b31d640

Observation 9780936c-0ab3-4128-ac96-c6b42d4a7b81 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Measuring Massive Multitask Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.308527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.308527Z digest=sha256:5a1366264598aef3fa2066d115ad2546d6ef80055c34f6ab169f56eaa9fe1ff2

Observation 7690da8f-5cc7-425e-b58f-771d6de5c9e4 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.313111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.313111Z digest=sha256:ce06590e5d31f49158f15a2d4f5161807fc613f60a31ea173600b14b8a39f1b4

Observation 4f7f3766-cdd2-4412-9a80-13b9463cdabe · outbound

This paper cites Aime 2024 dataset.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Aime 2024 dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.317033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.317033Z digest=sha256:caf74ab51e810410ed389e6729898f0c6115a675f6a253822fbe60af32397a3f

Observation eb0957be-cb23-46cd-a2e5-4dc409b753e7 · outbound

This paper cites On manifolds homeomorphic to the 7-sphere.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models On manifolds homeomorphic to the 7-sphere

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.508438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.321186Z digest=sha256:34f450e9f7f3a2b63928dc14a6cbf70f4090d082ae27451654211be1314c36ce

Observation 0f0ae241-59d6-4897-883c-b4011385568a · outbound

This paper cites O3 and o4-mini system card, 2025.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models O3 and o4-mini system card, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.496563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.324522Z digest=sha256:0c2a672d9940a4607bdafc22a35db8c30ca418eaacb4da33419048d16c3c5129

Observation 33def663-f585-4cd1-a77b-8c762ccdecdd · outbound

This paper cites Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.327994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.327994Z digest=sha256:740ca261c121a0be76e0d3cb21dbe1405495f64a8d6e8ea4612365e5288e0492

Observation 6f2f3467-345f-4bb0-9b6c-076498c4c52b · outbound

This paper cites The V alue of Science.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models The V alue of Science

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.486148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.331732Z digest=sha256:dde6ae2612a16812d36b0a9f82c2ede1665eb57525979bb8b83dad93bbd20936

Observation 9f94f927-d9ba-4378-8eeb-99381142dc6f · outbound

This paper cites Comparison of three large language models as middle school math tutoring assistants.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Comparison of three large language models as middle school math tutoring assistants

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.474120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.335033Z digest=sha256:694ef9605af93633cfccc50263e4d026177c3454bacd43720ceb7b5697843170

Observation 0403dbf8-d42e-42d5-82cc-923d5a4eb6af · outbound

This paper cites Trinh, Yuhuai Wu, Quoc V.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Trinh, Yuhuai Wu, Quoc V

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:51:21.462162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.338015Z digest=sha256:38987ffe7b8f6b4c9a09ced88a8b6ad4011146c98c86d22ab4a9a815ae490a2c

Observation 470fa610-0cb8-42c2-961d-20a5cd940218 · outbound

This paper cites Über continuirliche funktionen eines reellen arguments, die für keinen werth des letzteren einen bestimmten differentialquotienten besitzen.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Über continuirliche funktionen eines reellen arguments, die für keinen werth des letzteren einen bestimmten differentialquotienten besitzen

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.450594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:51:21.341349Z digest=sha256:4cc0a6160788c3d686613a636b653593b8b7631b239434a8590e3c5fc6b7272a

Pith citing papers

Observation 5c7aa15c-feeb-44fa-a3c2-7bafb674320d · inbound

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions cites this paper.

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:32.301871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T09:45:18.385925Z digest=sha256:473f01e4930ba21af3606fb6697e32bdd9e8318ad502eb5ec340bf6b1e0bc073

Observation c28464a8-26be-4c94-af5b-35a30cf38bd5 · inbound

Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing cites this paper.

Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:11:02.502588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T16:09:08.317910Z digest=sha256:2b289b36279305f40d6526ee8fda608eb8f6a1aecafccd9e97cd1e874f51d569