Pith. sign in

Paper Citation Record · LEDGER

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

As of 20 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2505.08744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08744 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:51:21.341349Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T09:45:18.385925Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 07b94738-330a-460b-8a3c-abe63e78e2e6 · outbound

This paper cites Counterexamples in Real Analysis.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Counterexamples in Real Analysis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.527925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.291082Z digest=sha256:0b01b871fad56dd60755cef69c6aaad0d0e74baa905b6736a68fe3802046e9b7

Observation 03b1ffa7-01f3-41ff-9ea6-d6ac25abcada · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.296102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.296102Z digest=sha256:d7921026a3102e82aae5d156472e5d27ff2546bb225847b005785fdb55cb35dd

Observation 979e5bd1-3c6a-4f60-9573-ba7d0b2a1a11 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.300560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.300560Z digest=sha256:b41e20daac7081dd3f35139ff8338cf18e5cf8060fa506e08e27f06516069895

Observation e076a5ee-e88e-4890-be4f-2145f48bf88b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.304758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.304758Z digest=sha256:5b2b3260c28bf31046a150a20a939bab4e65a8f7d5efad1973f70a3a0b31d640

Observation 9780936c-0ab3-4128-ac96-c6b42d4a7b81 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Measuring Massive Multitask Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.308527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.308527Z digest=sha256:5a1366264598aef3fa2066d115ad2546d6ef80055c34f6ab169f56eaa9fe1ff2

Observation 7690da8f-5cc7-425e-b58f-771d6de5c9e4 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.313111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.313111Z digest=sha256:ce06590e5d31f49158f15a2d4f5161807fc613f60a31ea173600b14b8a39f1b4

Observation 4f7f3766-cdd2-4412-9a80-13b9463cdabe · outbound

This paper cites Aime 2024 dataset.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Aime 2024 dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.317033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.317033Z digest=sha256:caf74ab51e810410ed389e6729898f0c6115a675f6a253822fbe60af32397a3f

Observation eb0957be-cb23-46cd-a2e5-4dc409b753e7 · outbound

This paper cites On manifolds homeomorphic to the 7-sphere.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models On manifolds homeomorphic to the 7-sphere

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.508438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.321186Z digest=sha256:2622cb0b0635c1b84bcd0e56ad03daea614b9ea297f5465b53ffe45659e4823b

Observation 0f0ae241-59d6-4897-883c-b4011385568a · outbound

This paper cites O3 and o4-mini system card, 2025.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models O3 and o4-mini system card, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.496563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.324522Z digest=sha256:469f7e016e552ebc6d74c119a88ae22356b7abb55efa372d3a0844f9503200be

Observation 33def663-f585-4cd1-a77b-8c762ccdecdd · outbound

This paper cites Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:51:21.327994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:51:21.327994Z digest=sha256:2188cb4a8f424ad03a9bfe2654f9ce0f2d5a2f68358c8cd9cd0225b4aaea79b7

Observation 6f2f3467-345f-4bb0-9b6c-076498c4c52b · outbound

This paper cites The V alue of Science.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models The V alue of Science

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.486148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.331732Z digest=sha256:6a9bed8aa27103bec16d691aac87a491bef091619f25e8913516668e5c20d878

Observation 9f94f927-d9ba-4378-8eeb-99381142dc6f · outbound

This paper cites Comparison of three large language models as middle school math tutoring assistants.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Comparison of three large language models as middle school math tutoring assistants

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.474120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.335033Z digest=sha256:a233da7c04cbc4bc0a35ae1bb005559da8c2a3ea1a68a7b24f72439fd25fa790

Observation 0403dbf8-d42e-42d5-82cc-923d5a4eb6af · outbound

This paper cites Trinh, Yuhuai Wu, Quoc V.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Trinh, Yuhuai Wu, Quoc V

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:51:21.462162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.338015Z digest=sha256:bb8cd649fce82daca5e6c1b7ae1f833dda3514b1189f845095c939ee559ce9de

Observation 470fa610-0cb8-42c2-961d-20a5cd940218 · outbound

This paper cites Über continuirliche funktionen eines reellen arguments, die für keinen werth des letzteren einen bestimmten differentialquotienten besitzen.

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Über continuirliche funktionen eines reellen arguments, die für keinen werth des letzteren einen bestimmten differentialquotienten besitzen

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:51:21.450594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:51:21.341349Z digest=sha256:25d3c8f342c05ec91ef3fb0fdd39bf53b75601cca1209c0c06b6d37308274c13

Pith citing papers

Observation 5c7aa15c-feeb-44fa-a3c2-7bafb674320d · inbound

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions cites this paper.

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:32.301871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T09:45:18.385925Z digest=sha256:5b92dcf52da57c62a50de6b1c9e9b90ab3767be20974007565e7d2de7608fdca

Observation c28464a8-26be-4c94-af5b-35a30cf38bd5 · inbound

Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing cites this paper.

Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:11:02.502588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T16:09:08.317910Z digest=sha256:8ba89ca1ca4634cfd2510ecffebd182c4802cd39b854a62c051c94cbb24fd5ee