Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:51:21.341349Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2505.08744.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:51:21.341349Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T09:45:18.385925Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
14 of 14 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 07b94738-330a-460b-8a3c-abe63e78e2e6 · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Counterexamples in Real Analysis
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03b1ffa7-01f3-41ff-9ea6-d6ac25abcada · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Training Verifiers to Solve Math Word Problems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979e5bd1-3c6a-4f60-9573-ba7d0b2a1a11 · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e076a5ee-e88e-4890-be4f-2145f48bf88b · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9780936c-0ab3-4128-ac96-c6b42d4a7b81 · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Measuring Massive Multitask Language Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7690da8f-5cc7-425e-b58f-771d6de5c9e4 · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7f3766-cdd2-4412-9a80-13b9463cdabe · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Aime 2024 dataset
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb0957be-cb23-46cd-a2e5-4dc409b753e7 · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models On manifolds homeomorphic to the 7-sphere
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f0ae241-59d6-4897-883c-b4011385568a · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models O3 and o4-mini system card, 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33def663-f585-4cd1-a77b-8c762ccdecdd · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2f3467-345f-4bb0-9b6c-076498c4c52b · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models The V alue of Science
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f94f927-d9ba-4378-8eeb-99381142dc6f · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Comparison of three large language models as middle school math tutoring assistants
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0403dbf8-d42e-42d5-82cc-923d5a4eb6af · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Trinh, Yuhuai Wu, Quoc V
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 470fa610-0cb8-42c2-961d-20a5cd940218 · outbound
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models Über continuirliche funktionen eines reellen arguments, die für keinen werth des letzteren einen bestimmten differentialquotienten besitzen
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c7aa15c-feeb-44fa-a3c2-7bafb674320d · inbound
CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c28464a8-26be-4c94-af5b-35a30cf38bd5 · inbound
Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.