Pith. sign in

Paper Citation Record · LEDGER

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

As of 15 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 4 inbound Pith citation observations for arXiv:2501.13766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13766 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:40:39.950474Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:27:46.891540Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:25.670647Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 112c66d7-04a1-4d1b-a21a-e7d108d90a83 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.544581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.544581Z digest=sha256:26b5638486fc300850b102b5db1cb3498edfb1e90880821a27c7fb8b61f4f431

Observation 53e3c48f-fd35-40f4-b26c-8fc824980633 · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.550749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.550749Z digest=sha256:bce3b7c1445b8df0a7ed930c6cdd4a0c48d70cdae6754af16d8a66ac5a12d45d

Observation c91449ec-0e50-4a9c-b3b5-43239d8bd634 · outbound

This paper cites Llama 3 model card.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.556201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.556201Z digest=sha256:78b59aa43b23cdccbe5c41605c179d2f0d57d13a0e29ea1282d738d4d7bc5414

Observation 04d744b4-b66d-4209-832e-ef129bbd370c · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.561212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.561212Z digest=sha256:420797e65c7e99ee56c177ea4b25212ed01e02476c348da0caefc2770891f2d6

Observation cedec0c5-2d90-4b56-97d7-d3461e4a3b75 · outbound

This paper cites Claude 3 family.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Claude 3 family

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.307721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.566715Z digest=sha256:f3a94de07dc5ad5bb296eb4b821ba533df6947e9309f5f112d9b14240846a86a

Observation 83950808-6d63-42db-b4f7-8fd3b7983326 · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Llemma: An Open Language Model For Mathematics

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.572302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.572302Z digest=sha256:c453d71eda2412f1cf71d54a909340a193cf183a1c27e36064a3413005cd9b1c

Observation b537a3de-8faf-455e-836c-f489bb6c4a66 · outbound

This paper cites Numinamath 7b cot.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Numinamath 7b cot

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.577730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.577730Z digest=sha256:87e4f49d581ff649f7e1f176da3698f05b76001afcdcfe7fed6ec71b7bbc59cc

Observation 8f23a55c-7aeb-46a0-91cc-5793641ba7ad · outbound

This paper cites Natural language input for a computer problem solving system.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Natural language input for a computer problem solving system

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.281044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.584189Z digest=sha256:238d3383933941cb1ab4fc43c48e114f410244e8223d4835865e5b6bc450735d

Observation 0c29df4c-4659-4dda-bfc5-48162fea5bca · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.589484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.589484Z digest=sha256:7877a76a30f6594969e72ce170252546fa5f34874bacc6c15e85a9d6cf748774

Observation 2c239a0b-3c6a-4e05-9ba8-de94667fed2e · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.594235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.594235Z digest=sha256:a82bbccb5fd345961bcb27089dd2de2bab62e1a576b770165bf039a4646e4767

Observation b1c5f3d2-255d-4d38-9d90-b5c2442894f2 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.599408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.599408Z digest=sha256:a7721cc3c4ae4fed0bfaf3a0436338677906538b1d106a43466f4ad5e568051a

Observation 9ffa50c7-3651-4513-b136-2d80c55f9f43 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Theoremqa: A theorem-driven question answering dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.253873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.604841Z digest=sha256:f10ffd43a1b68e79af28092b1ca2e57ee6a76441eb6090dcca4d9de3231d3421

Observation 3d0553c2-64d4-40b2-981f-03f6f2cacf01 · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Premise Order Matters in Reasoning with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.609708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.609708Z digest=sha256:caf177fe1c4561e7a358642320de01e0e2df604b80bb78aa5939cecb92a4da9a

Observation 4b5d623a-009c-45bf-98fc-8027e215b48c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.614608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.614608Z digest=sha256:1544a623450dfdb61f1d8a30decf17e3b6718273b6599ec3c2b08c4cd41d284d

Observation 00d1b62a-186b-4e91-9f97-523cf32dc4fb · outbound

This paper cites Evaluating language models for mathematics through interactions.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Evaluating language models for mathematics through interactions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.236541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.619641Z digest=sha256:85edd7f2122ae679fd075501e4b180d259a486ce1be17592d55ed2e8e2bb2e4d

Observation b56fc7bb-e869-42d5-8e8c-bac6c4a40447 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.624353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.624353Z digest=sha256:b569f2419386996de616169b099388996d054230bbb29534ae76ecbb4d18481c

Observation 2ffca2f8-04ee-4faf-80dd-38c1b1d26ceb · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.629739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.629739Z digest=sha256:ceef7f1cba0afd50ee42b1cad6bfe9b0bf73cce80222a93afc4aeaac658ddb7e

Observation 02230e92-85ab-4b48-af74-053c66d8d60d · outbound

This paper cites Investigating Data Contamination in Modern Benchmarks for Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.634533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.634533Z digest=sha256:4f4a798f1d9a079e0ba0585d119fdc457910520ec7c59b67b062233d78634eec

Observation c3126efb-dc84-4429-b7ff-800f2f3c8d47 · outbound

This paper cites Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.639354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.639354Z digest=sha256:dc26a89d25af5e57d5c65712c6e5c70bf85fcad87ad474fa5128c4469f03bc2f

Observation 077e38fe-1be2-4295-a652-7ca1a967eb0a · outbound

This paper cites Time Travel in LLMs: Tracing Data Contamination in Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Time Travel in LLMs: Tracing Data Contamination in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.644278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.644278Z digest=sha256:76c4f96359785b36081621ea357d0e759e529212259419c76fd9c96bae0b082c

Observation eabd3cea-40ad-42db-a1e3-f89c85a61189 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.649051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.649051Z digest=sha256:2a3b7def3ca04847113b924be57cad404c864ce38de24d1b592d70f1e53a8792

Observation 88227e78-1063-4d19-8980-340995b03504 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.653756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.653756Z digest=sha256:019a28f4437917b814cb2f1f47f0dcd8f620c8f3935870e381414313ba630d7c

Observation 5880b332-0ed5-4872-b79c-90421489126f · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.658649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.658649Z digest=sha256:8538bf65b9e72e9b507d35dddc4b785cf58488b542a6e00cd1aa0382bc834daf

Observation f6f556c5-0b3b-4e02-8c07-d2fca9794773 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.663637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.663637Z digest=sha256:a7a89b4c042f1b1928d6b224a8649905ff218a104dfb6304667061ba553824c2

Observation c65a5993-c6da-4384-8b0e-fe79d8886ba3 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.668864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.668864Z digest=sha256:a4b3065fbe13bb8d695dfde6c443feaf28fdf53280be4a2f722bbdcbead3de01

Observation b3d0c280-304c-49a7-a989-ab3e5fc7b561 · outbound

This paper cites OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.673875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.673875Z digest=sha256:d51374a9b07d3b73de9aa5b6cc0982f515b427d82feaf1e6ad7eb68ab80ad770

Observation f8781234-6a3c-4196-bc36-f758951902b5 · outbound

This paper cites Mistral 7B.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mistral 7B

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.678599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.678599Z digest=sha256:b8e979a591583c688bf7ed73ece417722070186b2b3337605184c957490d9309

Observation 9db68948-0a87-4acf-9eee-84d5d90e3d3f · outbound

This paper cites Investigating Data Contamination for Pre-training Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Investigating Data Contamination for Pre-training Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.683641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.683641Z digest=sha256:e848e0a3f9ebd07aa60ee640450584ef7b8b4992db17dac0ed7ecef3a1322f44

Observation 9078c86e-1a07-4823-8cd8-d942510a60f6 · outbound

This paper cites Large language models are zero-shot reasoners.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Large language models are zero-shot reasoners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.688524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.688524Z digest=sha256:228908dfd0475d509567e30d29ea11fd4057f9b0b9cc5e21407c0acd0d681ba5

Observation 0dfad2ef-30b3-40ce-98e1-586cb125aa77 · outbound

This paper cites Mawps: A math word problem repository.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mawps: A math word problem repository

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.195779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.693132Z digest=sha256:21aee394a2a015f9bac73e4dfa1a0c652182a17098ee98fcdafd255a088b2469

Observation 74f166fb-caf5-4f6e-86cc-e6d481253606 · outbound

This paper cites Solving quantitative reasoning problems with language models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Solving quantitative reasoning problems with language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.697680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.697680Z digest=sha256:96b594b6ec2ddb31820ff86ad21ab1e2ec63147a91bfdbcd2feffcbe58364e64

Observation c921ab3e-8305-40ce-ac3d-f7a0b99bd94f · outbound

This paper cites Common 7B Language Models Already Possess Strong Math Capabilities.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Common 7B Language Models Already Possess Strong Math Capabilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.702855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.702855Z digest=sha256:4d2555c813adca3f2fe97b5c38e471dde2bb07ddfc2f3d4752ca24c790683beb

Observation 3aa96eeb-b20d-4d50-bdeb-84b6ff0231e9 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.707901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.707901Z digest=sha256:f9dac12c707f9b01b27da8f5fed19a4c71b74b6b5253150f5667529d2cb707c8

Observation 433d176f-0b6d-451d-8053-e732d9c6b62f · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.712585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.712585Z digest=sha256:463720185b6a1d84f03f9eb70994a3ba92fbc24480ef8db728822caec9092e8c

Observation 37d24616-1827-498e-a9f2-0b7911c0c595 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.717433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.717433Z digest=sha256:9f135c8d251a32951e7f3301d3dd7d76eb1dfb39cf558f9b99383868c06fe0ba

Observation ccbf91a8-5b0e-41c4-9b04-f03f8f0a8a8c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.722180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.722180Z digest=sha256:5f8ffca711c828fccb0daf87d42c6d1eca759c53cb80c610455cdcdec12ab249

Observation 14c2d6a3-7291-4e88-ade2-344cbdc376f4 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.727254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.727254Z digest=sha256:6dc26719114f6b4a74289affa9f571034132789d967a2c8d98c97e3cb449c2bb

Observation 599c6f31-cc86-45d9-93bc-32d5d4cbf542 · outbound

This paper cites Fairness-guided few-shot prompting for large language models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Fairness-guided few-shot prompting for large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.169782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.732105Z digest=sha256:efe20b67c9a4af4238bc5b95aff45fa6724df208f188c33c7264a38aeb028b98

Observation a38806a0-1aaa-4941-8bc8-f89d5f7abca2 · outbound

This paper cites Mathstral.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mathstral

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.736586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.736586Z digest=sha256:f65e45112c3f50c769440650404ec8d9c6cfcc6a44fcea451178d701fd0ea49c

Observation 2c317ae4-a445-4d8e-8806-c85647dfd5a3 · outbound

This paper cites Mistral-7b-instruct-v0.3.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mistral-7b-instruct-v0.3

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.142404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.741261Z digest=sha256:a65a9ef5b78b4b71957b50247d626bfd5074ba7fb90bc9b7697f3609ec0581b4

Observation 3b1b4e39-5e71-4bf9-94e3-4a461cc34d4d · outbound

This paper cites Mistral large 2.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mistral large 2

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.124881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.747037Z digest=sha256:65dba68fc94077956d3366d01431671b0b403fec4b67577cf028187bd4e2b32b

Observation aa9bbae8-0374-4506-9469-32bbc1d62eff · outbound

This paper cites The future of ai: Trends and predictions.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models The future of ai: Trends and predictions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.107384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.751486Z digest=sha256:050f34ff200191048b394ac91c42e36808107d7c62c21c29505413d15e23b9b0

Observation 459c3e02-77f1-413d-9cdb-203c0485619b · outbound

This paper cites mistralai/mistral-small-instruct-2409.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models mistralai/mistral-small-instruct-2409

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.090764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.756854Z digest=sha256:9249164e34a36b344fca4b47835f6f95cf61744cad52a0a7f73cdcc98e219c99

Observation d97bf66b-9ffc-40ba-b4a9-771b8196c7ff · outbound

This paper cites GPT-4 Technical Report.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.761299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.761299Z digest=sha256:a6436a113bc00bf4f30e8a09e7cf59bda3196429c169d1c8c4cc489825865160

Observation 457d8c92-56a7-459c-9e29-df681c5c9e25 · outbound

This paper cites Hello gpt-4o.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Hello gpt-4o

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.765864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.765864Z digest=sha256:b96c19c43844a91cc0d455339d2733d7d1253db3e8b968822fc56d402bd05953

Observation 21421130-e9a3-4aa4-8fe7-5d00c6be18b3 · outbound

This paper cites Learning to reason with llms.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Learning to reason with llms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.770409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.770409Z digest=sha256:00d5f0471333d6e1a2e2d9f8545a9ac60273ebace34e384747a021dfddeb6b04

Observation f11b28a2-617b-4ab1-af5e-66097d738444 · outbound

This paper cites Training language models to follow instructions with human feedback.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.775015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.775015Z digest=sha256:bd413a184f09a0bb947585f61d4da84555c57d81f578c5ff9da58d643fec0305

Observation 024796c3-3452-4db2-ad50-c22106483284 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Are NLP Models really able to Solve Simple Math Word Problems?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.780365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.780365Z digest=sha256:6d5b716ca4336e0bb1ee2e562c604e03107269a0c7b48727162eb5ab402fec49

Observation 68217314-83f4-4253-a44b-77855e4a987a · outbound

This paper cites VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.786212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.786212Z digest=sha256:faae639fd4a96294ac262c5b7a4dcc4535144d2c59268d636bf54a7b71e68891

Observation a0d1a685-4a18-4770-9069-0b01d45904ae · outbound

This paper cites Impact of Pretraining Term Frequencies on Few-Shot Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Impact of Pretraining Term Frequencies on Few-Shot Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.790994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.790994Z digest=sha256:20342d3fbd6cede134f1f09a61c1b0a0bd6247cae1a29167f1436efac4e3951f

Observation cce3f2c0-f7d3-4efc-bf83-9852a15f3771 · outbound

This paper cites To the cutoff.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models To the cutoff

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.042710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.795631Z digest=sha256:2512b12caafabce3268a8c092e241a3bd2b84ad5b5e22202238e4ca56e28b8fa

Observation 1a048196-af5d-4ebc-a63b-cc9e3787ae95 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.800241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.800241Z digest=sha256:49b771da40c76e901be90454c4af0dfdfbcbf110d9c7b9e54d9fb5a8f29958d3

Observation cf56f8a0-24e0-4295-98c7-21a16c8832ae · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Large language models can be easily distracted by irrelevant context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.805110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.805110Z digest=sha256:8ca70dc961c252735e84a7b0bdcbd504aea5b6a2b4f0bee5767172a6cb14c73d

Observation a7e5437e-8496-4c30-8a34-ceca47b43642 · outbound

This paper cites Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.809872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.809872Z digest=sha256:cf460468ba36d1ab998b85a576b752266415fb4e52478bce653403a31117f67a

Observation ad55903f-86bb-4535-becd-65104a268541 · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.814540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.814540Z digest=sha256:3c9fb5e740f16c31fa97c61c82ffa34bcaa82b72982f76335f05a2dd959aa972

Observation 934e076b-8366-41db-af36-00a957359448 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.820487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.820487Z digest=sha256:34025033160fe383b18e81b692f77af22465fcb37228d0025ddfec1116a3ed42

Observation f2f0a032-c36b-4daa-8248-92a6c1db513f · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.825147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.825147Z digest=sha256:ac9c0157b70372ce5d9cde4641add360f266285825fc44289b9d2761146479bd

Observation 2a1c0d03-b390-4700-83c0-5a4ed1f5084e · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.830578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.830578Z digest=sha256:a2d321702bdc701d7df6e8b4f77848f09e84a700e9084ebde2766964c0491f5b

Observation 89e4f12c-0e1e-45fe-a848-9dfbc2b60afe · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.835782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.835782Z digest=sha256:b81b8161897fca836ffc7b5749ca86960b4e190372d70f8bdf7d1eb95680ef20

Observation 1a947438-dea3-4669-ac75-a59931a82729 · outbound

This paper cites Deep neural solver for math word problems.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Deep neural solver for math word problems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:41.015819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.840532Z digest=sha256:b3f12fd0c6c7f943f51d08600fc86ed8cf0bed34439dde8f053c0038c191b6a1

Observation c01740dd-8ba5-46d9-bd61-ac4061053c35 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.845175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.845175Z digest=sha256:7d96c9f03c91ba0e8b6ad41376870082c3afec3c21bb655bbc110e4c66d5e787

Observation 2cb89d01-d296-4229-9d77-6f07b88bb2ab · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.850504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.850504Z digest=sha256:3767a460d6b806c470e5908d6f9cffdbf47958dbf8d3b9426510f91abf58a86e

Observation ec7cec21-656e-4d3c-bdc1-f3db0a5496fd · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Crowdsourcing Multiple Choice Science Questions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.855073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.855073Z digest=sha256:c5c957fa227acc815753da25ae9ee42949308818455da913440a5e63637a0476

Observation b7273789-96d0-423a-b457-a92f31e57fcf · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.860944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.860944Z digest=sha256:26d3b6e16d5ac489a90516425dec8289691cb23c064962c0d785d0c05d718eb6

Observation 9cd00303-8f6a-4d48-b38c-ded731bf435a · outbound

This paper cites Can We Verify Step by Step for Incorrect Answer Detection?.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Can We Verify Step by Step for Incorrect Answer Detection?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.866556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.866556Z digest=sha256:b4bf1cf1cf4d11bb02649f138b18d9b91f5ca316923b0c4097977cbf74ef72c7

Observation 23f0feda-a3e9-42b3-95f5-ecf2217b8a83 · outbound

This paper cites Can LLMs Solve longer Math Word Problems Better?.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Can LLMs Solve longer Math Word Problems Better?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.872020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.872020Z digest=sha256:e64408fab6d92b1b50f8f4a343de0c2f53816e4978f773c3a3c903327054abe6

Observation 2220d358-a814-4ef2-94e9-06802039777d · outbound

This paper cites S^3cMath: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models S^3cMath: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.876942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.876942Z digest=sha256:433d49f85ab61ef80cff8b51919619d9542edb46eda430a75dfe4f53ed85a007

Observation 026e68af-1ce4-442b-8fda-b46bd6829f18 · outbound

This paper cites Qwen2 Technical Report.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Qwen2 Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.882278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.882278Z digest=sha256:5344af50cda80a61801fb23b607d92cf768779daf398a50f01ee643504212208

Observation 47b04974-d71d-41ad-a51f-c9a1a33f51dd · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.887051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.887051Z digest=sha256:6f15bd2c90d3903779b6c9ef50e0630ebdfdb861b9364c1ddecd9247913ba69a

Observation 4e1cef28-7bac-4abb-a63d-fa80d8a41e69 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.891943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.891943Z digest=sha256:434ed463cb08f5c4dae2c81c1d519420adf318be975ba7eb9751acedf93fa40d

Observation b89a4fc6-2a52-4f64-830b-90fc643e0569 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.897050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.897050Z digest=sha256:2ae3edcf616f2fa925165f3a1f8c05a998aab25cbebbd6050a402a3425dd66a6

Observation a245528d-1133-4eff-8ad8-e3739da4b8f6 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:40:40.987548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T15:40:39.901836Z digest=sha256:dd222698176c52ce56f43d2e68594c1de407295998cd96f4561efa8e70a26cd5

Observation 927131eb-c1e4-4242-935b-097e1159a221 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.907302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.907302Z digest=sha256:30aabd07fa7f2541f8a16d881aa9f59e6a4884bd0638db24db1fd5e4092f70f4

Observation 29d7a177-961d-447e-b8a0-95c4f6d97ac3 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.912031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.912031Z digest=sha256:618ce87c14018731417859e12c93936870d1f79e47837c4f84d455225a52b48f

Observation 05349184-bc8f-46eb-ab58-95304d74812e · outbound

This paper cites Cumulative Reasoning with Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Cumulative Reasoning with Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.916871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.916871Z digest=sha256:3c187f8e94a98720eb7c03594566b4bebbae5076c7526920c97028c0019bc52c

Observation 6c3ec3ec-f72d-4321-bbf7-c7c786e381cb · outbound

This paper cites Progressive-Hint Prompting Improves Reasoning in Large Language Models.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Progressive-Hint Prompting Improves Reasoning in Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.921948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.921948Z digest=sha256:fbc3c7b1d4099b3d5c01b23fba1cec5cd67011dc7f692be32969b44ee9ecd136

Observation 23df265d-f769-483c-a362-c3dfa64b58b4 · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.926795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.926795Z digest=sha256:ad8e042e8bebde6268315f1df5dadae29635adba3a5f3ddc196624749a1a5552

Observation ef799b5d-3e13-4643-8f81-5b823d7740c9 · outbound

This paper cites write newline.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models write newline

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.931716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.931716Z digest=sha256:c363f53878d17534de902d8185451db32990058472edb56c814c0944f16f0ea5

Observation 00c936af-350d-401d-9171-b91ecdd63caf · outbound

This paper cites @esa (Ref.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models @esa (Ref

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.937620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.937620Z digest=sha256:e19e96f1d5fa2c3b94061f44fea23cc3503c1fae422661702c99ad14d6bfcceb

Observation fb76f3c2-b2cc-4f7d-afa2-9552682e12ae · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.944207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.944207Z digest=sha256:12a8cce3c84ca45af4cf42c89329714365bfe9f6afac5ec4c5fed9391717d7e8

Observation b46613b4-3af1-405e-95c1-bd144fb6d148 · outbound

This paper cites an unresolved cited work.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.950474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.950474Z digest=sha256:11712683af912242ebce1caa74baf88e45a357c94d06435fba8a25bd3ab4a56f

Pith citing papers

Observation db9ebfaa-48ac-4194-918c-b6c7ac841939 · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.891540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.891540Z digest=sha256:b1031c8fa9de427404acc46e177007d4287f60f3e996d0ff02cf0ffb0aac61a8

Observation 2d14513f-748d-4736-b5d6-42ef87186ab8 · inbound

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification cites this paper.

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:59.476120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:44:59.476120Z digest=sha256:da13cd49bb2afc091c5e906b7097e276de02324c5acf2bb7604d247d2e112fdd

Observation 0a6fec3f-c75c-4d2a-aa25-1949f32e86ed · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:25.330394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:25.330394Z digest=sha256:3510c1cba679aa8a806bd48f72071d147bd4ba272c659b04bdea4a49ba09f6e2

Observation e7eed9a8-b5c5-4b7a-8a56-6dc6792e5537 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:25.672136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:095238a6810a2e0ff12efd234826fb11704fbd4f89d743bad05c6949bf340332