Pith. sign in

Paper Citation Record · LEDGER

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.14683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14683 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:55:42.622845Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:56:39.504430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.510220Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f41dd2f5-c890-404e-8a61-91943603b5a6 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.505183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.505183Z digest=sha256:2a87aa6ce7668ca4c6da15d39a1c659a44c66b76fc086dcdbc661d5f313c8b24

Observation bf4c0288-0f85-4589-9c67-3dd807430791 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization OpenThoughts: Data Recipes for Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.523415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.523415Z digest=sha256:4d275dfa3df01c20610bf0ab42ca07d496dd4c9aaeb4d5844224102420cddeb2

Observation e3323e53-83b4-4e52-a678-a60a1bac7677 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.528683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.528683Z digest=sha256:7cf341cc5ad530de944e1f2c36f8ef20919a5edba6c3c6774d0ba16b99c497c2

Observation def4558f-e3ab-4f03-9a65-2dc11a981de7 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.534276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.534276Z digest=sha256:94aa009e6115cb63e61d4506140862ae407fc7cc184e5a30643d0cb88c80978d

Observation 851a1dcd-ee08-4c1c-bd10-434df931e55c · outbound

This paper cites AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.539416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.539416Z digest=sha256:137937160403a88befdbefcc32053f295cee0a04976b05172e992d4ef92a6093

Observation 95b5c68a-747b-40cc-a7fe-7e7cbf65108b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.546685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.546685Z digest=sha256:6c26265e776166b6b700a0bfe8d527ede66aab0fe0e417e98100a4e839d8f59e

Observation fe408ce1-a139-472b-86a2-cd9a40d73c78 · outbound

This paper cites Let's Verify Step by Step.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Let's Verify Step by Step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.551665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.551665Z digest=sha256:9d0eae7c1dea979a0044d73ef9707b22958ce5e2bab1a1a882369394442e48c0

Observation 9b541b4f-9a42-489e-8a94-581485138278 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.556087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.556087Z digest=sha256:05f523a8693a5a15d7341d70ec6d66e980e5667f9c19f26c55f64f409253fa39

Observation 55adf272-4e07-4ef0-975f-904dc0245558 · outbound

This paper cites Magistral.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Magistral

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.560228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.560228Z digest=sha256:01a1b6002b179e73e5921e7ade9a27fae1bc58d784687d53db53660c890d331d

Observation 26475905-cb27-4011-9119-0f0a525756e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.564588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.564588Z digest=sha256:c46dc898f601159f02aa678857c359e54befacc8d06beeb55135433d2c102f33

Observation bed715db-3950-435a-ba11-91f47d4e0237 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.572359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.572359Z digest=sha256:cd448e40f4a0d3a5abd42578fe5baea5ea131b8e082de07eaf6ec35317b76a0d

Observation 41314aea-683b-4c79-9ca2-95e98c39808c · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Gemma: Open Models Based on Gemini Research and Technology

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.577131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.577131Z digest=sha256:f00ca356fe9c0865ae30efe48233635bedce2507a39340b70db3291f70749dcb

Observation b00e935a-2c58-46b2-ab08-e8702f93d437 · outbound

This paper cites DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.580195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.580195Z digest=sha256:359a66abdcceafc6ab7cfc6136ce3985855e3ef566747a66d94739779a59ec3d

Observation aea26196-b430-4ef0-bc43-bcd375bec63c · outbound

This paper cites Measuring short-form factuality in large language models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Measuring short-form factuality in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.587682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.587682Z digest=sha256:05d5bb72b38630075ff31d0cd7633619ea204780ac764345aea3fe7ccc9e9bc5

Observation 45047020-db86-4237-9b92-81b3f553d448 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.592502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.592502Z digest=sha256:f9277858f6958e6a06e6bdb24d3aa52042ac7f40776434f26f99fc77a9b33481

Observation 0ef95cc2-b12f-4a6b-8a84-7b3cb7a23793 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.595925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.595925Z digest=sha256:90e1560faadff3e907bd2f7f9cc84ad54922a1f793f93381e6238fcd179a7ea0

Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.604562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.604562Z digest=sha256:3cfb580566f88e22857a1db773f82f19482911de3f6bc438ef348b49532ced12

Observation 7917bd7d-861b-4f80-9ded-b4b348468bba · outbound

This paper cites HARP: A challenging human-annotated math reasoning benchmark.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization HARP: A challenging human-annotated math reasoning benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.608109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.608109Z digest=sha256:2b17a7da5b8d09b7d812838b601092a01030d9035f77ae384711134c90b4a6b6

Observation 0815638b-df74-481c-8fa8-bdd6e337cb8d · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.611272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.611272Z digest=sha256:b29e26cf1cba7377f0c142c01bc05799596d08401b1d787646a09e51077d2373

Observation ea541bba-8763-401e-8d46-369a528fb236 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.614460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.614460Z digest=sha256:53f981a328a452459e7a8cfb0503c47e05ea4c891f59195e5ad33ba170765256

Observation 959f1a18-0010-458e-a947-05e298a39acd · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.618381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.618381Z digest=sha256:fbe185e72c78206540afe1ac6e5e4f3e535f35606419c50df888fa97ec678487

Observation a1a3ce8a-cfa6-48b5-8796-7fc38837fe9e · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.622845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.622845Z digest=sha256:d8fdb2df9846481d59ecba36d4f97ef9d76ade207c6581873b95dc4445dad704

Observation b0e8d495-3b3e-4664-b5a9-ae421033963b · outbound

This paper cites PlanGenLLMs: A Modern Survey of LLM Planning Capabilities.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.583833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.583833Z digest=sha256:3cbd9f3a5cf9d179543cb4e68b45b5df68e2d64018c17120634b0a579708746b

Observation 0c79cd03-d79e-4acf-a928-ee92ef571748 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Proximal Policy Optimization Algorithms

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.568123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.568123Z digest=sha256:24cb8a22da4c536ef6e3ddca4bdf7889723ea0971ad92acb9aa24a58287180fe

Observation 5fe026e6-8a8d-4273-9404-f310a2f38811 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.600378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.600378Z digest=sha256:7b79521cba9d128daa7f51e94c4c4f2b5dafe05e3ab5dde2ac1cb81db5430fc6

Observation 95c82986-bede-402c-81e8-de0a59f8ff2d · outbound

This paper cites The Llama 3 Herd of Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.512830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.512830Z digest=sha256:d283206772cc893f44ff529a40fdb01ca245a51c5e82493bb5daa20a34c7402e

Observation b5ca1b6c-54fa-4e40-a847-b45d0f7aee59 · outbound

This paper cites Reasoning Does Not Necessarily Improve Role-Playing Ability.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Reasoning Does Not Necessarily Improve Role-Playing Ability

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.518689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.518689Z digest=sha256:bfd7c70ed0c40f0e38c6e17285fea51948dd202ace63f35513c4a16988c3769a

Pith citing papers

Observation 5cb70732-447f-4e2c-aab2-faa64c78da79 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.767561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:1e3e577b23eb6f8f3cb04acca762c41d3d497f6bc0e83d46c13ee3896ade1395

Observation 9285862f-99e2-4b13-bd67-3bfdfc80eda2 · inbound

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs cites this paper.

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:06:00.557948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:14:38.371700Z digest=sha256:eb2a7d6c0d84afb6f98e3e849924965e223995d6232dab970220f9ff54ef725e

Observation c00b2957-89d4-47bb-b0ee-42962a5b5236 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.540542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:71e35352e1e8c13a7af7789cc254b04b1f19ea4fc83e0a15c47ba501d659a8a9

Observation 05ef2408-2264-4ead-839b-6940bffe0e8f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.511537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:f8c01c83e3347c2b4515ed89834f4edaa92892d116d6afcbcc7b27ca80a60a28