Pith. sign in

Paper Citation Record · LEDGER

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.14683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14683 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:55:42.622845Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:56:39.504430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.510220Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f41dd2f5-c890-404e-8a61-91943603b5a6 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.505183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.505183Z digest=sha256:3550277d92599291d440ad1bc4de3d1df737350e74704e9ddb0db3bf451db8c7

Observation bf4c0288-0f85-4589-9c67-3dd807430791 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization OpenThoughts: Data Recipes for Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.523415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.523415Z digest=sha256:07c119fcb50d3ddfed43f53ff96bb721be161fa8041e18b06de741f8d8ed4860

Observation e3323e53-83b4-4e52-a678-a60a1bac7677 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.528683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.528683Z digest=sha256:c3a14673b3f8a2e1a168d816fe7763fb4448a9c9bdf08e27c744c1b380efb5c6

Observation def4558f-e3ab-4f03-9a65-2dc11a981de7 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.534276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.534276Z digest=sha256:c0eca9850b27771772000cc06d7132174b9e6e9901fbb33e95c7d41991af05d3

Observation 851a1dcd-ee08-4c1c-bd10-434df931e55c · outbound

This paper cites AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.539416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.539416Z digest=sha256:c35ed4e8809b347fbb9489c4abaa85b09b89076dee013f6179f76f6086861795

Observation 95b5c68a-747b-40cc-a7fe-7e7cbf65108b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.546685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.546685Z digest=sha256:099adb04f68fd78fdabde044031a8ddf85e42670700a5795a174de33a207ced9

Observation fe408ce1-a139-472b-86a2-cd9a40d73c78 · outbound

This paper cites Let's Verify Step by Step.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Let's Verify Step by Step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.551665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.551665Z digest=sha256:306b3a17fa544c3664a963fadcc3b18599b40075f1d5d3891f3e0522744ac04e

Observation 9b541b4f-9a42-489e-8a94-581485138278 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.556087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.556087Z digest=sha256:7d26d830ba756bd9d9927f3e6fc54c8e7bf4644fb8de3ef42b88554202262641

Observation 55adf272-4e07-4ef0-975f-904dc0245558 · outbound

This paper cites Magistral.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Magistral

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.560228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.560228Z digest=sha256:b3681dc08980b9249b8c221820d921b768eafe53fbb47aeb109eb8a88ae263f9

Observation 26475905-cb27-4011-9119-0f0a525756e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.564588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.564588Z digest=sha256:7a71542a109047d76f4f90eb8b4362b2a3ccbe0312d7a2300b36a1e365309185

Observation bed715db-3950-435a-ba11-91f47d4e0237 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.572359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.572359Z digest=sha256:789c178911b3177b3069d223e5ad72664e068fb802a4aa58b4ea077ff67137d2

Observation 41314aea-683b-4c79-9ca2-95e98c39808c · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Gemma: Open Models Based on Gemini Research and Technology

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.577131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.577131Z digest=sha256:e26588d8df93bcfae6ff560eb277d53b7b85e9653cb8d27b96915b1e01bbfce1

Observation b00e935a-2c58-46b2-ab08-e8702f93d437 · outbound

This paper cites DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.580195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.580195Z digest=sha256:fb14ae61685c1f1eb230ef89376327135f9ef368299b54af188192a00a4dcef0

Observation aea26196-b430-4ef0-bc43-bcd375bec63c · outbound

This paper cites Measuring short-form factuality in large language models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Measuring short-form factuality in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.587682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.587682Z digest=sha256:1a0eac719059e8b4ede9ec3dea61e745747aa144dc96b0913dc9c61e0c9643b7

Observation 45047020-db86-4237-9b92-81b3f553d448 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.592502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.592502Z digest=sha256:3118dff6b0b79be64d06dce0c281774387e6949e67645d50e866870d3b326fef

Observation 0ef95cc2-b12f-4a6b-8a84-7b3cb7a23793 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.595925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.595925Z digest=sha256:b04c4bb54c4b506c80ec89cdb679c68fba586d18c7907c9c481b7e32db729f33

Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.604562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.604562Z digest=sha256:31ebdc74cbd596bbfc0cf3d0a96a4dbc31eafa6735715c0f318707b1fa9b9462

Observation 7917bd7d-861b-4f80-9ded-b4b348468bba · outbound

This paper cites HARP: A challenging human-annotated math reasoning benchmark.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization HARP: A challenging human-annotated math reasoning benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.608109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.608109Z digest=sha256:170d70fdbee3d85a2f91ce7d9b1cd8b1e1a8d4e6f2471e3099fd4679dedc3765

Observation 0815638b-df74-481c-8fa8-bdd6e337cb8d · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.611272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.611272Z digest=sha256:cbc448eb62d0a61cdf29cb7485d7d4c01a994218d14eb8538bb76aa13b71d3be

Observation ea541bba-8763-401e-8d46-369a528fb236 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.614460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.614460Z digest=sha256:9d8d0672b4f24e7b5ef76ee08cc3a2f17abc8e02f33a86838769908bb9b62de3

Observation 959f1a18-0010-458e-a947-05e298a39acd · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.618381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.618381Z digest=sha256:5030fcdc66695362df8e4bdd4359774d4a5f200b8d83d6a90723f0c7e4d53dc5

Observation a1a3ce8a-cfa6-48b5-8796-7fc38837fe9e · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.622845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.622845Z digest=sha256:6883d8564ab18e2f19095c16866a5378fd7de368fb577ed790d48c8f441d471f

Observation b0e8d495-3b3e-4664-b5a9-ae421033963b · outbound

This paper cites PlanGenLLMs: A Modern Survey of LLM Planning Capabilities.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.583833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.583833Z digest=sha256:4d481971a7c83808300f5281203d469624a077d23c1e3902fffb8657bfa0db3a

Observation 0c79cd03-d79e-4acf-a928-ee92ef571748 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Proximal Policy Optimization Algorithms

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.568123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.568123Z digest=sha256:b4adc836d7fb28334c267c8fcefb6cc43219f2974ff240f1852044ccd7a5dc7a

Observation 5fe026e6-8a8d-4273-9404-f310a2f38811 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.600378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.600378Z digest=sha256:43dbdad582f2d33aeb776321b640743af55a6665ce40772c9c04714189ab52ae

Observation 95c82986-bede-402c-81e8-de0a59f8ff2d · outbound

This paper cites The Llama 3 Herd of Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.512830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.512830Z digest=sha256:08b4346fdd32866c791986c35f08afc67ee540339d98c3a80430739efcbf18d2

Observation b5ca1b6c-54fa-4e40-a847-b45d0f7aee59 · outbound

This paper cites Reasoning Does Not Necessarily Improve Role-Playing Ability.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Reasoning Does Not Necessarily Improve Role-Playing Ability

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.518689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.518689Z digest=sha256:c43cf28d287b4505a9801222797db3c7024d0ae59cbffe96d8fcd831266c7fa1

Pith citing papers

Observation 5cb70732-447f-4e2c-aab2-faa64c78da79 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.767561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:1ab694ff4acb3ac72bf18e7e27120fa3d58654bcd98d28f1b2f40171171f6a8b

Observation 9285862f-99e2-4b13-bd67-3bfdfc80eda2 · inbound

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs cites this paper.

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:06:00.557948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:14:38.371700Z digest=sha256:835ad91a12bdae0f92b502508b21cb8e9504ead2307e6c988f8b828c65d55231

Observation c00b2957-89d4-47bb-b0ee-42962a5b5236 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.540542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:71b1f2a60d123417759c3e57c4ef90858c225ee06831c1d80e376dc62d7d503a

Observation 05ef2408-2264-4ead-839b-6940bffe0e8f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.511537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a6eff9a4f300a16e040de5a0275540346c0a9a064a8cd58e8293b4b69b1ef146