Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:55:42.622845Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.14683.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:55:42.622845Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:56:39.504430Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:56:13.510220Z
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f41dd2f5-c890-404e-8a61-91943603b5a6 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4c0288-0f85-4589-9c67-3dd807430791 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization OpenThoughts: Data Recipes for Reasoning Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3323e53-83b4-4e52-a678-a60a1bac7677 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def4558f-e3ab-4f03-9a65-2dc11a981de7 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 851a1dcd-ee08-4c1c-bd10-434df931e55c · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b5c68a-747b-40cc-a7fe-7e7cbf65108b · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe408ce1-a139-472b-86a2-cd9a40d73c78 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Let's Verify Step by Step
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b541b4f-9a42-489e-8a94-581485138278 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55adf272-4e07-4ef0-975f-904dc0245558 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Magistral
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26475905-cb27-4011-9119-0f0a525756e2 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Training language models to follow instructions with human feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed715db-3950-435a-ba11-91f47d4e0237 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41314aea-683b-4c79-9ca2-95e98c39808c · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Gemma: Open Models Based on Gemini Research and Technology
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00e935a-2c58-46b2-ab08-e8702f93d437 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea26196-b430-4ef0-bc43-bcd375bec63c · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Measuring short-form factuality in large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45047020-db86-4237-9b92-81b3f553d448 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef95cc2-b12f-4a6b-8a84-7b3cb7a23793 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7917bd7d-861b-4f80-9ded-b4b348468bba · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization HARP: A challenging human-annotated math reasoning benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0815638b-df74-481c-8fa8-bdd6e337cb8d · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea541bba-8763-401e-8d46-369a528fb236 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959f1a18-0010-458e-a947-05e298a39acd · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a3ce8a-cfa6-48b5-8796-7fc38837fe9e · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e8d495-3b3e-4664-b5a9-ae421033963b · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c79cd03-d79e-4acf-a928-ee92ef571748 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Proximal Policy Optimization Algorithms
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe026e6-8a8d-4273-9404-f310a2f38811 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c82986-bede-402c-81e8-de0a59f8ff2d · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization The Llama 3 Herd of Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ca1b6c-54fa-4e40-a847-b45d0f7aee59 · outbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Reasoning Does Not Necessarily Improve Role-Playing Ability
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb70732-447f-4e2c-aab2-faa64c78da79 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Reference 287
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9285862f-99e2-4b13-bd67-3bfdfc80eda2 · inbound
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c00b2957-89d4-47bb-b0ee-42962a5b5236 · inbound
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05ef2408-2264-4ead-839b-6940bffe0e8f · inbound
Trust Region On-Policy Distillation MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Reference 288
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.