Pith. sign in

Paper Citation Record · LEDGER

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2505.12432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12432 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:46.508251Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:14.896841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T19:21:48.411112Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a505e807-a861-4c0b-98d9-8e1720f2b6fc · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:45.857977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:45.857977Z digest=sha256:45583335f2c0b36a9575475db1f710662fd4af94851f25200b9611e2af199114

Observation 1da66c8c-df16-4f25-9390-592bbbca1659 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:45.863816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:45.863816Z digest=sha256:a00ee08a3ff29bc8fb4c7e8d0aef44fb0df62ee87e8092fbf4319f3f57634bab

Observation 78eb0315-3194-4bc4-a610-f33084568247 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.003693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.003693Z digest=sha256:08218eba28ad9eb483e34992c992eef26cf4563a89c311a1d2381d7f32b411f3

Observation 01ce5095-d6a1-4c50-abb3-1c032426e652 · outbound

This paper cites OpenAI o1 System Card.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning OpenAI o1 System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.053819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.053819Z digest=sha256:19a330d573d5e1d9a8c7a54b9e4336b4c472323e0a66f9363ea75eb45065fc4b

Observation 3de2749a-c9ef-48ca-8a11-c16cdaf572ac · outbound

This paper cites SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.058897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.058897Z digest=sha256:58b456c9b734f24775702c1f25fc8d9aabe1000b940e1b878ff950177e84dfdd

Observation fd9a489e-3d24-478e-8cbe-8ca7ff4f0224 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.063233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.063233Z digest=sha256:e2cc985b3a5848d5f41b0f6bcb497dd25dcb203a312cc7fc42e55fbf0546bb4b

Observation d3a6211f-9b00-45ea-92a2-ad3aadac7c63 · outbound

This paper cites s1: Simple test-time scaling.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning s1: Simple test-time scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.067524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.067524Z digest=sha256:603eead53ddc37f76ec5a3700742fb52f39fb71a45ebfff8e512c5b2a8ee3ef7

Observation 735177b5-6545-4778-bf20-2bc984e92764 · outbound

This paper cites GPT-4o System Card.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.071775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.071775Z digest=sha256:96af26c5cf8d18e4f69f28e7ccb2ce23d50aa8f6113d1442ec72ad8e50076d35

Observation 24df8388-fb31-43b5-8374-2447c667c5d1 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.169738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.169738Z digest=sha256:1b05b09af7255875cfae3565e0cfea4153179bdd691ecaae056782223605494f

Observation 673daf41-b42f-45df-bc96-62a88921f2fc · outbound

This paper cites Dast: Difficulty-adaptive slow-thinking for large reasoning models.arXiv preprint arXiv:2503.04472,.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Dast: Difficulty-adaptive slow-thinking for large reasoning models.arXiv preprint arXiv:2503.04472,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.260365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.260365Z digest=sha256:6e1c2723dc99242af5c2360bb768ffed9a2a0eb9954ac9c69c257c7e78d47bff

Observation 6b05100e-6e67-42f7-bccf-94e02c4e7266 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.265254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.265254Z digest=sha256:94244a6ce3f33074752094ba40c2d70a011ba0f2dc95923b9745fc7e9fe18534

Observation e115d9fa-dfea-44cb-baf5-0a547975168d · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.270585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.270585Z digest=sha256:af0e606307403b3fb6d7f0ad96b65464fd3f3af453d6ad7a1622b05a624169a4

Observation 7f599363-13ac-4945-8cbc-1c801bfcf3c3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.274359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.274359Z digest=sha256:74af320a978c611fcda513e0619e25f02918af24b88af5fa1f32c55343b5f968

Observation a6ed0821-378e-4180-a4dd-925e76b47924 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.329461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.329461Z digest=sha256:edba307bdfce18baf67b2ddf2cceddb9dc059371f3079f55e6f707a6be7be449

Observation db346497-f6f2-465a-b9db-98e7eb84f9c7 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.333604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.333604Z digest=sha256:37be0af41a155a0d1a2fda258f03ceec2063f9aa7ba412e9cc7c1aac9c4eaef2

Observation 1ba20f66-acb9-4b26-9a8d-c9b26048fe75 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.337631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.337631Z digest=sha256:c1ba2a012d4a5598c082dedd79a494f57e93cbaede84475710d43a96a674e104

Observation 51e63674-0375-40a4-ac0e-fe2c5a6716fe · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.342233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.342233Z digest=sha256:1c235aa4d385c298b48cca37bed7ae52e28785fb1d04298cf6e2ede8533ef5fb

Observation ba12231c-599e-4856-8da1-a59661153281 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.438147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.438147Z digest=sha256:a38158742efb97348e4b6aa89a09785d6fd2415d591c4bd29930b12e9624dd26

Observation d8ece693-0bbd-448a-8bda-e65ecbde9a7d · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.508251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.508251Z digest=sha256:cac3d4f0297b6933e1292588ccdf24d5902c30ec2ff2c4b16fd70c2ccd0edc1d

Observation 6514ab55-bde0-4c99-8342-b8488a7798b1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.252382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.252382Z digest=sha256:e04032b55e1f2315edfb496de42f496e8a952e3956428010e937dd0767be2c10

Observation 1b8bf07e-fdf7-4fd9-84f2-e45c3aaa964f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.256869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.256869Z digest=sha256:1af02be569fe9b4500b7b13c41d104dd16ecbd4cc6cd16037bbb9cb862d18f17

Observation 12cb828b-d076-4d14-a2f5-f8e6f59452d9 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:46.324995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:46.324995Z digest=sha256:93a27d30aa618fb6c5c3e9cfc785723d4bbab308c337ddb8138086ce8e6ce947

Observation 5b74eb5e-bf75-46a1-9f50-7334544e6f31 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:45.868304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:45.868304Z digest=sha256:d7f49039d8c62c711041584d7f88fdcf036777c29c07984c73521abc881f51a5

Observation 9ed0cbac-6ab7-47ad-802e-2f06fb49c99c · outbound

This paper cites Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:45.873197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:45.873197Z digest=sha256:8f5e46626b93ba0dbf50637df7e73fc16a5a29351d42627abc5dcd8fe8da97bf

Observation b59c9d22-f03f-4e1d-bcb0-92607a77b814 · outbound

This paper cites Qwen2.5-VL Technical Report.

Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:45.735671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:45.735671Z digest=sha256:90aa90ed42ae936e1e0c65231a6145f5adbdecb9766dc071f18708410e912e61

Pith citing papers

Observation ccc0b7ab-2e2b-4db4-ad07-1df3b15e635b · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.896841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.896841Z digest=sha256:01baa8eb6d25edbd63a99644b8019327ed10d2022fdcca757aaa3da2b32b83c0

Observation 5fdd68c9-9932-4818-b4ee-3ff99ee3e845 · inbound

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization cites this paper.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.988706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.988706Z digest=sha256:6e2f05aab7fce9c490dbaa1e8a21ed492c6c1fe41427577ac64ec3fb5b3eac22

Observation 21d7b7ee-9adf-4259-bc08-f4f631342e59 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 241

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.413537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:787c45e2a9c84134dea7c28c6f353e5a7c039197349759800632021254bc953a

Observation 083c6a0b-ebfc-4fbd-8cd4-7ace85e552bb · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:13.888889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:13.888889Z digest=sha256:ec00c8da53b0b3818ccbb2d2a97a85b04bf89037b45bcc5a6434e198099d0f30

Observation 77d9b4ec-6780-4bfe-96a7-700501be8b6b · inbound

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning cites this paper.

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:40:13.534500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:40:13.534500Z digest=sha256:12d5c697f955748d5136cfe8ce31ceb0d13a09f2ad617b7d04cbfcd3858a1a9c