Pith. sign in

Paper Citation Record · LEDGER

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:2506.21655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21655 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:29:58.706228Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:06:38.215137Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:03.800004Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c96e234-31ec-458c-8f5d-8bb24e67d18e · outbound

This paper cites Qwen2.5-vl technical report, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Qwen2.5-vl technical report, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.781638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.888597Z digest=sha256:e92ffe6401f2ee60f209f6e176668bf56bb797882829bcf16a986b3f4112f485

Observation 327722d8-5644-4639-9b27-3bed28a57ba1 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.735122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.904495Z digest=sha256:54d7a23e27c36bcfa9cfe476aeab297cbe8e00bfeaed8f1e3ff5a3217b0e5a54

Observation d09779fb-616a-4f88-9d59-238be952b6ae · outbound

This paper cites Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.924914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.924914Z digest=sha256:04cac03a0e9cb9502441e5362cc8a7e57be19f0808343455aebdbda4b1ca9f8e

Observation e81456dd-7cec-4c91-9ef0-abfb61a3f24f · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.932417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.932417Z digest=sha256:97c2efcfbc7e4ded3f4ee3f5d5d9c3c88f5737e86f28fcf282d692591b89454d

Observation 016fa379-0fcf-4893-ad9f-fe665d6a571c · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.652681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.955233Z digest=sha256:8cddaaea054ca94d62d0844bd332cc84e2135a8ef405d40c08341be0b40fdb1d

Observation bbd749a9-6891-4608-86d9-ae3e8ea20915 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.974697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.974697Z digest=sha256:c519f0b65c46681cee2697d4806cc9b22f71334efd3da2b50093943b16c1bae0

Observation 5fdd68c9-9932-4818-b4ee-3ff99ee3e845 · outbound

This paper cites Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.988706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.988706Z digest=sha256:ce86a31c6c9da2083d1be5d8ecc4335d13967d17d60e90366ba585b4f5b2fe36

Observation e818cc55-f350-4991-8670-0a8a355a27a5 · outbound

This paper cites Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.614767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.994310Z digest=sha256:7fc8abe767d9086222e74e5d324b3b8d842d0513c5d99c5e95fa88c7d9966b28

Observation dbc04aea-902b-41ec-8568-3dd0883b44c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.009078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.009078Z digest=sha256:1eaff6b1ad1d63ce8700a1070f23b13c085ec2865f4b10fb774bca1dafeadc46

Observation 80de8bee-487e-4ee6-9186-82df041b8cbc · outbound

This paper cites OpenAI o1 System Card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.028456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.028456Z digest=sha256:b51b96b6e2fba7c06a26b1f9553c3508649b33d8075958b18b36913f7b9a63ab

Observation 3ec7c38b-6828-4b97-88fc-9fa899b7ad1f · outbound

This paper cites Figureqa: An annotated figure dataset for visual reasoning, 2018.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Figureqa: An annotated figure dataset for visual reasoning, 2018

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.544330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.044738Z digest=sha256:7b29e0e3ecff459397f905afaf837a42579e20ed9cfb688ae3a697b390eaff3e

Observation acf61fb4-74fe-426d-b3ee-ff087bf9ce08 · outbound

This paper cites Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.480854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.050376Z digest=sha256:e3cc6f7538db5b96df432256d57bc5e8b9b961b4f3b80fb7a95b441a9b23c49b

Observation 12dfd25f-c2a2-40f5-a76f-5eb6dfb29b1f · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.429329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.060256Z digest=sha256:20e54900326618f63fffc806f26e3bdc18fe02a60bb4e615a1bb685daa258a65

Observation 0f5e4e9e-3093-48d0-a5bb-64c12784dc1d · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.079295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.079295Z digest=sha256:b1b889df718244d7fb822f05bb36d06f6b03ab3149bcc49c24ea315ed1669cf8

Observation 56df06e3-2475-4ddf-ad5e-8574e58b8645 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.095636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.095636Z digest=sha256:496238694cb58d3877391448779bb4b0523b8c0d10f2798a6a5442125c5619cb

Observation 206b3455-9ff0-4cf2-bd4a-da9688c92399 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.113236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.113236Z digest=sha256:b051b94d3863296fc52fcd6f71418b5250e39e3a7a31bdb02ed422b8915a2417

Observation f3aac8ad-e753-4ebc-b15c-ffb64e941b4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.132413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.132413Z digest=sha256:ef1fce07275f499acb260c7cb23b49b3bfdfcf45c07f3e950d4c9cce22d753d4

Observation 03d23385-d3e8-4bdb-8bf6-8e49c2346561 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.361633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.149495Z digest=sha256:523c8a840cb56eda97ddb14287f83262271acd503a591371619048a1bba5dd85

Observation 71e362d7-4cd3-4a1f-a580-8f5c901f3d14 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.158314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.158314Z digest=sha256:147e248e53f44cd7dd3a11cf1b425b615dbbe3b0abfc50af6cc13c606dff3bb0

Observation f952b1c0-258f-48bb-aeb8-4b3a4d6ae0ca · outbound

This paper cites s1: Simple test-time scaling.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.186779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.186779Z digest=sha256:25d4f63da8b96d2e7836db0c76996b5ca91196b8420cb2eb49a19bbb2a1e1ca0

Observation 86fc2708-ff5b-4708-8c98-7bd8c78a9eb8 · outbound

This paper cites Gpt-4v(ision) system card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.211930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.211930Z digest=sha256:617fe0c075761020808d60c5aa98ecba1a55d0701c0ecbfb3247df606eb3ec2d

Observation dc33ca82-101c-41c5-aaa3-3e8c96efbe77 · outbound

This paper cites Gpt-4o system card, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4o system card, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.226495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.226495Z digest=sha256:d095293056e2038d607432d6eef7f03abda63f90d204686c81a35ac4b52286c4

Observation 396e1984-423c-4197-943a-369eb298496f · outbound

This paper cites Training language models to follow instructions with human feedback.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.258630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.258630Z digest=sha256:6837708602c5e9eb5a49dcaf27738e057d44648316215da97e3d583b3d1a119c

Observation aa884deb-08a2-42ee-8cf7-3c4eb220165b · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.285695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.285695Z digest=sha256:42f0a5184f93d9d91f6389406268039dc37382ff0fb2c6580a6be7e460fad72a

Observation d3222508-3b10-4c05-9fcb-1cb1ca70d0b6 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.314845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.314845Z digest=sha256:55ef9a601731f569aee2553a73a6b84ab681fb3557ddc7d5915c180e240e3a9f

Observation 9ebabd7b-508f-4a76-a38e-41f1264898a9 · outbound

This paper cites We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.270019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.332879Z digest=sha256:3593be3b7f4b4166eb4de2f4e21a538e5444e0bd8f432c91db0e0fea4d402c4b

Observation 39c9e3c3-3bf3-4f1f-a155-c8be6a65e26a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.350191Z digest=sha256:0f8c6ee2bdfe8b45341fb4e4006ea312eb286c1b72cd3efcb1a17a23a441317b

Observation 28c8ddea-9165-4496-a0d0-ea487f9a260a · outbound

This paper cites Proximal Policy Optimization Algorithms.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.369712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.369712Z digest=sha256:f386fa018057f57ee31e977a5bd68667b9639457f61bc96dbb2a7ce38abfd264

Observation 09744d3f-4f29-4648-8b1d-fa749b4b2543 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.391054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.391054Z digest=sha256:30d032621cd988f2c0a78f30d37e4b1d070130d3b09186cc490e696d9a7f8eba

Observation 3fb3d1a6-01cd-4bbb-aa48-a8eb1b7bc719 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.205422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.405993Z digest=sha256:7f9bb9d6fadbae3005bf31b321540300f2590bd6e979608b460967e48e2d8684

Observation df3d18f0-709c-4db1-a285-a23af471f3f5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.422267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.422267Z digest=sha256:ea174bc63d7994cdfaa42e4dbadbf24dd4ee770e8053e41c4a434d4be41ba391

Observation 6696a32d-4608-4ef0-8de1-58e824fae9d1 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.444374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.444374Z digest=sha256:378e9b139a6e1dfc95f732338276a56b898576a8b3cf1d756dcd872fccfb6d3b

Observation fe87a75a-7355-42d1-8540-f8e5d4d0c008 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.463737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.463737Z digest=sha256:ea54f03f8e989af596b23ceb76400c970f9aada02b11372c161b59510192d4ac

Observation ee98e613-7690-4797-9504-94d2f83e0d75 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.485026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.485026Z digest=sha256:3745ee5826a71eb02c3631dcfa672bce35a01d73d8d66bf14ca245c585ec3cba

Observation f7352591-2e0a-4fc6-99aa-2ddd81b2977a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.505793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.505793Z digest=sha256:93697d50e16eb31234980e5893239d6dad4eb0ffefdcad1c063d3e5ca6d02f86

Observation b32addc3-6660-4748-8b35-0f3b99ab08f1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.514846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.514846Z digest=sha256:7ba4180258aa76d938e5833b0aa9b0e686d54ced6addefaf93321ede76a76fa7

Observation f559c9d4-cc7e-4725-96c6-91dcc2de487d · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.645747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.645747Z digest=sha256:e111a92daee5dd54dc75efadc2106920f2ab070f8de1a663ba5aa6d26e4d0a4e

Observation 4f21f1cc-1561-4ae4-b7f2-662c6d7cd173 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.670138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.670138Z digest=sha256:18ff62917dcb554b9b8282d6bde1f59b260f23ed46e89fb1062dea0a4843dbd2

Observation b9d0598b-3945-4573-bf7d-2516c8da6e38 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.699089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.699089Z digest=sha256:63d2284041a4f28d27698b23a4d105a9ac50bf9733929cc2edde4227f7742a84

Observation 9cb4bd50-cb27-4851-be18-111a1b7de6e1 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.144612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:58.706228Z digest=sha256:1560a4a9c9a7ebc6cb81b7c86b438d15e6fa36d06b1509ed367dd92482d4da77

Pith citing papers

Observation 42811ffa-b70c-4621-bfea-976bd4e8bfe9 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.389824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:97eb8317f8727824c6aedd8af6c878707c3bb9442dce70fd3bf1b837fd045c47

Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.261864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:8f2cfb6091dc31a2528fed9e286f13534bfe73676ff209ad8b055cffe56c0d3d

Observation a858a48e-99e8-4637-bbd6-e9b8dc097144 · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:29.016519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:2c2f4e98ab9b050402a6dc091e56c9081cac664d3dc2f42cbd9a08924a94d209

Observation 16209c2b-dcd3-46ce-befd-a980d05c49f3 · inbound

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning cites this paper.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:03.801892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T15:06:38.215137Z digest=sha256:7ba765e8fc7bbe699d2c26a071e491cb498ea7bbb2460442ab2b3997496561eb