Pith. sign in

Paper Citation Record · LEDGER

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:2506.21655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21655 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:29:58.706228Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:06:38.215137Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:03.800004Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c96e234-31ec-458c-8f5d-8bb24e67d18e · outbound

This paper cites Qwen2.5-vl technical report, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Qwen2.5-vl technical report, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.781638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:27.888597Z digest=sha256:55385d0ad1ab48022224a05a96df0b57b3bccd997f7c307330e03458047ed2d5

Observation 327722d8-5644-4639-9b27-3bed28a57ba1 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.735122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:27.904495Z digest=sha256:1d76f47df8f09335c807fd1ed30e1cb250533b8b931f1e880edc4b65d55b13d4

Observation d09779fb-616a-4f88-9d59-238be952b6ae · outbound

This paper cites Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.924914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.924914Z digest=sha256:a2839c63b9fe3f3bc1da92eb83ec769d64230b6ec38077571218fea41f9ba73a

Observation e81456dd-7cec-4c91-9ef0-abfb61a3f24f · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.932417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.932417Z digest=sha256:774d3c4e6b43ba02f0327d6650c63885abfb00e97c9cf3aa10d85cdc35bec491

Observation 016fa379-0fcf-4893-ad9f-fe665d6a571c · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.652681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:27.955233Z digest=sha256:514b8fee05a4e0e0d068982d7526412e1e87d3587cde125d7c326f4f6d88855e

Observation bbd749a9-6891-4608-86d9-ae3e8ea20915 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.974697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.974697Z digest=sha256:06b6d81a5d49dfbb995a067271cc3a7656aced41285b35764ba165cd945710b8

Observation 5fdd68c9-9932-4818-b4ee-3ff99ee3e845 · outbound

This paper cites Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.988706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.988706Z digest=sha256:6e2f05aab7fce9c490dbaa1e8a21ed492c6c1fe41427577ac64ec3fb5b3eac22

Observation e818cc55-f350-4991-8670-0a8a355a27a5 · outbound

This paper cites Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.614767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:27.994310Z digest=sha256:3a21e3c4a7a7c5323e53bfb3909f5399523820d1a04811d69fff2df7840e4107

Observation dbc04aea-902b-41ec-8568-3dd0883b44c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.009078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.009078Z digest=sha256:9fcdb9daa7ce6e6034163c51b6355c3fc6a6c804118fb8d067f33a8998cf1967

Observation 80de8bee-487e-4ee6-9186-82df041b8cbc · outbound

This paper cites OpenAI o1 System Card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.028456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.028456Z digest=sha256:ff6156dc1fa77ddb8f66e5690b91d574ae9aeed09fa60b9e9a11af8cacb0f55e

Observation 3ec7c38b-6828-4b97-88fc-9fa899b7ad1f · outbound

This paper cites Figureqa: An annotated figure dataset for visual reasoning, 2018.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Figureqa: An annotated figure dataset for visual reasoning, 2018

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.544330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:28.044738Z digest=sha256:d94adf3bfb0f21a584f275fd85c3c3a55705de9b53d127915e4324d03e508e22

Observation acf61fb4-74fe-426d-b3ee-ff087bf9ce08 · outbound

This paper cites Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.480854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:28.050376Z digest=sha256:7a8ad793667aecab48a50f93c4dad256ce54e3acf83ed8eb8abcf82411722dde

Observation 12dfd25f-c2a2-40f5-a76f-5eb6dfb29b1f · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.429329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:28.060256Z digest=sha256:ddccf6ea1d085e0bc9827d16c37c2971f97b3973810775dfa9fd6d3177116bf1

Observation 0f5e4e9e-3093-48d0-a5bb-64c12784dc1d · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.079295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.079295Z digest=sha256:c8f89615b2c7333d1d70a1ee93af24eb526563858fd3c9e25aae731093e0def4

Observation 56df06e3-2475-4ddf-ad5e-8574e58b8645 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.095636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.095636Z digest=sha256:552336b15cdbfc2cd129f8353b6c04c8f843ec2b62c087d69a75674efb326f46

Observation 206b3455-9ff0-4cf2-bd4a-da9688c92399 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.113236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.113236Z digest=sha256:8028852053738a5cf28852c91d1fbd900ae4d2fe0aa9aa28ffed98b95cb5b2c7

Observation f3aac8ad-e753-4ebc-b15c-ffb64e941b4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.132413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.132413Z digest=sha256:143928c294f2c5ebf10d8a9b904fb5a2ebb1de956e3924139a1060ee7228818d

Observation 03d23385-d3e8-4bdb-8bf6-8e49c2346561 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.361633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:28.149495Z digest=sha256:1629acb8c1857800c2199094035358f07007c5a1a1acb3f36a5977fce7c08fc3

Observation 71e362d7-4cd3-4a1f-a580-8f5c901f3d14 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.158314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.158314Z digest=sha256:0d614344ff09cb68727832ab749a48ae131bddcae5d7aa3b1719f584b09d76dd

Observation f952b1c0-258f-48bb-aeb8-4b3a4d6ae0ca · outbound

This paper cites s1: Simple test-time scaling.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.186779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.186779Z digest=sha256:2dd47eabfb33455e9dd1bc1580fef63aad0c6048c5f7e3f828a2cfa49651412b

Observation 86fc2708-ff5b-4708-8c98-7bd8c78a9eb8 · outbound

This paper cites Gpt-4v(ision) system card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.211930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.211930Z digest=sha256:e46af38b96c7c89adc0bc5785a4a9166fb5eb9032a28e6e80d215fdfa14a283e

Observation dc33ca82-101c-41c5-aaa3-3e8c96efbe77 · outbound

This paper cites Gpt-4o system card, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4o system card, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.226495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.226495Z digest=sha256:ba850eeb567f4cae43f307bddc8021f2958358ca0ba4fc370a6f1f103a4a840d

Observation 396e1984-423c-4197-943a-369eb298496f · outbound

This paper cites Training language models to follow instructions with human feedback.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.258630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.258630Z digest=sha256:74a7bb3fb7e4fa2fb616e9c73b4e5c877a6d9fca2c9372a6ff138e9ddec675de

Observation aa884deb-08a2-42ee-8cf7-3c4eb220165b · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.285695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.285695Z digest=sha256:5abc8b7d7f75a7d794b4d1c329b42fef61dd2c8fc74d92c5dac4c51f9ae805fa

Observation d3222508-3b10-4c05-9fcb-1cb1ca70d0b6 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.314845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.314845Z digest=sha256:83850fac07fefbd364e8b851b0c9c0d95642c01097db8a7cb3f52a1ddadd111d

Observation 9ebabd7b-508f-4a76-a38e-41f1264898a9 · outbound

This paper cites We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.270019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:28.332879Z digest=sha256:6128dc1649d7bac27a933828fb2df40823fd6b752d76bbae61e9e0201794d284

Observation 39c9e3c3-3bf3-4f1f-a155-c8be6a65e26a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.350191Z digest=sha256:cbf8f77590586d0583f0eadb902b1b050aae44483408facb1a246c02aa28b18d

Observation 28c8ddea-9165-4496-a0d0-ea487f9a260a · outbound

This paper cites Proximal Policy Optimization Algorithms.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.369712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.369712Z digest=sha256:390ebbaf4a402c67989f35f64ae0320946d0fe233346d9014d282c5d811f7e96

Observation 09744d3f-4f29-4648-8b1d-fa749b4b2543 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.391054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.391054Z digest=sha256:23e81e23b62d37bc461ead0e6bc52ea2645e5fc0772579d75cd2039dae3b796d

Observation 3fb3d1a6-01cd-4bbb-aa48-a8eb1b7bc719 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.205422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:28.405993Z digest=sha256:7b8f7fa16ecf80ebe9985ef0025ff98549c0705c35d756b463fba30d91c626f8

Observation df3d18f0-709c-4db1-a285-a23af471f3f5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.422267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.422267Z digest=sha256:052e11aa83c8f1b7fed22afa3a3d61cdb3937d72ec20df6ebfc479e082a69fbf

Observation 6696a32d-4608-4ef0-8de1-58e824fae9d1 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.444374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.444374Z digest=sha256:b59cbdf8b0b420e41045226ee0b4d1196794de1719fbda604225e2c2920c8712

Observation fe87a75a-7355-42d1-8540-f8e5d4d0c008 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.463737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.463737Z digest=sha256:a1fc068fa1ca9c95ccd70dd9a93d5af09a211d42d05a5aca97bea6138ef34e67

Observation ee98e613-7690-4797-9504-94d2f83e0d75 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.485026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.485026Z digest=sha256:a4126bd76020adab9a158ee4adaf392d31a82bdb2a9e34585e95b6426f8231bf

Observation f7352591-2e0a-4fc6-99aa-2ddd81b2977a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.505793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.505793Z digest=sha256:14e5deee667a0c7220310626ca46110e14a554d550b3ccac729f1631633b884e

Observation b32addc3-6660-4748-8b35-0f3b99ab08f1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.514846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.514846Z digest=sha256:feb926e02be139ef0ca02f786ae7e77c14aa6d0d8dcc7142888657453e79deb7

Observation f559c9d4-cc7e-4725-96c6-91dcc2de487d · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.645747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.645747Z digest=sha256:7315157a4862cd91d436fb76d84cc843f06a1ac5180321ee352c56272e4cb856

Observation 4f21f1cc-1561-4ae4-b7f2-662c6d7cd173 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.670138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.670138Z digest=sha256:bc07efa30eb69e6c797066966f3cb46bcf6ba0907ff0899f0d00e45e287e0378

Observation b9d0598b-3945-4573-bf7d-2516c8da6e38 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.699089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.699089Z digest=sha256:62e8cad5e85e4585d26a7e524ea6a0154f206de9121c10f1eb4310d38d7ba41a

Observation 9cb4bd50-cb27-4851-be18-111a1b7de6e1 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.144612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T22:29:58.706228Z digest=sha256:56bf225b03cf0ac58a9264ba0b6d2c08d833bde9cc83b2fb70a835085e5d9d3b

Pith citing papers

Observation 42811ffa-b70c-4621-bfea-976bd4e8bfe9 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.389824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:1d92e6edf26740632479faff60121b68df05aa29d2b6c4ca094fd014feee2a88

Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.261864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:e7aa9ab0fa1ecd4f67d053e2c420c227f105d5bf714d63a6a6ac84e5facd21b6

Observation a858a48e-99e8-4637-bbd6-e9b8dc097144 · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:29.016519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:6be0a80ee224965c1214b6bb3ffacb309b081ff2fb6d691ba69c91b1a715a05a

Observation 16209c2b-dcd3-46ce-befd-a980d05c49f3 · inbound

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning cites this paper.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:03.801892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T15:06:38.215137Z digest=sha256:00ff3e067ba423e819c33ba00e422a94260df65c59471ed6c75016ec474c1d9f