Pith. sign in

Paper Citation Record · LEDGER

Multi-Branch Policy Optimization for Multimodal Large Language Models

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.07581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07581 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:01.084911Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcc858e9-2c85-480e-b2b0-86ffc1b3ffec · outbound

This paper cites GPT-4 Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.968747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.968747Z digest=sha256:72421e2493dac6b92a85880a28a6b92451e5888677b3205fca90c63ccbd1226d

Observation dfd166a4-ba88-4690-96f9-c1abb37e5db3 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.557255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:00.972707Z digest=sha256:5fadd85c5b7b1a7e8ac7a5bc03dbc6c0c67f6f575b9d07eb63074e7f1cafff6f

Observation afdffb96-1086-4ee0-ac21-d9691563c3d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.975962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.975962Z digest=sha256:9bd9b683661c7b20a7b1352319b824ad80929271a96cec661f8080563315cf7e

Observation 42ebee1d-af1e-4d55-962d-4173f2797270 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.979284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.979284Z digest=sha256:705d8e9cf93e67026872117eb39e747d96a6e054204e4c526013303e4c426f22

Observation b48126ed-e5b0-4c55-9e1c-282650858c1e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Multi-Branch Policy Optimization for Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.982550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.982550Z digest=sha256:eb7724c633598c54cf549453f6726968c17a476bbf3040c73136a83af85d12ff

Observation 253aa727-6c76-42a7-95bc-4e52d27df8c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-Branch Policy Optimization for Multimodal Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.985579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.985579Z digest=sha256:db2d6bb1577680f7da95083c4ddc51e111ddfa63d56ad704552828a7b5d7975c

Observation f7d123c6-4c8c-45ee-a00a-f1511d78138a · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.989085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.989085Z digest=sha256:203bf2e43671528bdd482641ed6f45d9009e7b2814d59d758f2bb6c60afc18f8

Observation 267c2edd-0536-43d5-abcb-1d548312f3af · outbound

This paper cites Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward.

Multi-Branch Policy Optimization for Multimodal Large Language Models Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.994456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.994456Z digest=sha256:27c82e24e5a74f5d8adf5f596317b1a8476c8018754ff30f7242ce2ce6eb5943

Observation f6e568da-3383-4219-aed2-934e79408e10 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.545458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:00.997566Z digest=sha256:0d11a19ba59c62b505e3053c76cd6d6f3c265c839e844c6269165d5b3939e20a

Observation 4a4c56c3-cc2e-4222-986b-f8c8e49df15d · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.537651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.000236Z digest=sha256:22df181151921b891cc5b743fd30e1f77631687cf1c612d3fe0abd249ea9b50b

Observation 113237e0-7488-4f53-ba33-eab608327b9f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.003086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.003086Z digest=sha256:cf61db9be6fa00991cfbbcce69e9a1b07b3fd546027245ece7f049bb5eee5203

Observation 97492cea-090f-4037-a82d-ea53535adefd · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.529428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.005837Z digest=sha256:2ece167872a458133a71a608fb132e6208a6c172112e16008e420cf044b150e7

Observation 23605e2c-8ddd-4081-9b79-0d8fc15a111f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.008576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.008576Z digest=sha256:08c6afd7f45bdf16b4e396221132074f0a6def813c94300129fe50a841638cd0

Observation b7060bc9-2abc-4d04-9b92-81cc05cc0c4a · outbound

This paper cites OpenAI o1 System Card.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.011369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.011369Z digest=sha256:bc4a5dedfa534e4b5a7357ac90ce7c05cfd77a3fda61f6e5b2b8209cb0c696c9

Observation aaad2af5-b7df-4a2a-bb56-63acccfec38b · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.521879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.014447Z digest=sha256:8e965a8c48a7f8671abfedd145efbbcdd818b9d1ff9b86ea505fab7dd0bb8866

Observation 3c826a07-a80f-46a4-b267-3abea1d0865e · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Multi-Branch Policy Optimization for Multimodal Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.017092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.017092Z digest=sha256:e2ec0f8adac928317aa368562d2c04bf97fc0790e0ed3e873f7087de37ce91bc

Observation e9d5b877-3a00-45f1-be99-124b6b875cb9 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.019846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.019846Z digest=sha256:585669671aeba3683087d3c1f68801814e8ed4b9134538077c2683b796b4fc01

Observation ba49155b-49b8-4ecf-87a5-ee1ca0475d31 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Branch Policy Optimization for Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.022292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.022292Z digest=sha256:61a820687059e56c1e2ebf2adc5df541297f44aea89cb5b7dd192a3e33d67f47

Observation 09a0a754-b403-4527-be31-829b8138caa7 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.025134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.025134Z digest=sha256:46b2878594ba00a6ebb2991f32f51d5c928ba9dd0caa45165a9a4657e825e1f8

Observation 21633ddb-1b6a-4adc-bebf-3f93f69a07f2 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:36:01.513081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.028118Z digest=sha256:3a4447b8b47b963ef2e9a82e77a0c1f48626019fe96d339e5444f1be6af78026

Observation 8b4c6a54-e0d3-4daa-86c6-0a10897c94ef · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.030834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.030834Z digest=sha256:5af84d4f65121e06e1bc6f951c74b08d05c9592d9ddf81d1a30f886a839a1769

Observation dd4ccb93-5a99-4f34-81cd-c8dca86021e4 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.033818Z digest=sha256:30b90d3e52dcdfa662b478db00534f488fedc9090581bd953bebe4e285c5bf2e

Observation 1fd55159-d00a-4060-885b-2c879ece720d · outbound

This paper cites When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation.

Multi-Branch Policy Optimization for Multimodal Large Language Models When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:36:01.312293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.036574Z digest=sha256:ffb5e23544b5880c9f58049fb481430e87ec51017083d960d8eb523038184a1c

Observation bf870af3-7ec5-4309-bf4f-d650250b6261 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Multi-Branch Policy Optimization for Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.039435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.039435Z digest=sha256:7219599a40edb6c2c866d1f6515de631839aac33234d4a3245a5ccd55322ff4a

Observation d30159e5-8b9e-49f6-b35e-03a37328e9b4 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.505262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.042321Z digest=sha256:8c9a9a066eb34b7b448eef57b739b57d5d3451c76de975592189d8def3d8a9fd

Observation 37162f79-8e37-45d3-90bc-e8e0baec79db · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.497435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.044931Z digest=sha256:e9299a40ff7d9bc16343374edcbe571cf55b7aee4394f2767007a6aa9337f0f3

Observation dd410201-fb9e-493f-82c3-dd805918cdae · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.047616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.047616Z digest=sha256:367e8943ff258f7bebf418cad87841cb3e4aec0f2534062dc8549e5329425647

Observation 6b83e7a8-4513-484c-9150-04db192292b0 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.050562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.050562Z digest=sha256:62c979623531935e4aa7a705634b1768f0bbb473449d6d369d1e280f6990c84f

Observation 21dea8d1-a349-4d73-afc9-6085a711eb7e · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Multi-Branch Policy Optimization for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.053166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.053166Z digest=sha256:34becd00b23bb50168891b807a0cd846c8140669300b617b4f917b1c119dab83

Observation 8403012f-6a3b-44ed-aa5b-92cdd066c3d7 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.056098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.056098Z digest=sha256:5685112ab929f55aa08cf3ca3f7d88d496e70d474cedf5bbbc1257f1c96c3ca3

Observation 4cbd1c43-68bd-457b-9c40-68ad9b565561 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.058958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.058958Z digest=sha256:09d1afa09a8f65659d672c0960c37a31ec43e4a0d8a48ba61f50335b0cf06508

Observation 7c5df10b-470b-4ac5-98be-8c1270150d93 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.061794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.061794Z digest=sha256:72fbadaa443d0afd3b7782b657d3b31bf30d52f37df90296d6728061554ac7d1

Observation 74ec1b88-93e0-49d3-af6a-f1c2b3b09c3a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multi-Branch Policy Optimization for Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.064518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.064518Z digest=sha256:65e72851dcac7fe16d4aa8b8b9891de4ba3ae0fc20b11150cb0aefbde06afee6

Observation 3f1e23d8-8ba2-4c1b-8af7-a4fded0586a9 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.067584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.067584Z digest=sha256:2bd793581d93431a4f5792792cea2074f072215dd93095461a90296668ef8a88

Observation 4032bd1c-5083-46a1-a163-72ec5eaf19d6 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.485564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T00:36:01.070552Z digest=sha256:2a3865c137dc5c4902cc492cea579a9bc34248b42b260af52930fc98912c0cc3

Observation abf73bd4-8eb0-4be9-bdda-dc5ad1175fe1 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.073147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.073147Z digest=sha256:d4d796cd02caad9c9a27602cc9d310d1c707deecad6d7d81df821b2647399357

Observation 170634ff-68b2-45db-8775-ec9d71c43d34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.079000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.079000Z digest=sha256:f01e471ca6835a912478b77c1adeb7ab32defb0139036860815e5917bdd49ecb

Observation 9d5a108e-0298-4566-8b53-50a1771d1f31 · outbound

This paper cites Group Sequence Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models Group Sequence Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.075961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.075961Z digest=sha256:518e64bf8c7884a95aa6a9238e005086bcc003fa77afeeb01274b7dcaa1de91b

Observation f6ac86ed-8a9f-4448-be1f-c4ba65ce1736 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.084911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.084911Z digest=sha256:71e0e2b9d68bf02f6e2277d8f19509c0b6dd8a068f7e89018cba6a8168df26c9

Observation 9b9b55b2-539a-4808-97f4-9198ae0094d1 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.081871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.081871Z digest=sha256:b7142542838d1b49c33b5f590d0631234b2687f2d37e75fd2333fa1d881092d9

Observation 6179ed9e-bd58-467e-a60d-6adb24ec082f · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.991604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.991604Z digest=sha256:0f8771ff93dc40fa1b5b221edc0712d2ccbbc1d3330cc6daaecb4d0c2737c6a9

Pith citing papers

No inbound Pith citation observations are available.