Pith. sign in

Paper Citation Record · LEDGER

Multi-Branch Policy Optimization for Multimodal Large Language Models

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.07581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07581 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:01.084911Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcc858e9-2c85-480e-b2b0-86ffc1b3ffec · outbound

This paper cites GPT-4 Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.968747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.968747Z digest=sha256:30f6ccb496213f5c772eeeb5d2c2b6d7b2ce004a3eae12cd86237fb923c03443

Observation dfd166a4-ba88-4690-96f9-c1abb37e5db3 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.557255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:00.972707Z digest=sha256:ffcdc1c83da67c80aa5b1e4ea8d25a262c1931ba4f51de2f0bf80daf0b90812b

Observation afdffb96-1086-4ee0-ac21-d9691563c3d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.975962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.975962Z digest=sha256:f09f85ccfbe458b27986dd9f8e4c74cb50127356eb9c60083d62e4a16d5bfbb9

Observation 42ebee1d-af1e-4d55-962d-4173f2797270 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.979284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.979284Z digest=sha256:7c0e6b39ee0fb1eb840919de924b053c04910bb7c83179b7447fef042bebaca1

Observation b48126ed-e5b0-4c55-9e1c-282650858c1e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Multi-Branch Policy Optimization for Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.982550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.982550Z digest=sha256:cbc9be358c59ef5dad9d66344074895e3bf62cc1a6336ed528b8d64e7b390c1f

Observation 253aa727-6c76-42a7-95bc-4e52d27df8c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-Branch Policy Optimization for Multimodal Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.985579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.985579Z digest=sha256:32d4cd7998adcf697904ec4950d6429697803aeacee3d20da176f77a3c8359d0

Observation f7d123c6-4c8c-45ee-a00a-f1511d78138a · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.989085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.989085Z digest=sha256:689fa3e98fa099b53c048dba9409a88c9f67b3f141bb7107306f150f3f7f8fe1

Observation 267c2edd-0536-43d5-abcb-1d548312f3af · outbound

This paper cites Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward.

Multi-Branch Policy Optimization for Multimodal Large Language Models Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.994456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.994456Z digest=sha256:064a982f148a2d81267aa4a653d22dee99b5dd58bef3ddaa3718ea38f29bf8a4

Observation f6e568da-3383-4219-aed2-934e79408e10 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.545458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:00.997566Z digest=sha256:f32b745ce6a83369fba78be3fb9f9f1b31ffe1d000a4c28f1b338da5b1c771f1

Observation 4a4c56c3-cc2e-4222-986b-f8c8e49df15d · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.537651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.000236Z digest=sha256:0d90a3b393bc50584e8dcba2c727400f23ccd92b1a10f8c484997e9f70102125

Observation 113237e0-7488-4f53-ba33-eab608327b9f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.003086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.003086Z digest=sha256:e5aad2dbfdd205dd0d87ec32cdf523dbdeda0621cd4e96ffc12afce2a45a2157

Observation 97492cea-090f-4037-a82d-ea53535adefd · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.529428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.005837Z digest=sha256:f7e3c0708a1a0a8a041f5f5cd6a33eb60fffbd920ab48ac8f551ead398e1dd0f

Observation 23605e2c-8ddd-4081-9b79-0d8fc15a111f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.008576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.008576Z digest=sha256:ab3e5fd2acae51117bcd1418c422d2426fbcdef828fbdb0f71b453d2e86bd6ac

Observation b7060bc9-2abc-4d04-9b92-81cc05cc0c4a · outbound

This paper cites OpenAI o1 System Card.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.011369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.011369Z digest=sha256:7c13c0005be539b379e365d0eda2a61d0c073bebbc9fd99cc6b2274b21f45fe1

Observation aaad2af5-b7df-4a2a-bb56-63acccfec38b · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.521879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.014447Z digest=sha256:b9091079fe6a871bb7d1b261e6ae414f01b64a4ce156057cc047f41ed3d5e531

Observation 3c826a07-a80f-46a4-b267-3abea1d0865e · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Multi-Branch Policy Optimization for Multimodal Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.017092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.017092Z digest=sha256:639197b2efb6102d9a9d035c54eade2a2156a4c84210bf96af784ddeb1819e76

Observation e9d5b877-3a00-45f1-be99-124b6b875cb9 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.019846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.019846Z digest=sha256:b06804259a1109b2d095abd4eadf3af5af8a17399e6502edf49a38b853986b6d

Observation ba49155b-49b8-4ecf-87a5-ee1ca0475d31 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Branch Policy Optimization for Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.022292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.022292Z digest=sha256:b78811dc49f2e461e722e815ea278efbdfcc1bed6d786983bf6c18fcd5da7acc

Observation 09a0a754-b403-4527-be31-829b8138caa7 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.025134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.025134Z digest=sha256:7d61abb77c2c8dd1b6eb09df425e8982fbe8aabde411e3a4d425672609bad24e

Observation 21633ddb-1b6a-4adc-bebf-3f93f69a07f2 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:36:01.513081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.028118Z digest=sha256:52984141ab86cbe88211fe41c8306efef1c9830b0439f3d21cad075bfe17a9a2

Observation 8b4c6a54-e0d3-4daa-86c6-0a10897c94ef · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.030834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.030834Z digest=sha256:54b8dfa6c008cf3e58ebe3c9be6e016b9bb729e4443b72b5cce8f4b37292117b

Observation dd4ccb93-5a99-4f34-81cd-c8dca86021e4 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.033818Z digest=sha256:394bfc610fe33ff0b09af40adf564a602a77c043d34aeaf5b1a1955e275ddf37

Observation 1fd55159-d00a-4060-885b-2c879ece720d · outbound

This paper cites When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation.

Multi-Branch Policy Optimization for Multimodal Large Language Models When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:36:01.312293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.036574Z digest=sha256:26a28e6564b559ecaae0e743a184c6dc695bdc7af128a8d653804ffcbd9ed6cf

Observation bf870af3-7ec5-4309-bf4f-d650250b6261 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Multi-Branch Policy Optimization for Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.039435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.039435Z digest=sha256:5e86979620c5ddcf4266bf81a1207910f4fbc907274e35fb619116390b69df85

Observation d30159e5-8b9e-49f6-b35e-03a37328e9b4 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.505262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.042321Z digest=sha256:e3c342125bd8bb8f9d11f04abecba0ab04397a7b1c209e319b161c720eb7991a

Observation 37162f79-8e37-45d3-90bc-e8e0baec79db · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.497435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.044931Z digest=sha256:ed5d459140885ba1479a0401f467d73f8ad30ee4feeb6affac6324e27c38d2a0

Observation dd410201-fb9e-493f-82c3-dd805918cdae · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.047616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.047616Z digest=sha256:0abd41967b37595c74a9f8585f31ca38c51ffd534c9c5777eba1f03d4a83854b

Observation 6b83e7a8-4513-484c-9150-04db192292b0 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.050562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.050562Z digest=sha256:9bc1f45529ad9a3999d86a75f6da2e785b529fb7142b750b1e26c4c4e9b2c264

Observation 21dea8d1-a349-4d73-afc9-6085a711eb7e · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Multi-Branch Policy Optimization for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.053166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.053166Z digest=sha256:45dfeaf340a62a7b0fa3d0cde421c4e91fa2a9dc195bdb8bd769c8426400993c

Observation 8403012f-6a3b-44ed-aa5b-92cdd066c3d7 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.056098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.056098Z digest=sha256:29132c34db443cae626389186b4654ab7e6d8d68116173074ae18d8ad883dbc6

Observation 4cbd1c43-68bd-457b-9c40-68ad9b565561 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.058958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.058958Z digest=sha256:fcbe0d5c99c6e7be4122aa8196f260191d74b6edcce2882d7676c3ce5998981f

Observation 7c5df10b-470b-4ac5-98be-8c1270150d93 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.061794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.061794Z digest=sha256:016147ad7c09691acaf094cd03ccff47dd4392575c6d3557e89c210ad5b754de

Observation 74ec1b88-93e0-49d3-af6a-f1c2b3b09c3a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multi-Branch Policy Optimization for Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.064518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.064518Z digest=sha256:f381876ddc462a7e1e017abb807adc8af862a6ce0ebfe76b9e84fc02b8780d18

Observation 3f1e23d8-8ba2-4c1b-8af7-a4fded0586a9 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.067584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.067584Z digest=sha256:8e3c7e7d83778393f0a8ae44e38f171a5a4a51e3d62551ab2240dc145cb68dc8

Observation 4032bd1c-5083-46a1-a163-72ec5eaf19d6 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.485564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.070552Z digest=sha256:596c3b35b6b1053f316c635f148cecaa2c5d4a91d41a9ac768163c94cfbe558e

Observation abf73bd4-8eb0-4be9-bdda-dc5ad1175fe1 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.073147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.073147Z digest=sha256:b63063ace7a74a33acbb15b3ed6da9de8330f2d8fda5900d1de14b113fcadc07

Observation 170634ff-68b2-45db-8775-ec9d71c43d34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.079000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.079000Z digest=sha256:adf9307c8b2db4fa21d21682872150f7727c568b20e00c1abf30a1a43d2b8481

Observation 9d5a108e-0298-4566-8b53-50a1771d1f31 · outbound

This paper cites Group Sequence Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models Group Sequence Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.075961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.075961Z digest=sha256:54e4ac4eccf3e63ea9ca7fe8000ac100102613e580baa4ac51a516a3835ce997

Observation f6ac86ed-8a9f-4448-be1f-c4ba65ce1736 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.084911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.084911Z digest=sha256:d72e4d57b5530b742f7fd10142ccf034c8748b3276d84be347844b701ca8e8aa

Observation 9b9b55b2-539a-4808-97f4-9198ae0094d1 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.081871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.081871Z digest=sha256:bc60f37d55896813f2003493ef471d2a19cdf6baa5b16e1609dc6de90955c696

Observation 6179ed9e-bd58-467e-a60d-6adb24ec082f · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.991604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.991604Z digest=sha256:86e8b4fded95a920781c85002011777ecaba90125a264565253a51be0f347cad

Pith citing papers

No inbound Pith citation observations are available.