Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:39.981501Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 13 inbound Pith citation observations for arXiv:2509.00676.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:39.981501Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:05:25.733973Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:59:58.567326Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2fc229f3-c54a-4901-ad3e-1b51b820e4a1 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5727700-f78a-4c26-99df-12418ed65885 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5250fb7f-1788-47ef-9e57-2043d760ea48 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b34e68-9e88-4e54-865c-cba0aaa15173 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb47259-c3aa-4b8a-b303-e66ff9a16ecf · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e84394d7-224c-45a2-b0ab-6454bf1f3611 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fffd557-bb35-4f9b-bed8-0e1784643ed6 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3588781-8ee6-4566-963c-9e74225587a6 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71821fab-acc8-4ada-a812-97ad3ddfe4df · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling laws for reward model overoptimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f011b7d-e72a-4d58-9801-99b13bf437fa · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Interpretable Contrastive Monte Carlo Tree Search Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3971631-1abe-431e-9f16-cf3e1610741b · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa570e30-2544-4fe1-8f67-24262a221426 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac3de43-4733-4cfc-8917-fae8087cf7ea · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b349b044-87c8-4a74-8c06-fc34d9966ed2 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a538029e-3803-4a07-aa4f-de78abda9025 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ced648-191f-4b70-8392-a741a2529c46 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model A diagram is worth a dozen images, 2016
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef2a548f-1c98-4a1a-830e-a71a6d157ba2 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Process reward models that think
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 178119a2-420c-43e2-891c-510f2a3b9fbc · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de18d49c-57ed-402f-a32f-3de92e04252b · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c40fdb1-95b6-49e4-821c-730d39f4ef74 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Let's Verify Step by Step
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fd7618-4392-4cec-957b-96f152aab42e · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Let's verify step by step
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3305b9b3-48f4-453e-aaee-f65503ffafa2 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Visual instruction tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e770e54-9690-40ea-86c1-9bb263ec143a · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Noisyrollout: Reinforcing visual reasoning with data augmentation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07490fcf-1e7d-491e-92c7-5752b4acaefe · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31857890-368e-45ca-9821-07fb765e8b8a · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 723e0e19-0b3c-4c7d-80a3-904aee3178e1 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74715f75-17e2-4199-8624-cfb7477650ad · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Generative Reward Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a250e68-03df-442c-92d8-f8ca24badbbf · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c55f1b8a-1eaa-499e-9d5f-ff3c7c2de962 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34765e4b-2911-4aa3-8025-95466ac147c8 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Gpt-4v(ision) system card
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51c071c5-8514-4663-b398-99f275f20ad1 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Learning to reason with llms, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3fcc1a58-bc25-43b9-bb6a-09da4fb64832 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Training language models to follow instructions with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946504b3-bcbd-45dc-8439-217a274ad464 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9be0652-2f35-4f3a-9ca3-c9eccf814e1a · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dcf6623-c1af-4cbf-9a42-11066e888fad · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Vision language models are blind
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bc0b383-70dc-41d8-b6f7-78728df2e511 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0c2e3f-79b5-4094-bab6-81b11772e5cc · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d408ac-a580-4b8e-afc5-0296b78e67ed · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43047e17-8353-43d1-9bfd-bdee1207effc · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1283b163-a25e-4989-99da-ea011bde4706 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MiMo-VL Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879f3a98-fe4c-4a9a-8316-1e03eaad7e5a · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa75da25-bfe4-4e13-822e-b46c35599431 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6164018c-232a-4b4c-9561-770c7be1db9c · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0805543d-08da-4fe1-8fcf-2341a6090190 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7370178e-d63b-4e1e-a69d-d9e8cceec03c · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Measuring multimodal mathematical reasoning with math-vision dataset
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e30e8a2-f9a7-480f-ac96-142f40427ac7 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e5b6d7-ed0c-4f79-b7bb-65ec5234e508 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9baea096-df66-4209-84a4-cc84c740d227 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e2937b-eda5-4503-9d59-4ba53a5bd6b7 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3674e452-b019-4857-80a8-1c47202b94e9 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Chain-of-thought prompting elicits reasoning in large language models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad221a9-477a-4640-93cf-f9ec647b244d · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c446d9-6703-4203-825c-4baea7897aba · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768a6749-4d68-45da-9c1c-064e1b3fc2f9 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed8bf8c-fce5-4454-a622-c5bef9720374 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d55eb9da-3eb0-4479-8494-2aca89b2e0a5 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f57d251-a640-4578-b521-e592117281b3 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Llava-critic: Learning to evaluate multimodal models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b840d0aa-88db-4ee4-8d7c-aea8368770aa · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb46d933-46d4-4af1-a71d-31fb270055d6 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c80213e-e434-4764-b9a1-dd9b0fe19e70 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea25225-aca4-4e11-a217-6a2b89ddfc92 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5e40fb-ea29-4f77-b6e6-136cf38cf54c · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2b7650-77ba-4a4c-8f8a-f9cdadc85895 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ba5982-1ad0-4cc9-a6b8-e41474e70ca2 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d24ae44-7827-42b1-bf90-4ea7efc56487 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7afc8e-51e1-4898-9593-28a73edb2c14 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17e8c60-dd13-46b4-9bf3-6a46a0cd282b · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cc9626-792b-4408-b9c4-d300a389e698 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmvu: Measuring expert-level multi-discipline video understanding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93c03e7c-9b72-4505-99e9-f8fe3882c242 · outbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b27b417a-d61c-409b-9100-0b2d4f172ddb · inbound
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c68e6a67-1d50-4550-87ab-dd7ed69a86fe · inbound
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bdd290c-5381-40d6-bcbd-09890ecf0064 · inbound
Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1a125c1-20ee-4954-a735-d66748f1417d · inbound
Watch Before You Answer: Learning from Visually Grounded Post-Training LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d21428b6-34ea-44c6-b988-a1c1f5f3883e · inbound
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a247ef6-a82d-46aa-b669-1071141cac14 · inbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6e06dc6-c571-49a2-9719-de29a1412d78 · inbound
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb90c0cd-3b97-4812-80b6-710292ac054a · inbound
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc854a1f-8dce-4537-9616-c1aebfd183d1 · inbound
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b631b8b4-6629-4e9c-b17c-c41a0306dd70 · inbound
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15ec6608-c26e-403a-8397-d6582a6b4fc3 · inbound
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98c9f657-6b68-4eb6-8288-294bb38a9e9a · inbound
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a3c139-f196-4cae-aa08-cd8e0f2cc00a · inbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.