Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:48:37.725055Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 13 inbound Pith citation observations for arXiv:2506.16141.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:48:37.725055Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:35:11.261146Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
50 of 50 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 1f1d5302-beb0-423c-96f4-79e6943e7596 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24167e82-b909-4652-9040-1d7369cbefcb · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96bc01f3-3303-4e46-94cd-badd1d66a95a · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a8d6fe-123a-4013-9744-9ef7209e42e7 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aedcf7f-19ce-4c21-afc1-14879ba2e61c · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87668594-5f01-4f8b-87d3-d9f12a73318e · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba8fcbb-2e6b-49cc-a7f8-6b01284932df · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1dcbb12-835c-47db-b815-97acdf0acfd0 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d460403-1db4-4434-a36a-5a883226b3fc · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01639783-15c5-441b-99d6-50825f3a2196 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e860905-a7be-4343-940b-cca4e68e3f50 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-r1: Reinforcing video reasoning in mllms, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b678cb0-2ede-464c-a21e-6c7e4010fef8 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed62ce0-a034-485a-bf36-63085f2fbbd6 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caef7470-ac90-4e9c-8f9c-ee95ebbe62f4 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ff65fe0-73c8-4310-8e5b-87392040dbc7 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Ego4d: Around the world in 3,000 hours of egocentric video
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9c40393-dab8-4541-82f8-5cca8d1246b1 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Proximal policy optimization algorithms, 2017
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba36716a-bc1f-4057-bbe6-55210a24beae · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55eb96e-5451-4977-8298-c7ca82c2048c · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55dc7406-8b3a-440b-b8c9-44bd12456a59 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Let’s verify step by step, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0924cf59-4514-4f59-8d31-deddba17e0f4 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Solving math word problems with process- and outcome-based feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703c21e7-fa80-4e50-8baa-1a99f975d6ab · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Alphamath almost zero: Process supervision without process
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 210ad232-6d95-4711-a1dd-85c5fe2758f7 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0feb8827-418f-418e-8da2-a7028c875bb9 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e166373-3cab-4d7d-93c9-1559829a576c · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4249718d-264d-47b1-9722-2ef4abd5d1ae · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Evaluating mathematical reasoning beyond accuracy
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97e55492-936f-4c0e-a906-accf45d41bb2 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a53331-fc24-45c0-93c9-e2aad67d13f5 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Warp: On the benefits of weight averaged rewarded policies, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a7d7d8c-b443-443e-a03b-13792fd8bef1 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c747885d-dd23-4ad6-9ada-f9fdd4f839b3 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5527d3f9-a4c7-4a69-9141-87eaf4ece1dc · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Open-r1-video
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4b54ff6-e282-4069-af2e-5c4dee1f3d62 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4042c7ab-253b-407c-887f-31b9ae9694a2 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7c72394-096f-4d90-8344-71c77b1bb301 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dfew: A large-scale database for recognizing dynamic facial expressions in the wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0df6ca4-fdf4-447a-850d-7e9688d04544 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a0426f-ac6a-4cdc-933e-c641f0d5a4d0 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db357cd1-ea01-459b-945c-e6a9cfcee43c · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6b4a07a-446e-465a-b0b2-fc16fed427ae · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mmbench-video: A long-form multi-shot benchmark for holistic video understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8120ea48-5a60-4ed9-b67c-264cbe53d991 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Qwen2.5-VL Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1bee5ad-baba-454a-9bb7-b7384a0c64e1 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning GPT-4o System Card
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bd632d-3da6-4f60-8a1d-5a5eb892d476 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Llama-vid: An image is worth 2 tokens in large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa98f0e-5053-4c37-abe0-bb1c03d95e62 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8f8ce4-9400-4786-81bd-7c484317d5cc · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c136af-0e49-4990-a1f1-2c94e53f5172 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vila: On pre-training for visual language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 847d3506-46b5-4ed4-996a-c47b8e2c793e · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unhackable Temporal Rewarding for Scalable Video MLLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd5f36df-dc5c-4131-8cbc-28b81e2c7ef5 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6198fb0-460d-4309-8991-d026ef49699f · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2d0813-3723-406c-9f37-99613b9ef861 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb50d06-db4a-4122-b1a9-f8f7b1410dc9 · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93b68ad-4f32-4fad-90d7-69d5909b854a · outbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c1c08e-f4f6-4914-8582-a873fff20c39 · inbound
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8c59c9-589e-4ac1-9487-00d555b75867 · inbound
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea14883e-98a6-4fb9-9ef4-a8371292709f · inbound
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0bfcdd-be2f-4489-b396-47970bb67ef4 · inbound
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9c14120-8beb-4b12-8b0b-274c4bf4b9d0 · inbound
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3429f25-bb38-4d17-b2ff-70b78648a36c · inbound
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b56a2e3-91fe-42c4-a10f-3c2b1b66a312 · inbound
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14e8d7db-96f7-45bf-99ff-3b55098fb9db · inbound
Touch-R1: Reinforcing Touch Reasoning in MLLMs GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7dbb6506-44d7-4151-a07d-775cf052e55e · inbound
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ef5e2ec-7c53-4b66-8092-847994b7ba01 · inbound
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4b025b0-8a22-45d5-aae7-e2db5fb881f0 · inbound
ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d83f06a3-2c79-4a41-b0bd-fc3c6771115c · inbound
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad9f6d09-1dc7-4dd0-9c3c-8b7df3220910 · inbound
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.