Pith. sign in

Paper Citation Record · LEDGER

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 13 inbound Pith citation observations for arXiv:2506.16141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16141 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:48:37.725055Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:35:11.261146Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1f1d5302-beb0-423c-96f4-79e6943e7596 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:32.974881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:32.974881Z digest=sha256:f10b6c82457cb08f50bbf03ea69585ca6d1dde5fe666a59e7756c9ae6eb92a4b

Observation 24167e82-b909-4652-9040-1d7369cbefcb · outbound

This paper cites an unresolved cited work.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:48:41.128978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:33.018546Z digest=sha256:7afb160e8f47c6670004a8c58223e55c951b3e10841c10a7f9458689c87baedc

Observation 96bc01f3-3303-4e46-94cd-badd1d66a95a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.083995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.083995Z digest=sha256:33acbd9e8fb0bf02e5e76d165f75be8874cc981af251bc700b08a77f423e8660

Observation 07a8d6fe-123a-4013-9744-9ef7209e42e7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.164335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.164335Z digest=sha256:edffbb78cbf33f945ffc6a203d59f546e28231d5b23624f954f2c63008f230c4

Observation 7aedcf7f-19ce-4c21-afc1-14879ba2e61c · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.223889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.223889Z digest=sha256:32669c514a9dfacdbf70d1ff8734f160e95a3f377ad5e87297721fad83c78341

Observation 87668594-5f01-4f8b-87d3-d9f12a73318e · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.343894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.343894Z digest=sha256:2e94416c0482ff38f7cf30cffb69f268a32e3dd8faca88cfb9755629b567f75b

Observation 7ba8fcbb-2e6b-49cc-a7f8-6b01284932df · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.451030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.451030Z digest=sha256:6783a8eae14a4f87d1330613c3ccc149001282e393dab71e5d6e7c5e1617bd58

Observation f1dcbb12-835c-47db-b815-97acdf0acfd0 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.585809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.585809Z digest=sha256:421f9f8fb953e369d9a1d7f53f6a5f2408f2e2572701a7c6b6e335f5025a7258

Observation 4d460403-1db4-4434-a36a-5a883226b3fc · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.702470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.702470Z digest=sha256:ee32cea3981abc41eec81c9fb30d1d163579052d260e2b1c7ed5b24958ed6b95

Observation 01639783-15c5-441b-99d6-50825f3a2196 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.790717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.790717Z digest=sha256:604f533162381da75d2cd5ba107d228de00226995d36d0f3351f49e4dec10e27

Observation 3e860905-a7be-4343-940b-cca4e68e3f50 · outbound

This paper cites Video-r1: Reinforcing video reasoning in mllms, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-r1: Reinforcing video reasoning in mllms, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.869684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:33.849002Z digest=sha256:fe7844e0a8744a6739f64d3dc6c40831f608a7d6f28459a73815cc2de8767acc

Observation 1b678cb0-2ede-464c-a21e-6c7e4010fef8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.921347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.921347Z digest=sha256:97544f2ca4ad28289b9845f7bebf2fd7f50d2a363a58015d34877d4cecf6c6a5

Observation fed62ce0-a034-485a-bf36-63085f2fbbd6 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.993024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.993024Z digest=sha256:1457069abd86198780db1b2da4c0996f7a686d293d14f040e08966356b4f5d05

Observation caef7470-ac90-4e9c-8f9c-ee95ebbe62f4 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.652360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:34.042645Z digest=sha256:c4db82fdb5a09bbcc307d4a1497bd7e153047c7bbb6adcb276f0debe0cae1510

Observation 5ff65fe0-73c8-4310-8e5b-87392040dbc7 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.135877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.135877Z digest=sha256:afe6c5c796e1440791ec462bb4a83981d722fb95f135dbed94f2aec351830194

Observation f9c40393-dab8-4541-82f8-5cca8d1246b1 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Proximal policy optimization algorithms, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.372247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:34.212172Z digest=sha256:70d204ce9ee23ec62b719f29915b403e2eebb68f5cb9ee5cd306a93adf673ce7

Observation ba36716a-bc1f-4057-bbe6-55210a24beae · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.320914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.320914Z digest=sha256:45c18ad7efe27486ba34dcbc810217cfbe31b7594ba1042e79ccc7d2caa8e2dc

Observation e55eb96e-5451-4977-8298-c7ca82c2048c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.438236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.438236Z digest=sha256:ab371da2ddc6ab85c0772b94a9d7e935931662e04c0a9aa7a183e5cf5fab6dc7

Observation 55dc7406-8b3a-440b-b8c9-44bd12456a59 · outbound

This paper cites Let’s verify step by step, 2023.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Let’s verify step by step, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.522538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.522538Z digest=sha256:59c22c6f925ffa102444388c4f7a98633a90f07ae8b5926dd2d529bfabe2455e

Observation 0924cf59-4514-4f59-8d31-deddba17e0f4 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Solving math word problems with process- and outcome-based feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.603403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.603403Z digest=sha256:463318809478888145460e6876ae605126cce24bbf6571f04be55ebfd59846f4

Observation 703c21e7-fa80-4e50-8baa-1a99f975d6ab · outbound

This paper cites Alphamath almost zero: Process supervision without process.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Alphamath almost zero: Process supervision without process

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.165451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:34.660272Z digest=sha256:56a82c4e472820b8a0d5254d1935308e04dc09d54714641310594fda6a470192

Observation 210ad232-6d95-4711-a1dd-85c5fe2758f7 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.742575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.742575Z digest=sha256:15707d68034e5475fd2de6ca4b40f01add24e5063b36237319a49bc716580c39

Observation 0feb8827-418f-418e-8da2-a7028c875bb9 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.838198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.838198Z digest=sha256:e4f86f19f210d3f05700ea35a24f325cafa9f02d1548421c7e148a081b3c3b83

Observation 3e166373-3cab-4d7d-93c9-1559829a576c · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.894129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.894129Z digest=sha256:63f348673f955fbc752c1d39d1d285e37230a31cbd7a415a50046722694bacc7

Observation 4249718d-264d-47b1-9722-2ef4abd5d1ae · outbound

This paper cites Evaluating mathematical reasoning beyond accuracy.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Evaluating mathematical reasoning beyond accuracy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.929136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:34.955288Z digest=sha256:cab929048c98ea30715a0c03ef723a9fc0b9917d567df3cd188c8a70b40aafbb

Observation 97e55492-936f-4c0e-a906-accf45d41bb2 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.066396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.066396Z digest=sha256:b2e843e25f8ff06a1d9a1631c02847395e97fcd5ba927493af6cfef4827e4ec9

Observation 82a53331-fc24-45c0-93c9-e2aad67d13f5 · outbound

This paper cites Warp: On the benefits of weight averaged rewarded policies, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Warp: On the benefits of weight averaged rewarded policies, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.715436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:35.168865Z digest=sha256:5dc133d158add3814aee60270f9bdc6a7075ccdab13afdcbee384d03d40e6643

Observation 3a7d7d8c-b443-443e-a03b-13792fd8bef1 · outbound

This paper cites Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.276903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.276903Z digest=sha256:d4b19cd892b208179890dbd90a3de9f05b59ad6b20db57ef8d04fc33133e5704

Observation c747885d-dd23-4ad6-9ada-f9fdd4f839b3 · outbound

This paper cites Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.435084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:35.459609Z digest=sha256:c0fad3bff6df8f3b6bf04be38d253bb7e0024924c10fa5047af770b9fdcf0b4c

Observation 5527d3f9-a4c7-4a69-9141-87eaf4ece1dc · outbound

This paper cites Open-r1-video.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Open-r1-video

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.241098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:35.524033Z digest=sha256:474ee1e245d5acd2f4fb76652faaf0214157a0579584844e22562ee1151300f3

Observation f4b54ff6-e282-4069-af2e-5c4dee1f3d62 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.612776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.612776Z digest=sha256:06e4c179d1f8ec207e7a2bac0e5adb962a393f177c2994af40dcc2f439606899

Observation 4042c7ab-253b-407c-887f-31b9ae9694a2 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.023697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:35.679709Z digest=sha256:157543b98977af9e4a05573431b4b08e0de22291f4e40b149c0f42e4db98d6c7

Observation e7c72394-096f-4d90-8344-71c77b1bb301 · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dfew: A large-scale database for recognizing dynamic facial expressions in the wild

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:38.821362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:35.767077Z digest=sha256:95fffe41a27ae23a865a4254ac117533110268287e9304d9ef5ae26b56411884

Observation a0df6ca4-fdf4-447a-850d-7e9688d04544 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.861496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.861496Z digest=sha256:2561de419202c7c8244a52728f7b28a51188f716346c2f92c72e4b6a064309f4

Observation 22a0426f-ac6a-4cdc-933e-c641f0d5a4d0 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.947030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.947030Z digest=sha256:c1e299441dac5358cdddbaa5445b4482822a7473fa5110f9729cce69358d373a

Observation db357cd1-ea01-459b-945c-e6a9cfcee43c · outbound

This paper cites Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:38.593156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:48:36.039772Z digest=sha256:4b2cbbb6b5cad3b2541d83a53b81bdb72fbdc7415b43e91bb8aa796d3eeae278

Observation b6b4a07a-446e-465a-b0b2-fc16fed427ae · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.245848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.245848Z digest=sha256:2edbff2971d2d32696eccb0c2b4909cd2ee3072a0c4bc5c7e1fb87fcdac74df5

Observation 8120ea48-5a60-4ed9-b67c-264cbe53d991 · outbound

This paper cites Qwen2.5-VL Technical Report.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Qwen2.5-VL Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.371774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.371774Z digest=sha256:24c7a9477176594a2658b7b27030f463560210e7f40e8a188dc92c9bb6ce62a5

Observation d1bee5ad-baba-454a-9bb7-b7384a0c64e1 · outbound

This paper cites GPT-4o System Card.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.504016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.504016Z digest=sha256:31b8a3b4a1d380f7613e1049865ae06b4427358a416f0d1d42e63fbbe1eae869

Observation 52bd632d-3da6-4f60-8a1d-5a5eb892d476 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Llama-vid: An image is worth 2 tokens in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.600274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.600274Z digest=sha256:8c9851df630dc9d1f38bf7b9731b9e372187b20dd0a9e9a5983a5ac19562b946

Observation 9fa98f0e-5053-4c37-abe0-bb1c03d95e62 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.680253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.680253Z digest=sha256:6163f87129856e736c8c5cb03747e7d50cb80585a64dfcbff58e206b4a2a1023

Observation 0b8f8ce4-9400-4786-81bd-7c484317d5cc · outbound

This paper cites Long Context Transfer from Language to Vision.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.814956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.814956Z digest=sha256:72eca2edb8b9c06f7bf3f7918090597bf66713bfa6a5b9ea7cea2d953937dce0

Observation 94c136af-0e49-4990-a1f1-2c94e53f5172 · outbound

This paper cites Vila: On pre-training for visual language models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vila: On pre-training for visual language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.999749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.999749Z digest=sha256:66b0ff9519f9d93f6da40f4e8c43c09de59fc66795642364730c64df0b39449d

Observation 847d3506-46b5-4ed4-996a-c47b8e2c793e · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.181614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.181614Z digest=sha256:730ce3310619d25307b28215a27210467418a94e9145610a13d63d64c3e60602

Observation cd5f36df-dc5c-4131-8cbc-28b81e2c7ef5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.284629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.284629Z digest=sha256:5d4b24a29b05ccb878685c38c69e5cb8d077c18cf9d8167fee3d9942c05d87a7

Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.373080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.373080Z digest=sha256:7fd847ee42b369abcd0acad821c43c41c8826d36b9ea3a30d3902a64a3e26085

Observation a6198fb0-460d-4309-8991-d026ef49699f · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.464542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.464542Z digest=sha256:79cb299b836ae97ecb768069de7546afaf1c5b00448c655354d2aab6ce278388

Observation 7e2d0813-3723-406c-9f37-99613b9ef861 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.556314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.556314Z digest=sha256:022a3bc7e4d7bf015ffc4b7a9459235cc952f14344844e7ce08bdb8e83e084ba

Observation cfb50d06-db4a-4122-b1a9-f8f7b1410dc9 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.656075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.656075Z digest=sha256:eac172d325e3797658220881ea144c5f56b51cf2e811de3d62f62e2a0f5945d1

Observation c93b68ad-4f32-4fad-90d7-69d5909b854a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.725055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.725055Z digest=sha256:31e8eb570c52941201585df4203ab1c071e97b60afd065ef12951e4abe602971

Pith citing papers

Observation 98c1c08e-f4f6-4914-8582-a873fff20c39 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.261146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.261146Z digest=sha256:58f1caa27bc79c57f84d7bcafffc5a450dddfe1795a833c207fca36fb7600c98

Observation ac8c59c9-589e-4ac1-9487-00d555b75867 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:03.419373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:03.419373Z digest=sha256:1ca37a0d3ea8a4cf6551b8d0884967a3e6e732c4da27355bb9ec061a1e7916eb

Observation ea14883e-98a6-4fb9-9ef4-a8371292709f · inbound

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation cites this paper.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T00:15:23.806078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:15:23.806078Z digest=sha256:8d95cfea458769c1879531a04dc1663ddc4238630cb11b17b125eeee7a82c111

Observation ba0bfcdd-be2f-4489-b396-47970bb67ef4 · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.673718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:c2385eeb7e0db655404f9ea0eabebf4321d61cdd080e9273bb9f1aca818d75db

Observation e9c14120-8beb-4b12-8b0b-274c4bf4b9d0 · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:27:51.300574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:ec1388f6794be047558fc5922ae68a8d4479811d8714472132d03d7c699d6f7f

Observation f3429f25-bb38-4d17-b2ff-70b78648a36c · inbound

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization cites this paper.

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.752486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:53:02.051453Z digest=sha256:bf40e76b3c113acf4d702098f76ecc17b2085f856e653c614c3377017d7d104e

Observation 9b56a2e3-91fe-42c4-a10f-3c2b1b66a312 · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.092273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:cb4335df337412886777da2e49d8b40cc45f3e54a84912a5e42171853ab6326f

Observation 14e8d7db-96f7-45bf-99ff-3b55098fb9db · inbound

Touch-R1: Reinforcing Touch Reasoning in MLLMs cites this paper.

Touch-R1: Reinforcing Touch Reasoning in MLLMs GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.639487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:55:05.919049Z digest=sha256:917b0cd01f0b2080cb927d5b98a47f27058e50afa464d0daa95ac9f75ca481f3

Observation 7dbb6506-44d7-4151-a07d-775cf052e55e · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.320768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:6bd1ed3e114194b35132b85e59fa0339eae2c5780887f7da52b1f5525866d669

Observation 1ef5e2ec-7c53-4b66-8092-847994b7ba01 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:44:36.816892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:521506a16fe88e359a5f748d1754de31e70ba18b079fe6d1b7f39711c667245c

Observation d4b025b0-8a22-45d5-aae7-e2db5fb881f0 · inbound

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection cites this paper.

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.252360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T20:58:30.855338Z digest=sha256:8f4c1a780f93c0757b18701f30eba0855fd312aabbb8608518fdec047944f28a

Observation d83f06a3-2c79-4a41-b0bd-fc3c6771115c · inbound

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning cites this paper.

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:41.169558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T05:55:19.083517Z digest=sha256:f274bd5d384ea7b75847f25c1bb3648be90a6c55689a5f38484f5d0e5b327281

Observation ad9f6d09-1dc7-4dd0-9c3c-8b7df3220910 · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:d5b8ef9b2e4b71681b6ca3728040323ea53d59b2d6e042417c5178fe3a27f374