Pith. sign in

Paper Citation Record · LEDGER

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 13 inbound Pith citation observations for arXiv:2506.16141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16141 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:48:37.725055Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:35:11.261146Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1f1d5302-beb0-423c-96f4-79e6943e7596 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:32.974881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:32.974881Z digest=sha256:6728d5db3021a0ac9fff6d2d53adb1a5c6b5185f8dc2c5d72557f369694d850b

Observation 24167e82-b909-4652-9040-1d7369cbefcb · outbound

This paper cites an unresolved cited work.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:48:41.128978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:33.018546Z digest=sha256:ea7f1a234cdc0b99918e2c9ec255709b3369cc2c16305cd93e2bea8e579271d5

Observation 96bc01f3-3303-4e46-94cd-badd1d66a95a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.083995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.083995Z digest=sha256:95313ec08d75b7c9d47ac281de3ef71335bbb487a09a4f12210630537b0c16be

Observation 07a8d6fe-123a-4013-9744-9ef7209e42e7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.164335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.164335Z digest=sha256:9332313932acd9a1a9df228de1e0b23d6841920f02d3ed0ed2ce00f51aed6c2f

Observation 7aedcf7f-19ce-4c21-afc1-14879ba2e61c · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.223889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.223889Z digest=sha256:84f246eb7b58c98ce7224657b37bb1ea89a55e6ea5236c4037a5174aa698b0c6

Observation 87668594-5f01-4f8b-87d3-d9f12a73318e · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.343894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.343894Z digest=sha256:8402640920f403ac285bf6458bc855321c41d938b0510bed70f2b1d5213fcd9c

Observation 7ba8fcbb-2e6b-49cc-a7f8-6b01284932df · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.451030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.451030Z digest=sha256:f1584a6caa898304c1b0863a8474b947a54b1056122fb1712492f46bae6d3e92

Observation f1dcbb12-835c-47db-b815-97acdf0acfd0 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.585809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.585809Z digest=sha256:a03d1e1f9600a193c153d162cd837c82ffa5f48db3b5da1cd0fbbbc30956ba0e

Observation 4d460403-1db4-4434-a36a-5a883226b3fc · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.702470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.702470Z digest=sha256:1298e1deed6a59ac7bda8ea5953e38ccd7dc8056ca6d6be0ae2e7c2634dd4393

Observation 01639783-15c5-441b-99d6-50825f3a2196 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.790717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.790717Z digest=sha256:c8ffa7f90bac81cc8826c04010c3374e46dddd7d9a511bbcc214618fb329de80

Observation 3e860905-a7be-4343-940b-cca4e68e3f50 · outbound

This paper cites Video-r1: Reinforcing video reasoning in mllms, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-r1: Reinforcing video reasoning in mllms, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.869684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:33.849002Z digest=sha256:635486b180658e7d4df2ed3843e56cbb6ff6c32be68fcc07550d4bcd9b52d3da

Observation 1b678cb0-2ede-464c-a21e-6c7e4010fef8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.921347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.921347Z digest=sha256:ac6d1a44d7f441d6903a60813431da7dc02bb2048e8ee271603a4d87d4dc97e8

Observation fed62ce0-a034-485a-bf36-63085f2fbbd6 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.993024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.993024Z digest=sha256:cf1cde1a4ce905a8019d7e84fa50ec96e1dff1b870383729f9c3d0ad6d55bcd6

Observation caef7470-ac90-4e9c-8f9c-ee95ebbe62f4 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.652360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:34.042645Z digest=sha256:23c591a219cff3bd5a8991f0a2d40afc47d7c54b0027dcfd629b575bf35054e2

Observation 5ff65fe0-73c8-4310-8e5b-87392040dbc7 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.135877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.135877Z digest=sha256:a58909f99b985051bf07900d05088e6721e1ba484794fd539cf1d2b479871aff

Observation f9c40393-dab8-4541-82f8-5cca8d1246b1 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Proximal policy optimization algorithms, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.372247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:34.212172Z digest=sha256:d557b1688720f8c2346dc18784002c6f961f6976a24c4ac43d62c6eb7bd8bcb5

Observation ba36716a-bc1f-4057-bbe6-55210a24beae · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.320914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.320914Z digest=sha256:902d8ef48b07bd5bba37e96b01665c41794d4d27de3c0e2c62bd9a35b6c8ba64

Observation e55eb96e-5451-4977-8298-c7ca82c2048c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.438236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.438236Z digest=sha256:980120a21c0f4bbf5c8960e96b891daa071327390fa9147f0dc781a136ea1551

Observation 55dc7406-8b3a-440b-b8c9-44bd12456a59 · outbound

This paper cites Let’s verify step by step, 2023.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Let’s verify step by step, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.522538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.522538Z digest=sha256:f76cb155e616af5379a908b86f0f41dabdfab0854553cf84ef38fe8fb618d7d3

Observation 0924cf59-4514-4f59-8d31-deddba17e0f4 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Solving math word problems with process- and outcome-based feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.603403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.603403Z digest=sha256:00a213d9de605c6937e7772018c7abc647f7f08ad416a5c03a4c083e8459b962

Observation 703c21e7-fa80-4e50-8baa-1a99f975d6ab · outbound

This paper cites Alphamath almost zero: Process supervision without process.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Alphamath almost zero: Process supervision without process

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.165451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:34.660272Z digest=sha256:cfc57516ae73655049451d9afdb9a94d0f5e7fc184d31638108e1722a0c23c3b

Observation 210ad232-6d95-4711-a1dd-85c5fe2758f7 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.742575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.742575Z digest=sha256:6f0421b0ca2593d2ba4098b0cc5c3069055d712437747fe3be94b7771d891f3f

Observation 0feb8827-418f-418e-8da2-a7028c875bb9 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.838198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.838198Z digest=sha256:601cdb4053fe5a2566ee8b2ba58cd5171a3b693fe169ba499003c538ac861699

Observation 3e166373-3cab-4d7d-93c9-1559829a576c · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.894129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.894129Z digest=sha256:fd6c014429e373b42847365b5f385a63a57d3c535405e445c754e4c9ec1c0603

Observation 4249718d-264d-47b1-9722-2ef4abd5d1ae · outbound

This paper cites Evaluating mathematical reasoning beyond accuracy.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Evaluating mathematical reasoning beyond accuracy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.929136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:34.955288Z digest=sha256:ebe9911c1089be0dd6384c441c7f1953310f5dc1595ec3a1efc7a03da87eba36

Observation 97e55492-936f-4c0e-a906-accf45d41bb2 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.066396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.066396Z digest=sha256:af80af4585f639d62d2fff65b6fe385516de70f9423478b988a84ec2dcb29121

Observation 82a53331-fc24-45c0-93c9-e2aad67d13f5 · outbound

This paper cites Warp: On the benefits of weight averaged rewarded policies, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Warp: On the benefits of weight averaged rewarded policies, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.715436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:35.168865Z digest=sha256:2f5fca82feb526f6636ce2838ac3fd958af60ef26e0a8ee8a42e3b00310d09dd

Observation 3a7d7d8c-b443-443e-a03b-13792fd8bef1 · outbound

This paper cites Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.276903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.276903Z digest=sha256:1cca10f90443300db6a33736933d98c302f7ebd53fcc0ba4b3b31db6621afd91

Observation c747885d-dd23-4ad6-9ada-f9fdd4f839b3 · outbound

This paper cites Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.435084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:35.459609Z digest=sha256:6bf702822e4589048a7b57e23c33f78e3b97eb14ca4495b00bf33b26794efdf0

Observation 5527d3f9-a4c7-4a69-9141-87eaf4ece1dc · outbound

This paper cites Open-r1-video.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Open-r1-video

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.241098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:35.524033Z digest=sha256:f0967757853c9584aabd58ad1659b7b46841c94c8e0639b9363ba9f81bd9bd60

Observation f4b54ff6-e282-4069-af2e-5c4dee1f3d62 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.612776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.612776Z digest=sha256:be0c4c83a4134aa4762e072946579dc1ff0c8a0b90e680805b4de0e02a6ed789

Observation 4042c7ab-253b-407c-887f-31b9ae9694a2 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.023697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:35.679709Z digest=sha256:f0515c84f4620057cf2671d9559570c4ce014f57aac93d782cc08a99bc10cbc2

Observation e7c72394-096f-4d90-8344-71c77b1bb301 · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dfew: A large-scale database for recognizing dynamic facial expressions in the wild

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:38.821362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:35.767077Z digest=sha256:a0cf71410c793c4dfe1e3b719a9348148a8b2554f52876623da1b176ded8c681

Observation a0df6ca4-fdf4-447a-850d-7e9688d04544 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.861496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.861496Z digest=sha256:d2d83c7277602890ec52e98fe93ef9ec0ee82c6d61c98b677f3ba2eccd541a3e

Observation 22a0426f-ac6a-4cdc-933e-c641f0d5a4d0 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.947030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.947030Z digest=sha256:5c1da4ae8ba7f8eb1e2a1ac9131bda777369cb0323b950f6f48988fb1996d303

Observation db357cd1-ea01-459b-945c-e6a9cfcee43c · outbound

This paper cites Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:38.593156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:48:36.039772Z digest=sha256:e6713ad0aebf5f31b4d12064f954511c1251c628958eb0ee69c0bd5f31808cc9

Observation b6b4a07a-446e-465a-b0b2-fc16fed427ae · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.245848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.245848Z digest=sha256:1ff767abf10353bdb4dce4817757cf9b6d6699a665ab4c31060f0be4e6b643bb

Observation 8120ea48-5a60-4ed9-b67c-264cbe53d991 · outbound

This paper cites Qwen2.5-VL Technical Report.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Qwen2.5-VL Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.371774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.371774Z digest=sha256:91a887775dd28c1e42bc67a4f76f79cc50ea5ce4d285031a85c96cbdc065e430

Observation d1bee5ad-baba-454a-9bb7-b7384a0c64e1 · outbound

This paper cites GPT-4o System Card.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.504016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.504016Z digest=sha256:166832e10880787aec413915277565b0a8e0dd8cc06c495dc8696d6ae4c13913

Observation 52bd632d-3da6-4f60-8a1d-5a5eb892d476 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Llama-vid: An image is worth 2 tokens in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.600274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.600274Z digest=sha256:eb72858764d5dbc2cb910d31c90687156391ebed699c172e56809a7054119ca2

Observation 9fa98f0e-5053-4c37-abe0-bb1c03d95e62 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.680253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.680253Z digest=sha256:b3ad0d9636ac60efdf3785fb6c1b1d8d755aa9e18fa413b2d0d173f564e47f81

Observation 0b8f8ce4-9400-4786-81bd-7c484317d5cc · outbound

This paper cites Long Context Transfer from Language to Vision.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.814956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.814956Z digest=sha256:45f26f7696a0fbbc1f2fb75ab7e8572c7b4ceb143609a6203bdf898ae3715c16

Observation 94c136af-0e49-4990-a1f1-2c94e53f5172 · outbound

This paper cites Vila: On pre-training for visual language models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vila: On pre-training for visual language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.999749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.999749Z digest=sha256:e68374a674e9c55d7c7a5b133812e01876c67a01fb0cb1cbb56e968fd17fc60f

Observation 847d3506-46b5-4ed4-996a-c47b8e2c793e · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.181614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.181614Z digest=sha256:8b5f763db71b6d484b17607adc9a501ffdba1c1616c48fa8ea236b0b1790e207

Observation cd5f36df-dc5c-4131-8cbc-28b81e2c7ef5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.284629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.284629Z digest=sha256:a4474ef7e987e855eb82f5c52a06f07510a063376c2cdc2cc6ce6380125bbbba

Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.373080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.373080Z digest=sha256:3e982e8d211c2dc5353e7ba44360ea5b879ab37f2ab2082b63db2745b1e34d3d

Observation a6198fb0-460d-4309-8991-d026ef49699f · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.464542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.464542Z digest=sha256:016b371b265e64fc534adc7f5f143f6801869d13ed844e907b605d8a8055e60c

Observation 7e2d0813-3723-406c-9f37-99613b9ef861 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.556314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.556314Z digest=sha256:cc3187e64912a43583f0bf28644dfc101a76943253fa51d531aaaf13cdb82e78

Observation cfb50d06-db4a-4122-b1a9-f8f7b1410dc9 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.656075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.656075Z digest=sha256:20712ca94f55eeb07070afbbcafc7560ac4a61f51d4178bf20a0486e1ac966c5

Observation c93b68ad-4f32-4fad-90d7-69d5909b854a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.725055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.725055Z digest=sha256:37e8bceab437b1bd609bd65a132e190e46ec93c0c8f61fbae9e97976455c12a2

Pith citing papers

Observation 98c1c08e-f4f6-4914-8582-a873fff20c39 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.261146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.261146Z digest=sha256:5e1e67ba9dca63fa3d710d00ea87f1c3079da7843bb06f09b24e3fe0cac2c293

Observation ac8c59c9-589e-4ac1-9487-00d555b75867 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:03.419373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:03.419373Z digest=sha256:9f81185664f02bd0a265d74961f9d01a96def95f6708adc8262922e0b1da8ff8

Observation ea14883e-98a6-4fb9-9ef4-a8371292709f · inbound

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation cites this paper.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T00:15:23.806078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:15:23.806078Z digest=sha256:1da8eadb25b59bb089cc879c061d4374288ce7552b51475bdd2f7392a50259aa

Observation ba0bfcdd-be2f-4489-b396-47970bb67ef4 · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.673718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:5bcdf32eed647a7b00217004b0f3fa6b582bf5fcbbd8df60a46491fbaaca7ccc

Observation e9c14120-8beb-4b12-8b0b-274c4bf4b9d0 · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:27:51.300574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:46f0621dccb2f6e46e8978aa08b3c0f5727bb667f3086846e29ee9b84244430c

Observation f3429f25-bb38-4d17-b2ff-70b78648a36c · inbound

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization cites this paper.

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.752486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T08:53:02.051453Z digest=sha256:6dd94c47ad17a6925687e642fa9cf55cdb9450b069c615423522359ba1e288a8

Observation 9b56a2e3-91fe-42c4-a10f-3c2b1b66a312 · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.092273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:e7db60b8a0c1f7129228db50443e9800e66851b7ffdde5553e5e643211a8ac1b

Observation 14e8d7db-96f7-45bf-99ff-3b55098fb9db · inbound

Touch-R1: Reinforcing Touch Reasoning in MLLMs cites this paper.

Touch-R1: Reinforcing Touch Reasoning in MLLMs GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.639487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:55:05.919049Z digest=sha256:5af55b2ca52c7de19d65b32e8fbec9222d1572d78ca00022e5e9b65d99e2b1f3

Observation 7dbb6506-44d7-4151-a07d-775cf052e55e · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.320768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:59a39d2c34e423a2ee317db26def537822808c7cd74ad1db5135d57536216404

Observation 1ef5e2ec-7c53-4b66-8092-847994b7ba01 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:44:36.816892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:5f0face0bae84c9b64ad10f9555a91b60ab582b1bfe425e1e9c1bb0cb21382ff

Observation d4b025b0-8a22-45d5-aae7-e2db5fb881f0 · inbound

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection cites this paper.

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.252360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:58:30.855338Z digest=sha256:7e7c65c3e992b091cb5e5bc0661b5ca766568699e271dada622a90235b84379d

Observation d83f06a3-2c79-4a41-b0bd-fc3c6771115c · inbound

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning cites this paper.

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:41.169558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T05:55:19.083517Z digest=sha256:acc768df10ed1a862af9f5eab9589d62b733e440d2bef27901297a1de5bda254

Observation ad9f6d09-1dc7-4dd0-9c3c-8b7df3220910 · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:2297ca9998cffff7278d890847e497743895939140e2df686b29fc1cef8d40da