Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:36:08.401563Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 5 inbound Pith citation observations for arXiv:2511.19356.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:36:08.401563Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T04:29:45.017327Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:23:15.759179Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7a26b75e-ccdb-4541-94ad-957a160d7f01 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fbcfb2d-1706-4683-aed3-2eb00a406eee · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ffd453f-6682-4f08-b348-8ab98838e177 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Reinforcement learning in continuous time and space.Neural computation, 12(1):219–245, 2000
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c372b29d-e85b-44d5-88db-1a97d8c7edca · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c870be-03be-4cf7-884c-a0d416dd60bc · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets You only look at one sequence: Rethinking transformer in vision through object detection.Advances in Neural Information Processing Systems, 34:26183–26197, 2021
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b4da14-7099-4ac6-a8d1-ec0faacc4d3a · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f136264-316f-4b35-bea6-c297c3a28190 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Hierar- chical process memory: memory as an integral component of information processing.Trends in cognitive sciences, 19 (6):304–313, 2015
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed3e48a-db57-4fd3-b1a0-0585347694f0 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9dcbf46-a30c-4650-906c-97ad8217fc24 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Non-negative matrix factorization with sparseness constraints.Journal of machine learning re- search, 5(Nov):1457–1469, 2004
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b19eb94-d592-44b4-bf51-c4ef1596d038 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets A reinforcement learning-based automatic video editing method using pre-trained vision-language model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb29fbe3-b2a9-4ce3-beb3-4f54649712ed · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Vbench: Comprehensive bench- mark suite for video generative models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe8ecbf-bf06-499f-8aac-c389af856352 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets istock video dataset.https : / / www
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe3484d-e872-494c-8a88-fb3fa698eb21 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf740fa4-ed8d-4eb3-b7d0-30c45294ce47 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Early language acquisition: cracking the speech code.Nature reviews neuroscience, 5(11):831–843,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad24454-862c-4482-8ccf-81fa676d4ba9 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db676d6-2ab4-4175-bdba-39f7c6d3850f · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf2ad4b-590f-401f-978b-c3089fe149a8 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Spatial-then-temporal self-supervised learning for video correspondence
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcce184f-2472-4bae-9d17-de9bcbb1404a · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Vila: On pre-training for vi- sual language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec0061d-2857-4c57-aad3-9bb6de700600 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Flow-GRPO: Training Flow Matching Models via Online RL
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154dd822-2a02-43a7-99a6-6f69ec2495f7 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Improving Video Generation with Human Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cf8f20-695f-42bc-84dc-dc927d85c6f1 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets When the future becomes the past: Taming temporal correspondence for self-supervised video represen- tation learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f458605-b810-4ace-a867-b7761e23c5cc · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Enhance-A-Video: Better Generated Video for Free
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c0482e-3c03-40c8-b80a-e1abebd19b69 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Liv: Language-image represen- tations and rewards for robotic control
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1d7051-8449-4d88-8b13-ffdd7cd19a90 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Video Diffusion Alignment via Reward Gradients
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3afd3ad-b9b6-4d74-bbbe-462bae08cb7a · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Learning transferable visual models from natural language supervi- sion
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e65b5bb0-a5c6-4c4b-bef3-69403e135409 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdedb479-18c1-4fa4-b884-2e3eabbf9305 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ec4549-e869-4fa7-860b-df6749cf4024 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Predictive reward signal of dopamine neu- rons.Journal of neurophysiology, 1998
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7478f871-005d-4492-8f96-b64fe333e03d · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09bffd77-3d7a-4723-9087-5940c2d7594a · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7f33f3-eb8b-42b9-8c58-d3ba12d758c5 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Wan: Open and Advanced Large-Scale Video Generative Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1148c233-707d-42d5-8580-ef6b65130e5e · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da874b86-f657-4855-9bd6-5e66fe2fced6 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b698f3-cd13-49ca-985c-5b610c9a5b47 · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets DanceGRPO: Unleashing GRPO on Visual Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5292d51-25b2-4bed-b186-ec7a93c6829b · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets Self-rewarding large vision-language models for opti- mizing prompts in text-to-image generation.arXiv preprint arXiv:2505.16763, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2800b94a-0457-477b-b420-fa151df9509f · outbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets {text_prompt}
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1482365-50cf-48cd-a368-7d3c056db61a · inbound
Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db3f62fc-c43c-45e2-9240-efef6d48f3fe · inbound
Video Models Can Reason with Verifiable Rewards Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f6f0326-c438-4795-a64c-3eb17e0c1d8d · inbound
CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a91ca6b-51ee-4564-8968-651b80867ef4 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a221d75-1be6-4883-ac0e-b10d67f04d23 · inbound
Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.