Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T19:47:48.545820Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2510.21583.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T19:47:48.545820Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T20:38:28.328436Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T06:39:38.334217Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1498f7d0-76ab-40db-a34f-6b65a354e068 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27003a4a-ebf9-49b7-8eff-585d8faee425 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a03c794-795c-4347-9dc4-4bcd9f75bf65 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d19ea9a-fc35-4144-92dd-8c6d4e76cb84 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc55de41-fc45-41ee-aaea-bc17659fc0b9 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a334ff87-f36a-48ed-88b5-8e58dd71a754 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1df560dc-494d-4e6d-91f2-299a60fda3f0 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Flow-GRPO: Training Flow Matching Models via Online RL
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2677e3f0-2f34-4f8a-ba04-f4c16eed30f8 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization HPSv3: Towards Wide-Spectrum Human Preference Score
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1265ee82-1043-4370-9bce-1c21f002e641 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5cf6a0ac-d990-47de-b3cb-753b6f1aea9d · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Proximal Policy Optimization Algorithms
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 24cb4f52-dabc-44c7-929f-868e2f7e2cde · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd645fd4-0f70-4601-a2e1-c52f3e27111f · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9161386b-c4f9-478d-afcf-2bc15d5c4f0c · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 30aa5399-ec83-4c02-9f71-e3ac89260494 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 767d40ab-f748-4528-9391-6c07f5686b69 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84c58d68-1553-4fc8-be1f-25082105195e · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Qwen-Image Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e86476be-345b-4391-96ca-9f607b5a9488 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 681582c1-61a5-474e-840c-bc1aad5d9923 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization DanceGRPO: Unleashing GRPO on Visual Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d40b59ed-b513-4e56-957b-0574b8368374 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf2a5567-0ec2-4da8-81ab-b5f980e92265 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Group Sequence Policy Optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2083c3d1-4b32-4c3c-afba-ec9404300a98 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 765f4c4f-1728-4acf-8fdc-ecbd5668cff3 · outbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization (2015; 2017)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 245acd5f-38ec-4881-b65f-4f557a80cd45 · inbound
HP-Edit: A Human-Preference Post-Training Framework for Image Editing Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3be852d5-5e50-4965-9537-90c03b6a0132 · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 416e3438-4e53-4013-85f0-d1c8b83e3f3d · inbound
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation abf3bc47-381f-44e8-8d90-401f9d2bbbf2 · inbound
Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.