Pith. sign in

Paper Citation Record · LEDGER

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2510.21583.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.21583 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T19:47:48.545820Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:38:28.328436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T06:39:38.334217Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact16
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1498f7d0-76ab-40db-a34f-6b65a354e068 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T19:50:33.854785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:ec6b27b3bd6768bdf1d3fa9c67df744b61f26596fda2c61e9ff7bd64f5bf8ab6

Observation 27003a4a-ebf9-49b7-8eff-585d8faee425 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.864804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:9c52779546f2ed92f65dca652a5a5f27e76928d7a2bfd5ba44412afe75b29f7b

Observation 4a03c794-795c-4347-9dc4-4bcd9f75bf65 · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:52:31.656662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:ad3ce798a64960427736978e3ca53d1f6757b69bb489e48fc65b19ce1f30eb29

Observation 2d19ea9a-fc35-4144-92dd-8c6d4e76cb84 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.868682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:424ea06b4ec019539ffd2ea8ba4ed5a1b531d9150bf507d304eeb975afbe7e4e

Observation fc55de41-fc45-41ee-aaea-bc17659fc0b9 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T19:50:33.858347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:955b4b241b50454f537128fddae55da3beb09dbd15d8537aa298179028920351

Observation a334ff87-f36a-48ed-88b5-8e58dd71a754 · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.861376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:ae709ce611396124686a06edcc92dc482b331405ca847b1a8eb5e97f4135dd98

Observation 1df560dc-494d-4e6d-91f2-299a60fda3f0 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Flow-GRPO: Training Flow Matching Models via Online RL

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T19:50:33.803085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:c05952704640d0a433a43858a29aa45061ed07d43e11b0fb451f9c1f9f855073

Observation 2677e3f0-2f34-4f8a-ba04-f4c16eed30f8 · outbound

This paper cites HPSv3: Towards Wide-Spectrum Human Preference Score.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization HPSv3: Towards Wide-Spectrum Human Preference Score

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:50:33.819081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:c5d01e1c1b9e56811f18c809c7a9e6dbdcbb04b7b419fb444d079fbad8afe062

Observation 1265ee82-1043-4370-9bce-1c21f002e641 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.846959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:455b5c18b650bc8cae504a0895352cb5f4e7a8dc602a15bd5b824fe19b39dd62

Observation 5cf6a0ac-d990-47de-b3cb-753b6f1aea9d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Proximal Policy Optimization Algorithms

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.843772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:3009a8738ea7fbbc8d2e0d668b5f7ad2ff2d417aaf9b5cf8dd9edddca080d2d5

Observation 24cb4f52-dabc-44c7-929f-868e2f7e2cde · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.872470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:6b0be91267950b0c3e9f7b56900462ae9066f278985421aff84be2ce89d31a54

Observation dd645fd4-0f70-4601-a2e1-c52f3e27111f · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.836808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:8e108ae6f699bb5ef72b102321ef6bef21a3dcb3faafb265591c11c99aa1d637

Observation 9161386b-c4f9-478d-afcf-2bc15d5c4f0c · outbound

This paper cites Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T19:50:33.810618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:d401eac9b5f20db1436d983d2d80ad1a778ab34463c77a886fe12435bcddc469

Observation 30aa5399-ec83-4c02-9f71-e3ac89260494 · outbound

This paper cites Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:50:33.833248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:d71a140d324150ab91dc08073e3afef75a0f134c841168007da940358f570b74

Observation 767d40ab-f748-4528-9391-6c07f5686b69 · outbound

This paper cites Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.840331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:1dcc27a990ed9678b20806f5439d7f954747b5f92781036f7be41a544eb357cf

Observation 84c58d68-1553-4fc8-be1f-25082105195e · outbound

This paper cites Qwen-Image Technical Report.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Qwen-Image Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.822100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:f44ba24978a8e9c89190da7fff425006ab78cbf9efa005e1b0a62eb694c9a782

Observation e86476be-345b-4391-96ca-9f607b5a9488 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.829300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:a5e7b14845092d4484bc4653f291eec0f20dc069cd6262a83d9a5653f7fb639f

Observation 681582c1-61a5-474e-840c-bc1aad5d9923 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization DanceGRPO: Unleashing GRPO on Visual Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.806287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:147bf0df9c75e23d8317236f732d391bb4cab26731c287b9dbadf3620060e250

Observation d40b59ed-b513-4e56-957b-0574b8368374 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.851348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:56ca71fb847e88434f5f17f8fa2f6cb677ff9b93a310326f8b09436b57a5c862

Observation cf2a5567-0ec2-4da8-81ab-b5f980e92265 · outbound

This paper cites Group Sequence Policy Optimization.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Group Sequence Policy Optimization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.825525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:401a259e0761aacb42e86dcd2f874ecead8d9522110ed017199baeb7f7a48fc7

Observation 2083c3d1-4b32-4c3c-afba-ec9404300a98 · outbound

This paper cites an unresolved cited work.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-21T19:50:34.567791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:2dd1dd96dd0529e7c71709f8125dee9c20a83c53db9251866b33399728ca5d04

Observation 765f4c4f-1728-4acf-8fdc-ecbd5668cff3 · outbound

This paper cites (2015; 2017).

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization (2015; 2017)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:50:34.565627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:1ff4da32bb0228b8333033a5d75861a76ebbb8f71fd72c08fea1f05f30a1b325

Pith citing papers

Observation 245acd5f-38ec-4881-b65f-4f557a80cd45 · inbound

HP-Edit: A Human-Preference Post-Training Framework for Image Editing cites this paper.

HP-Edit: A Human-Preference Post-Training Framework for Image Editing Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:53.918218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:26:15.525307Z digest=sha256:866e9f0de694e4470574bb84ffba32a6a71f9c7937bf53e88d69f371d18eb172

Observation 3be852d5-5e50-4965-9537-90c03b6a0132 · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:53.918218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:364a77d381dae2d93ce9382d2e9aae3ea6e1e5bcb00ca49ae00ad4c378e88a65

Observation 416e3438-4e53-4013-85f0-d1c8b83e3f3d · inbound

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO cites this paper.

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:47.194096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T20:38:28.328436Z digest=sha256:f6730bbed51b05a241fdf2ef181fcf2138c98eb93f5abeb19a446add02ad2ad5

Observation abf3bc47-381f-44e8-8d90-401f9d2bbbf2 · inbound

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training cites this paper.

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:39:38.335537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T14:14:45.547957Z digest=sha256:ca77a16edfc62f4391deb2eec13df251c99a80ca8b94172609dd6c3cbfe5742c