Pith. sign in

Paper Citation Record · LEDGER

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03872 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:46:28.386276Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6a5e125-819a-4c7c-81a0-a85a03ff8c17 · outbound

This paper cites Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.464484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:24.994176Z digest=sha256:f8e246a64e0c5dd3528eeaa3b3ab0d34a612d71bf0888158d6422a1d4986e1bb

Observation 3dbb61d6-3dbf-49e8-a7ff-f65676dc94be · outbound

This paper cites SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.455706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.057055Z digest=sha256:9c8b77044b4aa5d7d35c334ec174d8303d6b183b3cc8f01b32b8b08ae5635231

Observation 07bbcc99-6cde-46cd-8fe9-77c30a925837 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.447346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.109193Z digest=sha256:8bdc573f33caf69136be8e56b8250a1d98d2ff5416c08b8fe80ccab3b5194a4b

Observation 7970e09c-8226-4a0b-84c3-cc097a6ca283 · outbound

This paper cites Reinforcement Learning in Robotics: A Survey,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning in Robotics: A Survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.438802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.182950Z digest=sha256:de9c984fb3291e7de060f673921910c3021cc6f6dd73d91ee9dcf0ab646215b1

Observation ee003f70-77fd-41ec-9553-583f223dc2f9 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.430466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.255645Z digest=sha256:c8291ce4fde819dcee773f1c04897a1bdbc0ce8f0114f7abe7975324f75f3ae0

Observation b3096895-cdc3-49b6-a3f4-ac879ef16933 · outbound

This paper cites How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.421970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.360263Z digest=sha256:806adc7a4a869ce9d9688c9b4f484202a37921f79ba795d317633aa504536ae4

Observation 00b39d8e-ce23-45e3-bc7e-87952e9e97d3 · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Addressing Function Approximation Error in Actor-Critic Methods,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.413050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.452881Z digest=sha256:0dd36292311a39265e01cd3365a827e9824d3e66f6ee07957d3ad56534bdb622

Observation 40e94431-dc38-4d66-b60b-a0f90c99a8b7 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.404067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.495996Z digest=sha256:b01432d769b545239982614b3fbb57fae244df38adcd30de5ad682af1125ebb9

Observation 2a819dc0-b74a-4df3-8ca1-d924d20837c4 · outbound

This paper cites A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.395749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.561058Z digest=sha256:ccf91d2f4f2df0dd5778b22efd6ca1e424b56ddfebc6ae3f394671bf236d5322

Observation 34b44d0c-441e-4c62-8145-4a128c3a3356 · outbound

This paper cites HG-DAgger: Interactive Imitation Learning with Human Experts,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning HG-DAgger: Interactive Imitation Learning with Human Experts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.386727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.622269Z digest=sha256:cbbaa8ce11d5678464c5b122cb138308a23877d73254cfed8d12a50a0c79ab84

Observation f17c478f-2b73-4db6-9063-41c732a28843 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Reinforcement Learning from Human Preferences,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.378305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.690310Z digest=sha256:ae00cf2fabd400f18dc0b54309b50049a13cc142063c49ae97d16fd9455491e1

Observation 4a128a76-9719-4ecd-acc5-a861f86a7e73 · outbound

This paper cites Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.369472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.781019Z digest=sha256:689404b6bc42e9fe9303e35d515b39515c847a9a74d86c59b60def29a2ee3f4f

Observation 0483ecda-299b-411f-be2a-7c5315a3f4ac · outbound

This paper cites Positive-Unlabeled Reward Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Positive-Unlabeled Reward Learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.360738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.865527Z digest=sha256:0722f6e84be31e0aa1e47cd35d4242878cc8f15c433e92d4bca16c97c474348e

Observation 0cb43490-8acb-4247-9fd0-dd6f41a666e9 · outbound

This paper cites Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.350941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:25.947347Z digest=sha256:580b86697c9949086cc6863b58aa05a906f753be269c9f46dedcd92e145f8d5d

Observation 69bc05e6-d327-4c3b-9aed-b739e9514e80 · outbound

This paper cites On Calibration of Modern Neural Networks,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning On Calibration of Modern Neural Networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.194875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.034319Z digest=sha256:8e1eaf9719fe4228a71a22663aada0ffcab3461796753e09ce2cfe1854f1af3a

Observation 723355f4-b2c4-4eb5-b93a-4753b1fbb54b · outbound

This paper cites Learning from Imbalanced Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning from Imbalanced Data,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.989060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.110725Z digest=sha256:17b9e3fff67da5b33e890e140a159bf42ec9c091bf17c9ff1a6671b0540233b1

Observation 2ac28726-310c-4896-acc0-01e4ed9f700b · outbound

This paper cites A Survey on Concept Drift Adaptation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Survey on Concept Drift Adaptation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.758765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.162134Z digest=sha256:4f2bd2e7cf1a63ec42ca980f5c585c6d7f412d49291227f5711049d082898e2d

Observation 59d41c21-fca4-407a-857b-bdf9634e7584 · outbound

This paper cites Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.454474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.229088Z digest=sha256:fe4d21f4751aefb23554b35294a7e04cd54603865512cdad7517a2391f1d1178

Observation d5cb31de-eda3-46a4-bd7d-c5bb94bd9ddd · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.134034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.255218Z digest=sha256:132e2dfa80f0820d835e0bf2de6e8d843b19ee6dfce0a73a87e73d4cda8ac50d

Observation a22c2192-035a-45e7-90ae-1e414fee8020 · outbound

This paper cites Implicit Behavioral Cloning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Implicit Behavioral Cloning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.028914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.342791Z digest=sha256:1775d3b1c964a57492508c266dab9603ffbfb5f541038dbb2b3fef1036e259b7

Observation 32dbd220-c69c-442c-be16-d6c59c2bd1f1 · outbound

This paper cites Flow Matching for Generative Modeling,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching for Generative Modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.835593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.396151Z digest=sha256:af4c5f747593e2479322b8ef8b37528365ae6a6f28328f63c94bbab332bd1bd7

Observation 9a1dcbfa-d030-4544-b0af-da1932723aba · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.713808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.458504Z digest=sha256:8bd9dd086a41b654c3435e79d3b5a0ad34d961975a4d1780319ad06e338500ee

Observation dbb56229-7bfe-4981-8084-ad4297a29541 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.580894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.561614Z digest=sha256:fb851115248dabdfb78f5efbd4f1bd7f83c4e5394bfd9027a4323251b658b5c8

Observation 6d19b647-08af-4fa0-a909-304ec2ad1165 · outbound

This paper cites Reinforcement Learning with Action Chunking.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Action Chunking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:26.613179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:26.613179Z digest=sha256:421f6582e13237e6140ce671735c7a1a8ca4e8c08d84489848900f088ede7f8b

Observation 83299f6f-d6a1-4a29-8389-38c1d2a70f88 · outbound

This paper cites Diffusion Policy Policy Optimization,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy Policy Optimization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.513155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.667042Z digest=sha256:2be94445316b6ab14b5215fcd78b1a0cd02f5fd38d850bc51e8fd612d1c0aa3c

Observation d2f9b537-f368-497f-969b-1ce7105477e0 · outbound

This paper cites ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.409585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.770861Z digest=sha256:9839bdab273aac2bed12f5a37be71b7cfb16ac43dc07910059f18979201b3ec5

Observation 6923b95e-7b69-446d-afe5-258403ac8ab1 · outbound

This paper cites Flow Matching Policy Gradients,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching Policy Gradients,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.323687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:26.927731Z digest=sha256:615143f8260f167b8c54e74186d589dd02ec6c36821dd1fbbf281d6d9312a6a2

Observation 6cd7f075-b8fa-4009-92b8-47037a1cc47d · outbound

This paper cites SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.197362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.068200Z digest=sha256:7842171d2b9898423cc3e66250c64193efded07912dd0b0b3a5893a40e412d96

Observation 2517fc5d-0c76-4dc7-a093-946cd622f38c · outbound

This paper cites Reinforcement Learning with Augmented Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Augmented Data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.994478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.111855Z digest=sha256:f617ee18070ea076593e479e9af5a04c7a4385b1c5c7c68582092f307224858c

Observation e7cf9319-7bf0-4076-bc40-bdd9ec4bd6f4 · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.825849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.230961Z digest=sha256:e967ea3f18d23c3bab325ab2d402b11112f7be2b5efc104222e640bbb41c5fa1

Observation ccc354a0-1743-4129-8397-473e10f00ace · outbound

This paper cites Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.674788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.392105Z digest=sha256:305ae02b45edcd83beaf5b36d96793fd1d6ea7ff4dc1ad9828cae857b5d59857

Observation bc970c47-367d-4ff0-a7e2-d1f5bbaa39a3 · outbound

This paper cites AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.554995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.549821Z digest=sha256:b035903d4573a9a2ed327bd26b86fefe012e5c3453109f3515df06cca6a39850

Observation f97f4f2d-835f-4c2c-a10a-38183256577b · outbound

This paper cites Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.420052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.685656Z digest=sha256:3684f6a31bcfe405dc693af0ee3c8419bc41a72d0a4b983dab1cd62101e91944

Observation dc7d9fb0-6292-4c23-88a6-5756ed7bfc09 · outbound

This paper cites Experience Replay for Continual Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Experience Replay for Continual Learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.280319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.797275Z digest=sha256:a6a6e80f7f488abac3a0900c11738a985a998b1243e37afa57692e34c08ac921

Observation 474f25e8-3476-42bd-b255-702e7abdd32d · outbound

This paper cites Regularizing Action Policies for Smooth Control with Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Regularizing Action Policies for Smooth Control with Reinforcement Learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.166038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:27.907550Z digest=sha256:fd5a1dca77f09e368255662408a0743372796953d3958a13249c980c3cdb76b7

Observation e85a1045-5005-4159-8b4c-1f52b905c0ff · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.009607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:28.019240Z digest=sha256:e0655f4260d7fa80d143892760980f63626a9c3622db4cd2ded9f695bd9ffd0f

Observation 7c8b414a-ce74-4b4e-82e7-1ad220966588 · outbound

This paper cites Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.840976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:28.152810Z digest=sha256:a86eebd2d1a0bb83d43e7685e499dd7f24292bde84c9f166650cfd3f3027f507

Observation a2cf5f71-6aba-4db1-acfc-d02d664a6107 · outbound

This paper cites Imitation Bootstrapped Rein- forcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Imitation Bootstrapped Rein- forcement Learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.686166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:28.219528Z digest=sha256:d6c94d1c4dcccf4cd48c6726901fa23e3a1187b5b4e2f3ef49d4e46b41f43c5b

Observation 8d219ba2-7854-4e8b-bbaf-96a92aaa9b08 · outbound

This paper cites Deep Residual Learning for Image Recognition,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Residual Learning for Image Recognition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.536036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T10:46:28.301849Z digest=sha256:f593faca2ee2d2c45ced938799030a2ff17590f4ef62075c50ea5164dd0c8a51

Observation e6f0ca36-6b27-4b5b-9b06-5f9778184faf · outbound

This paper cites UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:28.386276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:28.386276Z digest=sha256:3b3e423b3a2487fe53d6e7e0dab9a9ec5830160acb1807a8bd8c76849e052230

Pith citing papers

No inbound Pith citation observations are available.