Pith. sign in

Paper Citation Record · LEDGER

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

As of 11 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2502.07645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07645 v3

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T04:01:18.219745Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:33:10.067648Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T23:26:21.964199Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact7
  • verified fuzzy51
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 964cd3fd-39e8-40e1-9d66-6a0c487874bc · outbound

This paper cites Implicit behavioral cloning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Implicit behavioral cloning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.108867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:420bf56cd9db4efeb1f2dfe36d40dab702825501b06e82ef5827dcf543a128fe

Observation b5e9a2a9-18f6-4870-ae7e-ae3248b98a45 · outbound

This paper cites An algorithmic perspective on imitation learning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback An algorithmic perspective on imitation learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.090664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:e3e3cb86377f2194a7430c7c01dd906974defdda3e14fe7e353647c2f69c6764

Observation 7a2b61ae-21ad-462f-b6eb-5bb707a46b7b · outbound

This paper cites Re- cent advances in robot learning from demonstration.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Re- cent advances in robot learning from demonstration

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.112086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:3b480145978eec6d297e6f4c4d28b2d8200f350463f683f71e6bef867eee1ca6

Observation 6e52e327-5476-42bf-a208-0babcff3e854 · outbound

This paper cites A survey of imitation learning: Algorithms, recent developments, and challenges.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback A survey of imitation learning: Algorithms, recent developments, and challenges

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.122768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:62c1c447650c2846bdde8174f0deb6782e82988883d107948419d07dd9a52a91

Observation 4d8cd410-1bf0-4a88-b347-4ce66a16bb1c · outbound

This paper cites A survey of communicating robot learning during human-robot in- teraction.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback A survey of communicating robot learning during human-robot in- teraction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.141544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:cc0701b6d9f74621a53c3ecb0dd3dd159e4470d4aa905fc2c62de97164e032f7

Observation f0102c90-28b6-4948-9b31-15004c8c25a9 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Diffusion policy: Visuomotor policy learning via action diffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.098074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:06e19feee652c4caa5c02afadd7935475fdaa0e447ba3352cafe3bcb3b25287a

Observation eb8c8d1e-5f02-44e9-b020-6b4a09b48835 · outbound

This paper cites Conditional Energy-Based Models for Implicit Policies: The Gap between Theory and Practice.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Conditional Energy-Based Models for Implicit Policies: The Gap between Theory and Practice

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.298733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:4f8d6062481b6e65f4853ca988efeb0c3a35614d9d281bdbf3252c83c4a22bdc

Observation 243dd2ad-2b1a-4ea1-b287-2df0d0d59de0 · outbound

This paper cites Goal conditioned imitation learning using score-based diffusion policies.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Goal conditioned imitation learning using score-based diffusion policies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.071851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:567e468dde2eed3fe2b14fee563db8cbd896cdc4c24143c0b48db43c315eccbb

Observation fc40ee59-04d6-4b29-95e2-0fbb9556b9dd · outbound

This paper cites Fast and Robust Visuomotor Riemannian Flow Matching Policy.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Fast and Robust Visuomotor Riemannian Flow Matching Policy

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.293197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:4fd89429680fab69fc06231196fd3efbbecb681ac8ef8c49cdbcf38b6bfa6262

Observation 42588c92-d57d-4c93-ad6d-1b9097a961df · outbound

This paper cites Deep Generative Models in Robotics: A Survey on Learning from Multimodal Demonstrations.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Deep Generative Models in Robotics: A Survey on Learning from Multimodal Demonstrations

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.288216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:b6c378585bbd9915578e7a14fb128468d218f9b760c2f2cabe2e6d6902e92810

Observation e668243f-4b19-4a67-96b7-3299202612fe · outbound

This paper cites Interactive imitation learning in robotics: A survey.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Interactive imitation learning in robotics: A survey

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.066606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:da0bbf95945257fc5ab574e2c8794e0b44512077d1d2e7975d0ae65a5037fee1

Observation 6466329d-8bb4-4ef0-8309-3849e1621ecd · outbound

This paper cites Reinforcement learning of motor skills using policy search and human corrective advice.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Reinforcement learning of motor skills using policy search and human corrective advice

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.083820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:8b6668d315ba7e1d14423fa17776cdf4a9badf45bfd7f99cb876936f148a139c

Observation ebb5835e-3f67-46bc-9850-5a563261e347 · outbound

This paper cites Contin- uous control for high-dimensional state spaces: An interactive learning approach.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Contin- uous control for high-dimensional state spaces: An interactive learning approach

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.053331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:5a20237d0881e07efafa8f62677f6682cda211432c3387e42e3bca82b7d441dd

Observation be6b9431-5d04-47dc-9664-42a41742c599 · outbound

This paper cites An interactive framework for learning continuous actions policies based on corrective feedback.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback An interactive framework for learning continuous actions policies based on corrective feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.058003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:360e2e968e58e2c40d88b66d4e370430667dca74469f7716bcbc15e6507af333

Observation ecabbe74-4cd0-461a-8196-136718e51157 · outbound

This paper cites Implicit generation and modeling with energy based models.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Implicit generation and modeling with energy based models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.039406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:15ed4bd860971f141b811494f7f45eb56eccd8c00ac1daf0854dccee1896080c

Observation debec787-66d9-467d-8f01-8681f52cba43 · outbound

This paper cites Towards tight convex relaxations for contact- rich manipulation.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Towards tight convex relaxations for contact- rich manipulation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.030105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:4ad198c02134e9e99f6f306d6a1fd2d1fb380f8e61e214692327e57aa9f9899b

Observation c1112610-36d7-4848-9694-c19430886e79 · outbound

This paper cites How to Train Your Energy-Based Models.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback How to Train Your Energy-Based Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.282133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:ed05a5572775d6796bc2470560c3eb8a77ed2b3e999ac2952d7c9313e8e85656

Observation ddb54be8-88f9-4e3d-8291-1ec9026aea20 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Deep unsupervised learning using nonequilibrium thermodynamics

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.033861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:a2ea22a7d3ebb5a74b56d504814ca7eda3418c35e2a41655e3bc4637b2b04651

Observation 092d489a-dede-47e1-ae9d-c139c4ed2cb0 · outbound

This paper cites Denoising diffusion probabilistic models.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Denoising diffusion probabilistic models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.129412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:f7c600c200ee60c2177280b17faf0ed4783dce40ab943be8e7858748daa51969

Observation 79639212-d07a-478f-a7a7-752d95b8c62f · outbound

This paper cites Score-based generative modeling through stochastic differ- ential equations.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Score-based generative modeling through stochastic differ- ential equations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.075677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:ae6078794123eeffd45726a1617402399f7b385ad6709a9c89b3642ff043d328

Observation 05895d16-1754-4783-8f62-238ec16da166 · outbound

This paper cites Energy-based contact planning under uncertainty for robot air hockey.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Energy-based contact planning under uncertainty for robot air hockey

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.087160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:1a9d8b3e5caec5e508c80ceb6f544113cde4d1cd5ae8a834bc183a5c1f9b7f64

Observation c1e222f6-0be0-4d4a-b2fe-e7c0d0ccdaa5 · outbound

This paper cites Using im- plicit behavior cloning and dynamic movement primitive to facilitate reinforcement learning for robot motion planning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Using im- plicit behavior cloning and dynamic movement primitive to facilitate reinforcement learning for robot motion planning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.094345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:6aca7d300f3f7f51c8a009d8a305661a27ccb25712abeec55e2d9112ee3262de

Observation 99896726-cd49-4fc8-806e-7a39687baf45 · outbound

This paper cites Iifl: Implicit interactive fleet learning from heterogeneous human supervisors.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Iifl: Implicit interactive fleet learning from heterogeneous human supervisors

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.015597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:0743f8784da67950eb1c72fde529a02ef9652c61b7b35aa5a9325ec5f0619a8f

Observation 7e26d9bd-6fed-4c3b-aa50-f3694479ca96 · outbound

This paper cites Diff-DAgger: Uncertainty Estimation with Diffusion Policy for Robotic Manipulation.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Diff-DAgger: Uncertainty Estimation with Diffusion Policy for Robotic Manipulation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.276465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:aff46871cce038b1eaca599498aedd2fb717fea433f3c62b3899dc74d56e6a12

Observation 4f779a3f-4ebb-4480-b4fc-6a46e56b45b5 · outbound

This paper cites Deep reinforcement learning from human preferences.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Deep reinforcement learning from human preferences

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.985882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:573ccab4087723c4add99b6cbb6c0713a3d1b2fbb0fb2c4bb3c9188abaa1e7fd

Observation 5ad75fd2-b316-49c3-aff0-73893d665a30 · outbound

This paper cites Learning preferences for manipulation tasks from online coactive feedback.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Learning preferences for manipulation tasks from online coactive feedback

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.001162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:54fc77907ef39d620c786b27f54be846de722526862ad075d7e4a762b6abe7a7

Observation b3209018-ef3b-478d-a0cb-f2e5751e542c · outbound

This paper cites Pebble: Feedback-efficient interac- tive reinforcement learning via relabeling experience and unsupervised pre-training.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Pebble: Feedback-efficient interac- tive reinforcement learning via relabeling experience and unsupervised pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.959128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:2be48ff37509513bdf5f40ad4d97b93bce45f6f0ebc5e161db8de2d4c2c2262f

Observation a0ec34fc-fedf-4edd-8ba7-b99e7ecdb6da · outbound

This paper cites Learning to summarize with human feedback.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Learning to summarize with human feedback

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.978504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:4ed99c0d7f80bc70b3de9695729d0c7177ac5b149285ab2e70bd65197dcd7f1d

Observation 65cfb03f-9577-43b4-a7d9-3b67403c2e53 · outbound

This paper cites Trajectory improvement and reward learning from comparative language feedback.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Trajectory improvement and reward learning from comparative language feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.993744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:a7d7b23a676a783e48af302c92ca953b6fdcaae212d33db13d2adfa6b11749b7

Observation e222255c-11a0-47dc-bb8e-15530259d435 · outbound

This paper cites Contrastive preference learning: Learning from human feedback without reinforcement learning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Contrastive preference learning: Learning from human feedback without reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.012334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:453f98ea648cf563800540b0658c678c15907b02b034092dff5ddf229645bd91

Observation 9c9c8b3d-1676-40c6-a291-6ed798be7e88 · outbound

This paper cites Calibrating sequence likelihood improves conditional language generation.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Calibrating sequence likelihood improves conditional language generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.018976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:133ead1e0fe55b1f62d9793fb1bf9149973bd00d76d7bf027da5637b8c5157cd

Observation 73525543-999e-45bc-bc28-8bbc904b8cdf · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Direct preference optimization: Your language model is secretly a reward model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.126075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:c1192346ff979d7533b9d4e2ee9aa7d2dd9fc922e17252b94669cc49925139c6

Observation f2335496-a74b-43c1-81f4-d65f5c6190e9 · outbound

This paper cites Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.997487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:cb6c5429025f4d6822242ad6c9f43db6ab2633815a608f74df71a7b992856fcd

Observation a9c171dd-bafe-4655-91d9-343dd65282a9 · outbound

This paper cites Batch active learning of reward functions from human preferences.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Batch active learning of reward functions from human preferences

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.005152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:6f63e8f011ce52e59a5fcb85801a2fcdd5afc64942c0d1b6eee9976d28861f2b

Observation a06d8981-b0ff-42b5-b9b7-568aa68d52f4 · outbound

This paper cites Hindsight PRIORs for reward learning from human preferences.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Hindsight PRIORs for reward learning from human preferences

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.022529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:44940a33ba1e41de66bfdf6b1020d66578773cb0af4de16bed61047fa98b2c15

Observation da478840-5636-4894-bcc8-cd74dda27ed3 · outbound

This paper cites Learning robot objectives from physical human interaction.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Learning robot objectives from physical human interaction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.008831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:3a885400ec6e2ecf33aea1e115b86fce5ba0950134377b62fc4eeeaa93dfc6ba

Observation 64de3116-b100-45b6-96a0-503bedadf55a · outbound

This paper cites Including uncertainty when learning from human corrections.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Including uncertainty when learning from human corrections

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.982276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:c4d773a39113b4bfb07293513e079a8be98c7547c20ca18cdd6484f4834f48c7

Observation 74b451cf-77e4-4641-a2d3-98b5d88d4b35 · outbound

This paper cites Learning from human directional corrections.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Learning from human directional corrections

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.975038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:29b78d60c36778c75dc8728887f57b5eb3657cd4a87bf3611cc6420fe7e2ac32

Observation e17b6781-8eb2-4ab9-bb3a-92467feb6b4e · outbound

This paper cites Interactive learning with corrective feedback for policies based on deep neural networks.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Interactive learning with corrective feedback for policies based on deep neural networks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.119290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:6491289e3f1b686163a7539daec5201700142572a1b0d9212f2da359ce5658ca

Observation 98b8ad1d-1520-4a8c-b0f7-8caa720d7426 · outbound

This paper cites Towards corrective deep imitation learning in data intensive environments: Helping robots to learn faster by leveraging human knowledge.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Towards corrective deep imitation learning in data intensive environments: Helping robots to learn faster by leveraging human knowledge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.137589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:2ef0f40dcb6c263530c9f2c72147f083c4341cf88174a200b093800787e0cac6

Observation 0d2477d9-1e3f-4768-a24a-2503d7cee6d5 · outbound

This paper cites Interactive imitation learning in 18 state-space.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Interactive imitation learning in 18 state-space

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.080071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:b4cd0b04de27ea23af8a396574412b4228e9829fc0164e69b6bc090c8329eca3

Observation 9f81d2b3-d00b-4db0-90e2-507acbb91217 · outbound

This paper cites Learning from active human involvement through proxy value propagation.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Learning from active human involvement through proxy value propagation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.990001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:62e1b5716f407884c31782005dcd14eb7d2555e74e227dd4f301a5b60032b764

Observation aa954b4c-ffd8-40f6-a8d5-34e68021ccc2 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Reinforcement learning with deep energy-based policies

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.967232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:13bac57a74fe5c6c31775dcf9440eaa72d3da23ad611d33f7ef238450794db69

Observation 25705914-2d45-4510-aa97-8b2063ef7060 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.963113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:0d0f9907238aa64f14804d91498e28c6fa32bc4d27de2dcd75bb45508332d648

Observation 55c9025f-9c63-4bca-a578-251354d22685 · outbound

This paper cites Aligning human intent from imperfect demonstrations with confidence-based inverse soft-q learning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Aligning human intent from imperfect demonstrations with confidence-based inverse soft-q learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.062479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:c417a271f5f0f779478074101e73f11afa6b6918d7131f26236f094684815c7b

Observation 3d2a8f9a-4dff-4cf4-812c-6bf2d94286fc · outbound

This paper cites Bayesian reparameteri- zation of reward-conditioned reinforcement learning with energy-based models.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Bayesian reparameteri- zation of reward-conditioned reinforcement learning with energy-based models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.133042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:6d3809fa88331e4c01ba6e23010ca02a3e8ca6613374c1575e434dc3d504d2a4

Observation e5dbdc9a-fb2c-4090-9407-851c218d5b01 · outbound

This paper cites Inverse preference learning: Preference-based rl without a reward function.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Inverse preference learning: Preference-based rl without a reward function

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.115367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:1e81354e693859eed6caa564589b09197cf93e2a531e3f47e40a2dc9f5fdeaf3

Observation 4a03b651-0daa-4acb-80aa-be86c190394a · outbound

This paper cites Learning from interventions: Human-robot interaction as both explicit and implicit feedback.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Learning from interventions: Human-robot interaction as both explicit and implicit feedback

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.105462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:655d34c2e15ca0d79bd4c24cf5650a69d0fe4db4fd632577695325305fb593a3

Observation 1f037782-f950-4156-a6d3-9ed14b01aa8d · outbound

This paper cites Flow contrastive estimation of energy-based models.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Flow contrastive estimation of energy-based models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.971063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:3562dad2b042dba292c62546f14b80af04427dcfd1540cbc25febac19a377a71

Observation 770ffb4f-80b6-4a50-8f1c-a06bcd9a1c38 · outbound

This paper cites Hard negative mixing for contrastive learning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Hard negative mixing for contrastive learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.026134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:3d3c104327ce831fd227e98474e0ee034a2bd5bc6b458f0fb15d860a25f9bb66

Observation a738eff6-1346-4475-adc3-bae0ff835626 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Representation Learning with Contrastive Predictive Coding

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:02:30.270965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:6cbf3d92d7e5f34eb5d569a511ce52276efcc76e14d934dfa9d260db39af0810

Observation c0a7a3bf-31e8-46b7-aa1d-4129418c2e1a · outbound

This paper cites Revisiting energy based models as policies: Ranking noise contrastive estimation and interpolating energy models.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Revisiting energy based models as policies: Ranking noise contrastive estimation and interpolating energy models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.940238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:aafc36444ef197df497564786327447b10dfe5efc5c7dff2e2349831e3c86ffe

Observation 9190a2c3-5710-4598-92e9-5f3212d7a73d · outbound

This paper cites A reduction of imitation learn- ing and structured prediction to no-regret online learning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback A reduction of imitation learn- ing and structured prediction to no-regret online learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.955193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:49f6fc2f5851da24fcbd9a0a97ffa06de6f60fb11a01ef36c3ca819db8c157d6

Observation 7f986c3d-15de-40c1-b223-7a2a946e923f · outbound

This paper cites Hg- dagger: Interactive imitation learning with human experts.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Hg- dagger: Interactive imitation learning with human experts

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.950848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:087b6c2ebafefd145c0179daa8a11eda78060ff95abe1b2238bd904b15dca43d

Observation 5c93ad7d-28c3-4f4e-b295-26d68cdb257a · outbound

This paper cites Ambient diffusion: Learning clean distributions from corrupted data.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Ambient diffusion: Learning clean distributions from corrupted data

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:30.944165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:9ae2bba066ffee64c454f40d222ed55678d5b4ac1beccea9afecb19cf3910b02

Observation 94a20fcf-5862-414a-9899-71cf91debbf9 · outbound

This paper cites robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback robosuite: A Modular Simulation Framework and Benchmark for Robot Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:02:30.265739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:41f9233d1cf0623437c8f6cc79eaf866694244800bab5bf494d28dd68cbf9109

Observation 3f9023d4-a2b2-4504-89e5-5742e15f9c2a · outbound

This paper cites Interactive learning of temporal features for control: Shap- ing policies and state representations from human feedback.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Interactive learning of temporal features for control: Shap- ing policies and state representations from human feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.101640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:afe30e42404b1af920aa11c207844b3ca822078000a97f56e03dcb8acf5b5f81

Observation 3b05a8d8-7cfc-482d-9554-525e7165c0c5 · outbound

This paper cites Bayesian learning via stochastic gradient langevin dynamics.

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback Bayesian learning via stochastic gradient langevin dynamics

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:02:31.144762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T04:01:18.219745Z digest=sha256:b13ffa12a5bd40ef57f11598b6b25736c0f1f61ddf24e910fc69bb39ce41a309

Pith citing papers

Observation b9f6d8ea-b5a9-4042-893e-8e764753b7b3 · inbound

Wavelet Policy: Imitation Learning in the Scale Domain with World Prior Memory cites this paper.

Wavelet Policy: Imitation Learning in the Scale Domain with World Prior Memory From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T20:52:06.320648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T20:49:27.459047Z digest=sha256:29d2515eab5ad410c38f01775582e3c92cabe57c7dcf627456df90c677a4d5dd

Observation 8ab9caa0-8478-4c6a-891e-f52a3913ffd6 · inbound

CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation cites this paper.

CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:33:10.067648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:33:10.067648Z digest=sha256:d9ad276a6c2b09482b017b29af9e5c771e2d80e484fc845fca048b19da55ad57

Observation 0a4d691f-7451-45c8-9af4-c098c2b28288 · inbound

Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections cites this paper.

Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:26:21.966351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T14:25:37.987916Z digest=sha256:bc35a174662df139b619296fd1e601aff6c9e8befa4b3b49ffdc590ae1c06f69