Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

As of 21 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2607.08647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08647 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T04:00:47.056185Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 208e9110-1738-43dc-8668-0ea9739d0691 · outbound

This paper cites URLhttps: //doi.org/10.1007/978-3-642-00982-2_1.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning URLhttps: //doi.org/10.1007/978-3-642-00982-2_1

Reference 1

Resolution
verified exact
doi, observed 2026-07-10T04:06:44.494310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:3af165c5863ca57d3d33f7d5b555c681bfba886a36faf8d3637b59491874bf44

Observation 724bce8f-8f9f-4345-bdfe-4a3af69c8bed · outbound

This paper cites Understanding the Power and Limitations of Teaching with Imperfect Knowledge.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Understanding the Power and Limitations of Teaching with Imperfect Knowledge

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.795817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:931e36647045fce230f672c017b9a0e5c9b37c3d6859dc1db77a32573ad4c5ee

Observation 89ae9dcb-9b29-4a60-bfdc-084100c778b2 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.799707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:3b59ceeefa9b24c0b38688bf3d2c0793dcc86bdeaa8a62dff30bf8ff8ae70e2a

Observation 0cfe838c-0d66-413a-bbbf-cf1ecf6df5d8 · outbound

This paper cites The effect of modeling human rationality level on learning rewards from multiple feedback types.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning The effect of modeling human rationality level on learning rewards from multiple feedback types

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T04:06:45.255728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:373cc1cb04f5f3b63fca4b27afda02e91cc83e62422093c60fdbd69b8c6137d0

Observation 4bbd38d5-4249-47c0-bb47-2893f642f869 · outbound

This paper cites Assisted Robust Reward Design.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Assisted Robust Reward Design

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:f9857ec7915b8b7a975ad33083aefc5a6805e98fcf7a44e16511ab7fe23db823

Observation eab54314-22fd-4d18-8bc3-a2d7e765b742 · outbound

This paper cites Interactive Teaching Algorithms for Inverse Reinforcement Learning.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Interactive Teaching Algorithms for Inverse Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.792911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:8adc0c755d10f2984b1be6ba386bcf7b7ff01b53af175ad67c818e8ac3d44451

Observation 165abf34-3469-4981-a613-e4638b6b1fc0 · outbound

This paper cites Mehta and Dylan P.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Mehta and Dylan P

Reference 7

Resolution
metadata mismatch
doi, observed 2026-07-10T04:06:44.496762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:96e1a251663ff64c3826c4762429da3d678490e43dcf416b86a39706c44f4b84

Observation f31d1d4c-ea9e-4b6a-856b-c57a00822c25 · outbound

This paper cites Effects of Robot Competency and Motion Legibility on Human Correction Feedback.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Effects of Robot Competency and Motion Legibility on Human Correction Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:06:44.789465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:6851fe3a785dace315a18b9f6c3b2fac7894f28e4eb316a313d5855feade2bbc

Observation 05e0804c-f802-43d4-803d-94241f7b81be · outbound

This paper cites An Overview of Machine Teaching.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning An Overview of Machine Teaching

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T04:06:44.803111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:ab7fb18c94804a74abee13a5c412ba9c641dc459d56a18175fce8e6f91bc78eb

Observation 395203d5-155b-44ed-8db4-dd4d596c1fd9 · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.253692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:25f0809c1879c7484ed45123a5fb4e355c771ea919baad907b3e228d4385f3d9

Observation db90686e-f1e6-4046-9935-59cf7cb9866c · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-10T04:06:45.250108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:126187c3e8cfe951779d2c379ffdd07e9fc7815785bdcd1f1d64c735139c47c7

Observation f4b88810-7a0b-417d-ab89-bd2594532659 · outbound

This paper cites We use2×3gridworlds with two cell features (drawn gray and white) and a randomly placed terminal cellTthat may occupy either feature.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning We use2×3gridworlds with two cell features (drawn gray and white) and a randomly placed terminal cellTthat may occupy either feature

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.246599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:0efa58c2cff624a8a7ffec0cb304320d3d6875dc9e395935dbeecea3624ab76d

Observation 8a7cab3b-aa36-43cc-a292-133c09631019 · outbound

This paper cites an unresolved cited work.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Unresolved cited work

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.248495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:98ff87fd6a448c1f1a2b4473040590ff13c98f74f40deb47542c9794681b0b99

Observation 5a762f50-e6a9-4847-9796-477fb4079fb4 · outbound

This paper cites S1) 2:Restrict candidate atoms to those in environmentsK 3:D←Greedy Atom Selection(K,U)(Alg.

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning S1) 2:Restrict candidate atoms to those in environmentsK 3:D←Greedy Atom Selection(K,U)(Alg

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T04:06:45.252086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:00:47.056185Z digest=sha256:8c0fe659e934621c4580dd996c52731d4262962f8f63bbc72f2e430f0df28c91

Pith citing papers

No inbound Pith citation observations are available.