Pith. sign in

Paper Citation Record · LEDGER

Toward Virtuous Reinforcement Learning: A Critique and Roadmap

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2512.04246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.04246 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T01:55:22.266761Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact5
  • verified fuzzy27
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65fc118f-8b8e-4d5b-be22-1152804a5203 · outbound

This paper cites Towards artificial virtuous agents: games, dilemmas and machine learning.AI and Ethics, 3(3):663–672.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Towards artificial virtuous agents: games, dilemmas and machine learning.AI and Ethics, 3(3):663–672

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.797474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:f367505a0135e3fcf9169d307cf8706c536dd25123d4cd3796a57bb45677ebcd

Observation b8184140-7a14-49c7-b189-e15938be5ce2 · outbound

This paper cites Reinforcement learning as a framework for ethical decision making.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Reinforcement learning as a framework for ethical decision making

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.774477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:def7a8d97e2b5bd55947aa36518731d6fbc83298904385962761c6b51b7c515a

Observation 0b603340-0650-4831-9dda-69d18fe323bb · outbound

This paper cites Reinforcement learning and machine ethics: a systematic review.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Reinforcement learning and machine ethics: a systematic review

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.927971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:130de41bb90aa0fd8c31689e438bcd51b6097756b2dc977d26db86cbce621f64

Observation 3d317330-0976-4308-acf6-f8405188ed22 · outbound

This paper cites Groundwork of the metaphysic of morals.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Groundwork of the metaphysic of morals

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.808489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:f396e577a02db3cbd4a6d3a22bf0b51531dcce0b7807c340592aef474bd1d5a8

Observation 165de463-5327-4399-bc58-858b22e40eba · outbound

This paper cites Utilitarianism.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Utilitarianism

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.751337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:c724cdce363a1c248891937c0a7b948d93d4925c5c39f39d0f8ecb646740e72e

Observation 6a602ff6-f407-412c-bcfc-0ddac7a374b3 · outbound

This paper cites Safe reinforcement learning via shielding.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Safe reinforcement learning via shielding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.777540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:e60e3cc7db3e7f2d3429ce2a038e9402a4dbe4f5b09ff5f55a8a6f90f4556090

Observation 93f96e0d-dead-4312-9175-88f88be0362e · outbound

This paper cites Can model-free reinforcement learning explain deontological moral judgments?Cognition, 150:232–242.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Can model-free reinforcement learning explain deontological moral judgments?Cognition, 150:232–242

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.754625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:e6ab357b626549f6fe04a91060f8c9b7cc37a655d647123f12eba76847b35dfa

Observation ea0ae0e3-ee69-4e20-8545-6b5890ae6943 · outbound

This paper cites Reinforcement learning under moral uncertainty.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Reinforcement learning under moral uncertainty

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.768431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:f7437a55499bbd07a202e54ca60a19d01b89f7a6d14d71fbd52f522097c58398

Observation cf320c6b-aa29-4dfe-bea4-e1f9f9dbb7c0 · outbound

This paper cites Q-learning as a model of utilitarianism in a human–machine team.Neural Computing and Applications, 35(23):16853–16864.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Q-learning as a model of utilitarianism in a human–machine team.Neural Computing and Applications, 35(23):16853–16864

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.804121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:d063b623bdf7c1066d5cc2ddfb56e2a601d285f4908763c70ec13bb4aa63e42d

Observation 0041ea57-7882-47b9-ac40-77df2b52780b · outbound

This paper cites Artificial morality: Top-down, bottom-up, and hybrid approaches.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Artificial morality: Top-down, bottom-up, and hybrid approaches

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.748504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:a6079c10c47829137511acf1c3cbb54919a27e9ed4a42bff8d4310251871c9f0

Observation 6e1235e9-c99d-4d1c-97b8-b5e20f6e2355 · outbound

This paper cites Building ethically bounded ai.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Building ethically bounded ai

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.757415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:f66398b8b6a3cef2b2db232963c5b8347e6d56ec2ecf6b1279f188d1f528cbaf

Observation b0b3d56a-599b-4450-8476-ab972085d163 · outbound

This paper cites Building Ethics into Artificial Intelligence.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Building Ethics into Artificial Intelligence

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:58:51.922355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:378c6c15bdbf550257167695496ccaf053e29237cf792426a73412b84ffc6ac7

Observation 4797add2-a982-442d-ab9b-39d7dd1263f0 · outbound

This paper cites A low-cost ethics shaping approach for designing reinforcement learning agents.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap A low-cost ethics shaping approach for designing reinforcement learning agents

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.762694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:974663ef21e779d415efdf818f8face3b4f3ab77e28d7611e7a01c46e4a85b57

Observation 5094bf44-3db7-4cc3-88d1-19f68825dcf7 · outbound

This paper cites Virtuous vs.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Virtuous vs

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.765337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:a5a5889d28e320ed9557e1a00b89612e44834c27a4be22560030e04cf54adcfa

Observation d044e971-4ef7-4f6d-970f-68b7082cc029 · outbound

This paper cites Teaching ai agents ethical values using reinforcement learning and policy orchestration.IBM Journal of Research and Development, 63(4/5):2–1.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Teaching ai agents ethical values using reinforcement learning and policy orchestration.IBM Journal of Research and Development, 63(4/5):2–1

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.771409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:ff207ba157a79bb47b0dbac78bb7553486b07f86b5edf3d4dbf385b3755dc621

Observation 29d42ca0-a2b1-4545-be7d-b809c8d04ff9 · outbound

This paper cites Cambridge University Press.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Cambridge University Press

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.759968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:a79329a164f46c31ff4f617e41b5109b7496f6babab22131b386f6e349e724bc

Observation f8e052f7-d963-4aea-a4e5-298b88ef6df0 · outbound

This paper cites Right action and the non-virtuous agent.Journal of Applied Philosophy, 28(1):80–92.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Right action and the non-virtuous agent.Journal of Applied Philosophy, 28(1):80–92

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.829283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:4207cce82021e1b9fecda959a92753c043289d4054a7ca6f6a0adc3509b9eaaa

Observation 7f1e9bae-cfcb-4c21-8af9-83b56572ba6c · outbound

This paper cites Introduction to Reinforcement Learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Introduction to Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.906720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:44d765a2c8ea24d09d012f836ad8678576c120c243bd975df03363741f8352a7

Observation 089947a4-0e68-4e35-9f5f-7441dcb8e630 · outbound

This paper cites Joint Attention for Multi-Agent Coordination and Social Learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Joint Attention for Multi-Agent Coordination and Social Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.917184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:de82522663487d4348cbdc3b61591c7a7895da17b4698d6f25e1184366ed573a

Observation 35039a66-c971-4ece-bbef-011ca266d317 · outbound

This paper cites Learning few-shot imitation as cultural transmission.Nature Communications, 14(1):7536.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Learning few-shot imitation as cultural transmission.Nature Communications, 14(1):7536

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.818867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:7452048a08334618c46af7dc24c3b73693c0858ea208ce4764c1c8d9f4b42119

Observation ddf573f2-442c-44d0-86d3-9496a13d9d89 · outbound

This paper cites An Efficient Open World Environment for Multi-Agent Social Learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap An Efficient Open World Environment for Multi-Agent Social Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.912009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:c0f47d61c7ee67bba356a79a9ba140f9836e0c5f29e2a9867f9efad88d8873d5

Observation 8efbd136-d07c-4f8b-b2a6-0b38994e95bb · outbound

This paper cites Emergent social learning via multi-agent reinforcement learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Emergent social learning via multi-agent reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.873772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:b16e042b639434c13b5af9a692038515af7027ee7b93360b9fdbd04966912e66

Observation 1a641980-ca1d-4a19-91fb-6e137c7a6a18 · outbound

This paper cites Social influence as intrinsic motivation for multi-agent deep reinforcement learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Social influence as intrinsic motivation for multi-agent deep reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.860647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:c860d4a22cf7f786a8e831b0f11bf6324e5803586e5a2a4b034afc5471d1d12e

Observation a88686ce-f456-4a6a-8516-9bd636c56849 · outbound

This paper cites Multi-objective reinforcement learning: an ethical perspective.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Multi-objective reinforcement learning: an ethical perspective

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.840830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:c827ed8c63c64eb07b2721eb14a713cd226500a56c540d30a779fb99dbbf2045

Observation 8a690175-566a-4676-a538-cfc137fbc855 · outbound

This paper cites Exploring affinity-based reinforcement learning for designing artificial virtuous agents in stochastic environments.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Exploring affinity-based reinforcement learning for designing artificial virtuous agents in stochastic environments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.813110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:5cc6e966c19f91b302cec963e6a3272a1d6c41d505d0fc6890c78035aaab5cbf

Observation 55ec86bd-7830-47c0-a72d-9ab0bf735c33 · outbound

This paper cites The core of confucian learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap The core of confucian learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.868522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:5eeb33cc0db3b5fdb79b98efd92cfc4ad97c785bdb1993498b99a456346b0f5e

Observation 98acc1b5-0419-47db-ad4e-c3ef28a39a9b · outbound

This paper cites The daoist thought of wu wei–action through non-action and its influence in vietnam.Synesis (ISSN 1984-6754), 17(2):55–71.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap The daoist thought of wu wei–action through non-action and its influence in vietnam.Synesis (ISSN 1984-6754), 17(2):55–71

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.864786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:415831e1783c8fd056e5ea37b6337a9a65da7dc184b0be6180f7bb3df6390806

Observation 789ad6f0-1e3d-4e3b-b5a9-c11f7170a1f3 · outbound

This paper cites An anthology of philosophy in persia.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap An anthology of philosophy in persia

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.851347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:52a37edd966cd2162c2504752a5082519389704e5c690a0e1c58dcf353329f41

Observation c8b9c0a0-1b71-403d-bbd0-000513057181 · outbound

This paper cites Composable modular reinforcement learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Composable modular reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.845494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:8562eaf9a66756688e160a2d4a4aea113490cab6c5d4c16b85a8f1549f584149

Observation 4baf5916-afb9-4490-8e2d-77a4c7d1e6fe · outbound

This paper cites Formal verification of ethical choices in autonomous systems.Robotics and Autonomous Systems, 77:1–14.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Formal verification of ethical choices in autonomous systems.Robotics and Autonomous Systems, 77:1–14

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.837274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:b472c770c322b1c4a83922d47ab40893df21309eb1e3480699ff512e55028ca6

Observation de968810-320b-43e0-b86b-1cbb1271b02d · outbound

This paper cites Ltl and beyond: Formal languages for reward function specification in reinforcement learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Ltl and beyond: Formal languages for reward function specification in reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.833813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:129a4f5cc372b09d216f316fb5d343edbfd974bc13e88dff61e91644780628eb

Observation 6dfc234c-3620-48bf-9429-82d71cbe5d65 · outbound

This paper cites Using reward machines for high- level task specification and decomposition in reinforcement learning.

Toward Virtuous Reinforcement Learning: A Critique and Roadmap Using reward machines for high- level task specification and decomposition in reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:01:26.824513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:55:22.266761Z digest=sha256:4bb8532745b3b83ecea426d7585dfc15ece00e28c646b111c2403f38450c2875

Pith citing papers

No inbound Pith citation observations are available.