Pith. sign in

Paper Citation Record · LEDGER

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning

As of 18 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2505.02483.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02483 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:54:43.782739Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc4e72d-99b4-48b5-9bb2-99c249f151c1 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.095329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.673416Z digest=sha256:3fa79a52460011934948511d36c4bfb649d711ff58ee2ca48b3eca79f8868883

Observation 7be07711-2ff7-4ba6-9a17-b9e057de5c98 · outbound

This paper cites Universal value func- tion approximators,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Universal value func- tion approximators,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.086766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.677982Z digest=sha256:0544e08167d23f4278cf446963fa55b0565c9777e166b602d520bf75f2063da5

Observation 9c494614-e082-4a82-91ec-90c93df8ca6b · outbound

This paper cites Learning agile soccer skills for a bipedal robot with deep reinforcement learning,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Learning agile soccer skills for a bipedal robot with deep reinforcement learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.077393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.681404Z digest=sha256:15009d65727514ce5a0467a6459f851d2b7b5165b2d71cfba717e1c5f2386267

Observation b8b3681f-b9c1-46ea-a421-4e073e45a213 · outbound

This paper cites Deep reinforcement learning that matters,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Deep reinforcement learning that matters,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.068740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.685433Z digest=sha256:ec3d5b576aa805d8496102d23143ed690338348c40a777fff8e0bd2f94d0a640

Observation 58fb2ed7-a9a8-48aa-b4ed-7966fd36e6a9 · outbound

This paper cites Discovering reinforcement learning algorithms,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Discovering reinforcement learning algorithms,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.059787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.689119Z digest=sha256:7ce826172861b4a41b68119d036310b083111d1ff72df06573a2190538453b7a

Observation 5ccd08d0-c3ea-4bd8-9f37-01073c09918b · outbound

This paper cites Legged robots that keep on learning: Fine-tuning locomotion policies in the real world,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Legged robots that keep on learning: Fine-tuning locomotion policies in the real world,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.692577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.692577Z digest=sha256:6633e0a8b306f017cec642dfc7352ca42e9aadb14110423bef2bfe7a05ee5af1

Observation 67179b22-4a15-4098-8083-a852534a9cc1 · outbound

This paper cites Reinforcement learning for reduced-order models of legged robots,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Reinforcement learning for reduced-order models of legged robots,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.043906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.696607Z digest=sha256:95505a988817622606866376b5eb48904adc45ec3979e3788206b347fb7e54da

Observation 022deb18-6bbb-4fc2-950a-1df922ac3a0b · outbound

This paper cites H-index: Visual reinforcement learning with hand-informed representations for dexterous manipulation,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning H-index: Visual reinforcement learning with hand-informed representations for dexterous manipulation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.033737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.700066Z digest=sha256:a51822dc917abd23c599c269198ee96b6d5f40c98bcf36e3e3f76ef943f4e0fe

Observation b44dcc3b-7eb9-4e7d-acfc-d256f11d4c80 · outbound

This paper cites Dexterous im- itation made easy: A learning-based framework for efficient dexterous manipulation,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Dexterous im- itation made easy: A learning-based framework for efficient dexterous manipulation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.023571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.703405Z digest=sha256:74d28af70c1b48ad62d8aa89f55e2896e9c30772a02e27e7c3595bf2151d9412

Observation 33dd97ec-da57-427c-9b08-377bbd68636a · outbound

This paper cites Bi-dexhands: Towards human-level bimanual dexterous manipulation,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Bi-dexhands: Towards human-level bimanual dexterous manipulation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.013923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.706315Z digest=sha256:bc706b43fa122d4f35cb17fd6ca7808d500ec3597b15079a2ea3eecad456bf87

Observation f0574433-8b35-4a70-bf97-801eb0f0fd29 · outbound

This paper cites Learning multi-arm manipulation through collaborative teleoperation,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Learning multi-arm manipulation through collaborative teleoperation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:44.003963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.709355Z digest=sha256:05377c5fb0c584bc4928a3f076b7d12995a9b3a27afd96d202cded161de3943e

Observation 11424681-f762-44b8-8b1c-2c377687f488 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Playing Atari with Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.712698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.712698Z digest=sha256:450df314752a4a379456e3340a21e9ad29bb2fcb473495a0ee7094d51b8de16b

Observation 05322926-24ee-4cbf-8c5e-e58a67702019 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.716066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.716066Z digest=sha256:eadbfd40a102972f9dd093237cb3be9a3fb2f7dad47626d31d41cd9dcee94085

Observation c2492b78-cfe9-4607-b515-4a97fea48006 · outbound

This paper cites Hybrid reward architecture for reinforcement learning,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Hybrid reward architecture for reinforcement learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:43.994371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.719428Z digest=sha256:80eae8063cc5b6fa272635eeb3b20f3d343ad5562824a8a8616b04fb477b0eb9

Observation 429525b5-7e31-4838-a362-a1ae9311e760 · outbound

This paper cites Reward- adaptive reinforcement learning: Dynamic policy gradient optimization for bipedal locomotion,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Reward- adaptive reinforcement learning: Dynamic policy gradient optimization for bipedal locomotion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:43.984412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.722488Z digest=sha256:16faeabe64c373f9c2af5d58195adfa0346cdfc843b38372833111e7d02bf8b0

Observation a99c9a18-1d7a-4935-a394-265f42c82970 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.725448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.725448Z digest=sha256:eb47e506edf8db021bd6d1e77a3dc8b04d7a32b1f1e6880ede882fe44d80c97a

Observation 40b369bd-4e7b-41ac-b9e3-1f1f630586b8 · outbound

This paper cites Horde: A scalable real-time architecture for learn- ing knowledge from unsupervised sensorimotor interaction,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Horde: A scalable real-time architecture for learn- ing knowledge from unsupervised sensorimotor interaction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:43.974390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.728949Z digest=sha256:6abccefb526064b43b2259edc183e7b3e620fd99b1c87de4078123b5ddb8581a

Observation 7f22f43d-f9ae-4570-9740-53a17689252a · outbound

This paper cites Continuous control with deep reinforcement learning.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Continuous control with deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.731667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.731667Z digest=sha256:49f3526821e47ba13dbd87f9861cee57f0bb6d3ffffb432982eb380588c29bea

Observation 41c75371-4942-4334-9b84-77984b53718f · outbound

This paper cites A survey on integration of large language models with intelligent robots,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning A survey on integration of large language models with intelligent robots,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.734906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.734906Z digest=sha256:675cac9b0e79fb3866d05ddc0bd4dc37b167e4e16c146989942034b5d7f3cbad

Observation 36caecda-11fb-4d92-9bf4-4ce3a694632a · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Chain-of-thought prompting elicits reasoning in large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.738097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.738097Z digest=sha256:0e69c75e35094414c7435b51b832bfa0d828d752010d5a343442b42201b6b2e7

Observation 1a183650-6279-4444-add1-1409a6ce4935 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.741164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.741164Z digest=sha256:2f75f3739dd529fd1ef8df2e8dda13d1519d7939e425382c18c97dbd8e9ca33a

Observation 829fcc76-1639-4aee-8eae-06d6b909b9ca · outbound

This paper cites Lever- aging pre-trained large language models to construct and utilize world models for model-based task planning,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Lever- aging pre-trained large language models to construct and utilize world models for model-based task planning,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.744323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.744323Z digest=sha256:7e5b0cc1ff0be392196b24a014f869f9aa343ddfbb0e33d0225059672cdf93df

Observation a745327b-295c-4406-b6d6-958ce422ff24 · outbound

This paper cites Text2motion: From natural language instructions to feasible plans,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Text2motion: From natural language instructions to feasible plans,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.747215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.747215Z digest=sha256:f3516d3d54f916678280b19d72b3d69ddf0ec0320c0f5d1d66f7540a76674445

Observation a33a2b9a-6296-42ca-b137-51740b31171f · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Progprompt: Generating situated robot task plans using large language models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.750524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.750524Z digest=sha256:9fe017b802b16ad666db6f0d2fc460a4def7da8dbad10ea9fe270d12af357809

Observation 961f2aee-d667-4294-8c22-2ef4fe8c6951 · outbound

This paper cites Code as policies: Language model programs for em- bodied control,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Code as policies: Language model programs for em- bodied control,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.753217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.753217Z digest=sha256:d8c03b0e032574f904ae161f5f9a6288228fd9a59d77d7c1446ecdcc338946f7

Observation 11a509e6-53b0-4e6a-8e95-fbe315ea49be · outbound

This paper cites Language to Rewards for Robotic Skill Synthesis.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Language to Rewards for Robotic Skill Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.755841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.755841Z digest=sha256:6a6d2f555e140a76145290f8d6d89df9c541c6cd7b3cdb215e6651199baf78aa

Observation 3a911811-6640-401e-a8cb-8621f6bce29e · outbound

This paper cites RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.759112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.759112Z digest=sha256:84f4ba4b946284698fdd909cfaa8b28c7e305007f59cbd24e97b30168a953b5c

Observation 7ad524be-6e0c-4b03-9e89-21f3ebdfc227 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.762359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.762359Z digest=sha256:3d9010a114654fc2a0b596fc52feb0edd939997ab6b06b32c3e030532cddca6a

Observation be46138e-83b1-42da-a143-f3fc8c672178 · outbound

This paper cites Language models as zero-shot trajectory generators,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Language models as zero-shot trajectory generators,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.765362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.765362Z digest=sha256:6cae25aab80e2e96c3dabb2bc3ec9011a0b4f697bbab6e3cf4ece395d537cd27

Observation 3a1b1363-fbb6-4bc4-b161-395902cf2e4a · outbound

This paper cites Roco: Dialectic multi-robot col- laboration with large language models,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Roco: Dialectic multi-robot col- laboration with large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.768185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.768185Z digest=sha256:4940dea6f2324762284d1255251ecae85de42620bb55bc432051f56658254fb9

Observation a414e3f0-d429-4d25-983d-aa9ac84349b1 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.770991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.770991Z digest=sha256:1891703632f029ec3bf1ffe90593738e26e795f76476c171200e7132961e23d0

Observation a5c51b19-5e81-4d5b-acb3-07f223edf92a · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Interactive planning using large language models for partially observable robotic tasks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:43.919870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.773737Z digest=sha256:c1dbdc058eabd24b6300589eb8f8b87d8e07dfba5785b6586689a4ac576922a8

Observation e925e5d6-4fe9-4282-a77e-9ecd7c08064f · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.776461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.776461Z digest=sha256:a76aa3376b1678af70fe2dcfe999cd0cfb3773edd8e000ab42092b409f87454a

Observation d6bdb82f-9553-40d2-a7f1-02d376909e91 · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.779796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.779796Z digest=sha256:e18e9e62d556dde1d14badc539ba1c2e55a4f1428789a47841a0de7196a5cd64

Observation fc26b991-d232-4ef3-b8bf-d16a1277c85d · outbound

This paper cites Explainable reinforcement learning via reward decomposition,.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Explainable reinforcement learning via reward decomposition,

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:54:43.910328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:54:43.782739Z digest=sha256:0f10c5fa272618b8a73cf4880d2cd2354f0597c4c0f38b1cc593d6c8cb383cf6

Pith citing papers

No inbound Pith citation observations are available.