Pith. sign in

Paper Citation Record · LEDGER

Language to Rewards for Robotic Skill Synthesis

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 68 inbound Pith citation observations for arXiv:2306.08647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.08647 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 68 of 68 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:15:06.101518Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

39
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4485a3e3-d36e-4933-8e16-4ef134db797f · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Language to Rewards for Robotic Skill Synthesis

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.346925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:0c31edb03e36c184750c50eaac80f10d2cc580ece0806c3eba227bfd72d562ff

Observation bbfd1523-8049-438c-bf99-3b29f4d0d1f2 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Language to Rewards for Robotic Skill Synthesis

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.072643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:b280590019fd9f31bd784e1ee3df95df9c9561b854523da71746c442dd3171c7

Observation 8072d747-791a-4e4a-8c8b-94b0528f86c3 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Language to Rewards for Robotic Skill Synthesis

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.511826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:0913296dec8baf2f2aef80b4ef9206ed4d557ae1fe24145ceca04a6e0d9d413d

Observation 59cd022e-4dc6-471d-94d3-291d033c800e · inbound

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery cites this paper.

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery Language to Rewards for Robotic Skill Synthesis

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:42:32.118153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T04:42:31.555355Z digest=sha256:882987b5a4847ac2752902efe4bbb5006892e14c1b6f348eac7f45a3fe4c347a

Observation 9fb7ffba-bd1b-4e03-a479-4de07da1c058 · inbound

Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement cites this paper.

Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement Language to Rewards for Robotic Skill Synthesis

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:03:26.377798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T21:02:03.017690Z digest=sha256:499a09d65320e508bbb32106ce213f95b9cf8742328ee6eff66012a62bbb8547

Observation 601b4ca6-7ad3-4ab0-b277-0ec3f2a6a0b3 · inbound

Agent Workflow Memory cites this paper.

Agent Workflow Memory Language to Rewards for Robotic Skill Synthesis

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:51:28.457802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T00:51:28.398311Z digest=sha256:9c2866c9569b056ba60944e94aec51930b3c6949846e6c5a2d7e9e727cd330f3

Observation 31cf42f0-ce24-4d41-ae79-4c2c5f62817e · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework Language to Rewards for Robotic Skill Synthesis

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:16.151245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:16.151245Z digest=sha256:b3198291e886e81439c63eaf5871bc81ad5199b58639039553f63594e6bbed36

Observation cc6d9f2d-48de-4635-a607-70d547252ef2 · inbound

PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making cites this paper.

PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making Language to Rewards for Robotic Skill Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:43:44.561260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:43:44.561260Z digest=sha256:275204d9aea9fdd2fa92d61cd3e3969173b4e611f1ab26e90ff8144370b78e45

Observation 3047f5e7-f089-47a7-a5b2-dff8849b393b · inbound

MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation cites this paper.

MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:03.899045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:03.899045Z digest=sha256:e150fcc55938dbf96cceea2f589e3bf26ae2c7c20a0ba4f2fc4100a43a2c5ce3

Observation d4e6acfe-f218-4db9-803d-0b821e9bd7d3 · inbound

Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning cites this paper.

Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning Language to Rewards for Robotic Skill Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:43:06.072105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:43:06.072105Z digest=sha256:0747076d6a53b98b96a2b1758257f184aa96ccdf7efee759d760e495ce7a6a7e

Observation 78bb1bfb-4b1f-46ad-9252-62cbe2a4ec24 · inbound

The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data cites this paper.

The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data Language to Rewards for Robotic Skill Synthesis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:23:12.104049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:23:12.104049Z digest=sha256:9db3cb72e3f7e4ef637be76d0ceb4e55f8603cf7985dbfc839d74b968a296b09

Observation 8840e9ed-b986-4ad2-af2d-8e4251dd001b · inbound

VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving cites this paper.

VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:15.422893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:25:15.422893Z digest=sha256:e63c46e0934eb8028ce91999d8ace35333526624d8c5591d08a68c5fc997d402

Observation d61a7231-fbe7-4fff-aa94-b9f870b33dd8 · inbound

CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning cites this paper.

CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:10:08.698596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:10:08.698596Z digest=sha256:c05ee806b55a1f6abefb71fc0f87456b27acd7aaaacc052a7633d9bd023bd5e5

Observation af46d44e-5259-47e5-8105-601d6c61ea28 · inbound

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning cites this paper.

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:52:27.291056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:52:27.291056Z digest=sha256:c2b33cb6092c133ba839ff918d9133d118c3ab2971a81dbca331b937577d436d

Observation 2bda9692-72c5-4e35-b572-8005643dc710 · inbound

Rapidly Adapting Policies to the Real World via Simulation-Guided Fine-Tuning cites this paper.

Rapidly Adapting Policies to the Real World via Simulation-Guided Fine-Tuning Language to Rewards for Robotic Skill Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:31:10.776502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:31:10.776502Z digest=sha256:f3452a37bad883ec76fdda683b5a94d695d2ffe4879c9f7bdbf41eeba5da1d37

Observation 64b65b33-7600-4aa1-b7b9-cf8a7245ee5a · inbound

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning cites this paper.

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning Language to Rewards for Robotic Skill Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T01:04:04.072614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:04:04.072614Z digest=sha256:d00414801fb9cddbcfae7f21a8176c513d1fcfdc14903ea12695d1e2faca4174

Observation b5c28687-3fbd-496b-865e-d78dc4aa2b1e · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Language to Rewards for Robotic Skill Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.484627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.484627Z digest=sha256:ed803db98e5c417836af80e88102ea191f9cb3dc5a2dbd29a6d50364e35e4ee9

Observation cd1dbb4e-a13b-408a-bce8-6d3f4c70a20d · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models Language to Rewards for Robotic Skill Synthesis

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:21:45.143391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:573d047d6851c8d9e063ca091f46439b81fe0052701cc37e40830fa4f05fb4ef

Observation f5a88b61-30fd-43b7-85df-92a144c2dde7 · inbound

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models cites this paper.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language to Rewards for Robotic Skill Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.101518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.101518Z digest=sha256:3a47a8a437157e79a2435ec9d2376de376a988208313bba4031e3ea010b43eec

Observation 0b208c99-aa67-4029-a99d-c19ef7265262 · inbound

PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations cites this paper.

PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations Language to Rewards for Robotic Skill Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:31:55.911543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:31:55.911543Z digest=sha256:a249209f019985a5e40dde31accd8d47e92b2df0d743008638cc55ad1e8d789b

Observation 11a509e6-53b0-4e6a-8e95-fbe315ea49be · inbound

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning cites this paper.

Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning Language to Rewards for Robotic Skill Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:43.755841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:43.755841Z digest=sha256:192542cf1b2b24b7bf38e866114108c74449a8bc9c322c95ca929e6df0e06058

Observation ae1adfe9-f324-4f75-ada1-2db330993a1a · inbound

RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning cites this paper.

RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:01:33.893211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:01:33.893211Z digest=sha256:c0a438018c1854a0b0d690993ebbfbb024e083ac56b552a4eceaa3acf4dfbb61

Observation 8fe68a62-3f43-45ae-ad6b-25f62ac3234c · inbound

Multi-agent Embodied AI: Advances and Future Directions cites this paper.

Multi-agent Embodied AI: Advances and Future Directions Language to Rewards for Robotic Skill Synthesis

Reference 233

Resolution
unresolved
no resolver link, observed 2026-08-15T23:16:16.038103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:16:16.038103Z digest=sha256:f7556cd7487c83df6e02e8926273031458ed13c386c62e5376fedd4da7e7ceca

Observation 09664b8e-bf87-4481-af80-1e13741bba5d · inbound

Motion Control of High-Dimensional Musculoskeletal Systems with Hierarchical Model-Based Planning cites this paper.

Motion Control of High-Dimensional Musculoskeletal Systems with Hierarchical Model-Based Planning Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:06:03.363467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:06:03.363467Z digest=sha256:52f66def504bb31fad9fb2dc7950a02c63b17658008bf15ae0342ad8450947f3

Observation 6ed62dec-8c63-4f13-903d-56de5e93add8 · inbound

MA-ROESL: Motion-aware Rapid Reward Optimization for Efficient Robot Skill Learning from Single Videos cites this paper.

MA-ROESL: Motion-aware Rapid Reward Optimization for Efficient Robot Skill Learning from Single Videos Language to Rewards for Robotic Skill Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:49.493036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:00:49.493036Z digest=sha256:0bc375b5c21d3d99801c619c7633cbb57864bd53930764f4b1e1a6af8e362739

Observation 8a8f4a27-2a9a-4875-9a67-1a308f1a5f39 · inbound

Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition cites this paper.

Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition Language to Rewards for Robotic Skill Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:13.568242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:13.568242Z digest=sha256:7fffc23701427fdf587d4599e15aba186ae7e044836b311dc6bb31a06dcc1572

Observation af353fed-4776-48f1-b44b-244220b05b17 · inbound

LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation cites this paper.

LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation Language to Rewards for Robotic Skill Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:41.067156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:41.067156Z digest=sha256:4044a16fc9afa20bdad65a605559ff9618e8b837797919dcf66b2ae1069ac565

Observation 4da63cfb-30ff-4a4d-b92d-d2132704f41c · inbound

Reducing Latency in LLM-Based Natural Language Commands Processing for Robot Navigation cites this paper.

Reducing Latency in LLM-Based Natural Language Commands Processing for Robot Navigation Language to Rewards for Robotic Skill Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:40.581010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:40.581010Z digest=sha256:0f1375866b245f5e2e82e54f001f82e59ba77a64ed50e10c0fe5406e38cb7941

Observation ae0b7640-be5b-4f5d-9402-c4eaba497d5b · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Language to Rewards for Robotic Skill Synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:37.188094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:37.188094Z digest=sha256:781037859b774863a4f49e31bc37d023d227b93bef38d573385a3839a560820d

Observation 46cd2feb-7740-4a6a-8b00-6b5ee804ad30 · inbound

FEAST: A Flexible Mealtime-Assistance System Towards In-the-Wild Personalization cites this paper.

FEAST: A Flexible Mealtime-Assistance System Towards In-the-Wild Personalization Language to Rewards for Robotic Skill Synthesis

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:22.321330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:22.321330Z digest=sha256:fda257eb22039a279b5c13b8f31f750c115353f05305e4622aae2ac84b9f4353

Observation 9ac9f634-29b9-4390-89fc-1d9dc1a80649 · inbound

ACTLLM: Action Consistency Tuned Large Language Model cites this paper.

ACTLLM: Action Consistency Tuned Large Language Model Language to Rewards for Robotic Skill Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:07.558279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:07.558279Z digest=sha256:255d325a00c7e5262c187d58b974000260d89458626c4dbd37b036fd962622b0

Observation f5278c0b-30a0-4f7e-a5ad-23a2f833e66a · inbound

Prompt Informed Reinforcement Learning for Visual Coverage Path Planning cites this paper.

Prompt Informed Reinforcement Learning for Visual Coverage Path Planning Language to Rewards for Robotic Skill Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:45.182765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:45.182765Z digest=sha256:647a25ce784aacfaca2813d1ce7db7033418458ab7346fd47d42fa0552430c5f

Observation 95a92f59-0f4f-49cc-9247-4211e82f381d · inbound

VLMgineer: Vision Language Models as Robotic Toolsmiths cites this paper.

VLMgineer: Vision Language Models as Robotic Toolsmiths Language to Rewards for Robotic Skill Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:59.846352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:59.846352Z digest=sha256:1835d1943528b2164808a42b6ddee1c066ceb796d852e20fe5a4205f0c84cce1

Observation c6c10f6b-9231-42e3-89c1-5a6461a56315 · inbound

DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning cites this paper.

DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning Language to Rewards for Robotic Skill Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:30.008389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:30.008389Z digest=sha256:d13eadd207d632d2f36b9c1525b2f0c3afadfcb33eeac658257e21688581ad52

Observation 0d2359ed-e86c-48bf-8be6-15a5c3fecf22 · inbound

A Human-in-the-loop Approach to Robot Action Replanning through LLM Common-Sense Reasoning cites this paper.

A Human-in-the-loop Approach to Robot Action Replanning through LLM Common-Sense Reasoning Language to Rewards for Robotic Skill Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:17:07.278935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:17:07.278935Z digest=sha256:02b3ca2ee0c8522bece548e2b70b549904744696e84ebe8eb432f16c6e540e27

Observation 6403afed-e422-4f5f-a27d-c2159ee75ae8 · inbound

GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming cites this paper.

GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming Language to Rewards for Robotic Skill Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:30:47.896896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:30:47.896896Z digest=sha256:d5530f13fad4650942a448c3c8f469a20b91e343ce11c37e110e25bac9002685

Observation 448ad640-4dc2-4e2d-9e0d-1f432652e372 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning Language to Rewards for Robotic Skill Synthesis

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:55.927393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:55.927393Z digest=sha256:f1fe2b310cedaaf54c4f8bca7ea5fcb2d773b97bc0f15933a313ae3a965db061

Observation d501f1bb-1e69-4b08-b6d6-2d3745075ebb · inbound

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation cites this paper.

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation Language to Rewards for Robotic Skill Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:41.754771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:41.754771Z digest=sha256:990f6862359187b5eb64e93056d2c7f7c718adcb6be7a1baf4c9403f980b6796

Observation 39a0252b-85b0-4774-b7ff-05f5fa6c793e · inbound

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions cites this paper.

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions Language to Rewards for Robotic Skill Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:16:14.200202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:16:14.200202Z digest=sha256:708251fbb455de88b75cf96d57d7f9bf0a4cebe52f97f7a81268f28458f3d844

Observation e07873b9-6cc9-4a4f-8d3a-b534774fef2a · inbound

Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following cites this paper.

Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following Language to Rewards for Robotic Skill Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T21:06:12.152973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:06:12.152973Z digest=sha256:a97a52dad725499c7672221afa39d584a4c192621dae1987027f275fbf0190d0

Observation 7db59076-6e9b-41d0-8479-5298be540dcb · inbound

ZapGPT: Free-form Language Prompting for Simulated Cellular Control cites this paper.

ZapGPT: Free-form Language Prompting for Simulated Cellular Control Language to Rewards for Robotic Skill Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:05.064762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:05.064762Z digest=sha256:a850ceab7da9b3144bf2a40f2310479f18e2d0597fcd03ac27d37b6cc27ffb2d

Observation 4bba9fdf-fd86-4a88-8042-e143e09511e8 · inbound

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning cites this paper.

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T16:06:58.564069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:06:58.564069Z digest=sha256:a1f1796408bbd59d382ea510b7141002b7ff6a02578ad30edc9d095325fafb11

Observation d6666f7d-6dfd-4b32-8e5e-e4c7304c732b · inbound

ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution cites this paper.

ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution Language to Rewards for Robotic Skill Synthesis

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:58:58.801754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-16T13:58:58.627748Z digest=sha256:ca8c1282ed3d75245ef9313909677c59166c8b86bf05a69dc9bb7bb1ae7db875

Observation 0e915944-7e6e-4289-9b4f-67d547764ee2 · inbound

Debate2Create: Robot Co-design via Multi-Agent LLM Debate cites this paper.

Debate2Create: Robot Co-design via Multi-Agent LLM Debate Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:27:03.953190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:27:03.953190Z digest=sha256:b2184feb0f1ad7ca2acddc38e0c13e44d52434c23342c4df1be0e002bef21cdf

Observation 7fdc6b1f-60d2-43d1-b0aa-1ecf0104c8d2 · inbound

ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning cites this paper.

ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T15:29:16.815306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:29:16.815306Z digest=sha256:7299a076f83d0b8595e09ad14f5a823b8ed3b76162903f25718d64f7c199848a

Observation 0824365d-b3eb-4fa2-8742-48835417d541 · inbound

PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models cites this paper.

PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models Language to Rewards for Robotic Skill Synthesis

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:07:20.574223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T05:05:41.709469Z digest=sha256:232dfb66539cc093faaba7cf1fd1c73da246f8a8e47cef4d38f57b2a9825a91a

Observation f7af31f6-665c-4958-9a75-37f05281bddd · inbound

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation cites this paper.

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation Language to Rewards for Robotic Skill Synthesis

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.866230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:38:49.399967Z digest=sha256:53620134738e764d60307a2fb5ce5f8a306d2bd76d892e544246c63ce6b4d973

Observation 4305bdb9-6863-4d46-b3b8-cc2b09eb4208 · inbound

Improving Zero-Shot Offline RL via Behavioral Task Sampling cites this paper.

Improving Zero-Shot Offline RL via Behavioral Task Sampling Language to Rewards for Robotic Skill Synthesis

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:18.595209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T16:27:15.347522Z digest=sha256:542971bfd76735aed2c0e4b314973a9cdc9c96f952676b4e04cd7af54b0a8b6c

Observation a8c43925-e26e-4f5b-bd86-c23a4441fc47 · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.200137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T19:14:22.162163Z digest=sha256:5a6dd8f3849490bd6c51e9cc67075a75433f8790f3bc807b01db3c55e241eab5

Observation ae951a6e-0a0d-441b-92e0-2a7dd7b6d549 · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:57.918220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:10:46.725354Z digest=sha256:7c23d2e735c735f18d5d316b25167bf9ed7fd33006126d0ef3ddc62bd799dfad

Observation d686d8f1-a8b7-422a-9fd5-875676b4e26f · inbound

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models cites this paper.

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models Language to Rewards for Robotic Skill Synthesis

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.896581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:19:38.587352Z digest=sha256:6f2a5b4b9667fc78318de13311530269307463092723a737ffaa4eb67af1d86b

Observation a88313d2-debd-4d30-a6c3-1e5b0c1ae666 · inbound

ERFSL: An Efficient Reward Function Searcher via Language Models for Custom-Environment Multi-Objective Optimization (Student Abstract) cites this paper.

ERFSL: An Efficient Reward Function Searcher via Language Models for Custom-Environment Multi-Objective Optimization (Student Abstract) Language to Rewards for Robotic Skill Synthesis

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:08:05.277461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T05:03:28.824500Z digest=sha256:0e90299a3b0d6b8155214deaadde7f567701a187c11382b76cdf3b1134b57e73

Observation 7787d8f2-017b-4a25-9c1f-8a0de03ab920 · inbound

Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations cites this paper.

Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations Language to Rewards for Robotic Skill Synthesis

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.294628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T05:32:17.780012Z digest=sha256:5d835f4d7762fae6acdcdcf876c2b3102f3adf5419c0c756aa2317b28ae22b0a

Observation 517f6d7d-4ab3-4e67-b090-1a83758dcff7 · inbound

PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management cites this paper.

PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management Language to Rewards for Robotic Skill Synthesis

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:23.904827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:17:41.705218Z digest=sha256:792e04959fad6e2b7a72b47d8b2565e03fbac39d1d03c8076250a23809b3113a

Observation 1da2b404-d440-4232-9112-da39b3a41554 · inbound

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation cites this paper.

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation Language to Rewards for Robotic Skill Synthesis

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:56.738569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T01:46:36.081851Z digest=sha256:66c5ad69aedad6066c5192723d977641e62b60dfce0ba2a0db961909547ce2f5

Observation edd7a058-da00-4aae-b0cd-11dd3f2f9eaa · inbound

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World cites this paper.

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Language to Rewards for Robotic Skill Synthesis

Reference 53

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T03:59:33.466912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T17:25:29.469359Z digest=sha256:004d8b6d4de24a6670a65b31a7fbf9f45b083762abb6c197fed10e44a003ab9e

Observation 0207fc36-28ac-46e9-a6bf-9df72ca2ab8d · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report Language to Rewards for Robotic Skill Synthesis

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:36.895230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:8bededdf42e6dfaf530c66cd48ecce38e062d3053c57c3b7673e5f5dc2daebfb

Observation aedced29-181b-4cb6-b901-e68cb4e20aab · inbound

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics cites this paper.

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics Language to Rewards for Robotic Skill Synthesis

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:42.632679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T10:43:35.766229Z digest=sha256:b5792134b99b2d2a1c353ed1f08a840b435ff757b97ef363fc2467b07730cd86

Observation a79d7254-6add-42be-bd93-2719168175a5 · inbound

LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation cites this paper.

LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation Language to Rewards for Robotic Skill Synthesis

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:54.290524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T04:46:12.483085Z digest=sha256:1289fb11799e91067b978d52e1508a1f651db7742280720334d002f0011b8a80

Observation 1fec7aa6-adc0-43fd-9d3a-d0cfdebf69b2 · inbound

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning cites this paper.

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:54.897603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-02T11:52:25.035266Z digest=sha256:43c8d1e5b77893a205e82fd4c345bfeb680277f19f113f361cc2a6f6a85442a2

Observation f293187d-0996-4e81-a3e4-de94ddd6f1bd · inbound

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning cites this paper.

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning Language to Rewards for Robotic Skill Synthesis

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.284136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-03T20:43:34.222848Z digest=sha256:2fb6fcd8b5fc67410d7a9bc2200e66bec140e6cbba6e3c41c26d0792ce1772f5

Observation c5067fdb-cc59-4039-bdef-edb1a229f603 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Language to Rewards for Robotic Skill Synthesis

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:53:50.182171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:7cafa416a8b6998f6650e66a6639937490dad007af0696e8d488c1d53a0571db

Observation 0e5994b9-f181-4ccf-8356-9aeb00c7c3af · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Language to Rewards for Robotic Skill Synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:395f08641c0e91c27c468c55a0ee317962254a5f2d4cc787fc7978033a4a989a

Observation 944e2d6b-e7be-44df-b0b9-35142ba50594 · inbound

Breaking D\'ej\`a Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning cites this paper.

Breaking D\'ej\`a Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning Language to Rewards for Robotic Skill Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:21:26.191346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:21:26.191346Z digest=sha256:b09a52dcbc755e384e0dbf22374c48975742d39b427efa2d14b946922c5fd362

Observation ced2984c-f95b-4194-8157-8f7abbb51cce · inbound

LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks cites this paper.

LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks Language to Rewards for Robotic Skill Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T20:27:55.866736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:27:55.866736Z digest=sha256:fffac576c6fc1d7758613329d9f23923508c78b5946b1b7aa8813d0c022114e5

Observation 9cf112c2-7a81-40f8-a41c-221863c23e67 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Language to Rewards for Robotic Skill Synthesis

Reference 288

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.417818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.417818Z digest=sha256:e2e17eefc02165f46b390313d251282b909fd16f2e7907755f5ad0c6b8d09729

Observation e4c2528b-55c1-4978-9fd4-6a10b37996a0 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Language to Rewards for Robotic Skill Synthesis

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.214411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.214411Z digest=sha256:1cf6310eab6e0e35a53ef4c0ab4fe7b2ba98a14331841c82f065cf0d598e5aa3

Observation be7ad28a-7900-4234-a019-b335c08756fe · inbound

Action- and Language-Conditioned Video Assessment for Embodied Control cites this paper.

Action- and Language-Conditioned Video Assessment for Embodied Control Language to Rewards for Robotic Skill Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.112288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.112288Z digest=sha256:954463853bdb4b31ddada7d184846b766319a8b99cb6c589cade32105b1ea2e3