Pith. sign in

Paper Citation Record · LEDGER

Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2309.11489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.11489 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:13:43.841984Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.316535Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8740558b-4808-4940-86a3-12562da04c3f · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.356658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:492d36ce5b0cc0ca35dec25287f7494649642c29f6cc0d878f568cf11f508be5

Observation 174f6f66-0c2a-45e2-9233-ca45e25201a4 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.816989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:456e392554fecd98d2fb263d325242c8c1ee709a4f9ffb3ae0cd361a812debd7

Observation 869f221a-4444-4f51-bd69-bb2e6a21cab0 · inbound

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving cites this paper.

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:43.841984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:43.841984Z digest=sha256:8dacd52c1ed025e3d06ad418192d48b05236c5a6ab6963c65e911080dbc3eaa0

Observation d0a81003-26d7-43ce-ba62-ae59e4f1d1e5 · inbound

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning cites this paper.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.188530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.188530Z digest=sha256:809ee937021b2568bc81ff83cacf1ba6362dac4107d4bd77485432b39535ab6b

Observation 7a4f9053-a688-482a-ae85-e31501427ad3 · inbound

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions cites this paper.

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:16:14.289687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:16:14.289687Z digest=sha256:1bfad721a6e8e495b1795fd1998f7e9c431a18a70eaa014a6d96691927f77754

Observation faf0f930-f3c5-4630-a678-b54619866c89 · inbound

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning cites this paper.

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:06:57.016648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:06:57.016648Z digest=sha256:c40a065b37e0c7fbc49c9dfb7d8eadd5f9ed277f1ad2808e9237f23e456b2798

Observation 9f5e5edb-2f76-47dd-ab22-4e6287629dcd · inbound

LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning cites this paper.

LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:21:32.862328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:19:30.609974Z digest=sha256:6fa5e89f763d6f57529c024c499c58a1c7ec8a5d98b1ce8a0a52fb92dba6385a

Observation 5fcb6402-7388-4175-8738-e944f9b01c1b · inbound

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization cites this paper.

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:37:26.211102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:37:26.211102Z digest=sha256:dbe6f58d97e053fea8b992a07d154d00633376f7609c4484c456844b4981aa6d

Observation 7b3121aa-73da-422d-9565-dc6e6dfe64e0 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.784752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.784752Z digest=sha256:0f9ed9ff28fa0d9ff3980cc41746472f5a6251dfa18651e2399978c045e1f2f7

Observation 5eddd61c-4bec-4df1-a17e-c6b1746cd718 · inbound

Automatic Generation of High-Performance RL Environments cites this paper.

Automatic Generation of High-Performance RL Environments Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:05:01.520674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:04:11.718672Z digest=sha256:d91926cb967106499104b95bffb8b33dbd2bba76ec799c29f0eb394b12911f14

Observation 8caf378a-ab71-4dc1-b7e9-0f224c756884 · inbound

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models cites this paper.

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:18:15.758972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T23:17:29.025222Z digest=sha256:94e5b957dbbf31c0c18908b0b94a7d0c500762c533e52baa4c018c089c45b438

Observation 38e963a9-50ad-4d1c-95f0-f9b0c1c28d83 · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.206079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:14:22.162163Z digest=sha256:7081c3604ac8b3403a7a8fa5c9618ad5f8acf8f6afe885393eda9582860fa839

Observation 72f94d4b-0117-4a26-8472-e567430873dd · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:57.895007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:10:46.725354Z digest=sha256:06eade210b0140dd9b6204c226ed44ac2b082500d203b43c3d8f2fec4abf3c37

Observation c1a87110-1785-4874-bae1-e61994a20dc9 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:26.415726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:8d4fceb41e94ccb8f5710bb20301b175973567d6e579e3800512739669ad073c

Observation e8e8a514-64ad-41ec-9936-201378a99ba2 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.551200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:69096c258e5063b5cad6b36b7659bc12cb474599e41bf8c4a8321bd68693cda6

Observation 8dbef211-ee88-4f67-8fe5-8d0785d58fdb · inbound

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models cites this paper.

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.909313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:19:38.587352Z digest=sha256:73b21e304695f0fba2243ba002ed16aaf56ed011e443c9980bdcd539d9297313

Observation 42126c6e-4e3a-4f13-a566-4228b2b66bac · inbound

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models cites this paper.

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.377018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T15:25:14.906470Z digest=sha256:67a8347f0b06ca4e51de3d58903e49ec79b583950bf1e6ade04c3e40170d783e

Observation 88e32e35-6223-45a4-a682-cb088f22245a · inbound

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models cites this paper.

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.875284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:04:13.554098Z digest=sha256:9c6ffdc7e9310453179c3f5abbf584ff5bd1c5f375ae807520c57cb3888a7aeb

Observation 6289b284-d96d-4818-964e-f2504f346da8 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.538848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:3cd6920ca9e25a3ad76f8360f43d94b114c7e41b81e0bf23086bf7490cce3db2

Observation 19b326fd-8551-41a7-91eb-26efa587dcb4 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:53:50.318343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:4358941e41839b780f010084c01ba435a5ef5ad8081f22719ba4ed4e10ff7b43

Observation 452ca998-7168-4afb-8c1e-428707d1d97f · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3c416a0d9cf8be358b8eb4960423911a61f1fb84ab92a2fe60afa1120a3f59a4

Observation 78f09d93-52dd-4add-9bea-8dfd62e6d254 · inbound

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cites this paper.

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T03:14:12.941468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:14:12.941468Z digest=sha256:a6f3bedc224ff856f0a0983de794d99d59c1f56b6ebd0a4cebac147f60fb2b8f

Observation f020b07f-da1a-4fd7-9999-e83625f0a3d3 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 280

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.370242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.370242Z digest=sha256:4b996b2c430b5a32439d118fba45852588c78a558e202932945745c809dad070