Pith. sign in

Paper Citation Record · LEDGER

Reward Design with Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2303.00001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.00001 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.302130Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.286725Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d23a82d5-3c8d-410b-b9e7-3dd456ea4d2c · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Reward Design with Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.564735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:72bdacaf178d61535a0b8facd750d95a091748db8e1dd18accd3fdbeb77fa512

Observation eedeb1d1-d01b-488b-bf83-d900cfbef269 · inbound

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games cites this paper.

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games Reward Design with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:04.302130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:04.302130Z digest=sha256:0a2f213e683aa00aaa661e008a5078e4be86c60f2c69c702eaecbf6718c94d01

Observation 20383376-18c4-4a35-860a-eda6c56fc415 · inbound

Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning cites this paper.

Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning Reward Design with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:15.052114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:54:15.052114Z digest=sha256:97968e5d2b870b701727d99e016e111e8872a0cca3e856585d5647cc52e15ae7

Observation e2c8dac2-a844-4e4f-a062-14f1f34ff23c · inbound

LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation cites this paper.

LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation Reward Design with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:38.843038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:38.843038Z digest=sha256:b37c543e5c836453622efea0e127ab09e7194c63f05ec98b76b913c21dd3e327

Observation 89208768-7722-478d-a081-5d8dfc6c7f02 · inbound

An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3 cites this paper.

An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3 Reward Design with Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:54.004545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:54.004545Z digest=sha256:ee4cc32646dd41d451a4c5d3afaed97e80052f076640eb71146e05ece1ff57a9

Observation cd972443-e59a-40ce-b06d-3303f40eea88 · inbound

Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively cites this paper.

Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively Reward Design with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:43.508406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:11:43.508406Z digest=sha256:7b0efd40db2712945a8c6d70302effdcb5b4dc4c8297fc5ec8185b22c4daae6c

Observation c49e6e44-5eb9-430d-b0db-fbe209d244c6 · inbound

Agentic Episodic Control cites this paper.

Agentic Episodic Control Reward Design with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:00.420509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:00.420509Z digest=sha256:c609fde4668983909fa79acd719e02766b70c81dc2b604eeae2908950ff25e2e

Observation de54f59e-6712-4905-8a83-7fe6714d6b9d · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reward Design with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.973403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.973403Z digest=sha256:b210f08aef1d4820b94d15986234956d004ab9352d12d15335590ac516d02c92

Observation 0e6cf36c-3d6b-4d3e-a5a5-6284597625a8 · inbound

ACTLLM: Action Consistency Tuned Large Language Model cites this paper.

ACTLLM: Action Consistency Tuned Large Language Model Reward Design with Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:08.994748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:08.994748Z digest=sha256:d7b019226eefe3adca55f35a76d4f2448bde73c59a6807860550a86a12a5aadd

Observation add81e89-0e28-4c1e-b6a8-8551fdc962c9 · inbound

Prompt Informed Reinforcement Learning for Visual Coverage Path Planning cites this paper.

Prompt Informed Reinforcement Learning for Visual Coverage Path Planning Reward Design with Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:45.103900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:45.103900Z digest=sha256:ff4cc29f9e9a65f34f62e0db5c0c8aa2b71b5b1264e461bc99947899207053ee

Observation 1c00c216-2a6c-4e00-bfd4-e78c84b992a8 · inbound

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making cites this paper.

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making Reward Design with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:12:28.663704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:12:28.663704Z digest=sha256:91f70b2458a1b07392e42c977a3e614e95d9a862aee321d4df0cb051535e2367

Observation 33e0d3f5-5f7a-4d3d-a8d7-a3ec2544814f · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment Reward Design with Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.257918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.257918Z digest=sha256:de90237849e00eea2e8ef369463e9c3a1359934c118373be5e85f1a427c073ed

Observation 158fe600-ef6b-423d-b6c2-3bc5a8a134f2 · inbound

Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN cites this paper.

Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN Reward Design with Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:30:39.980782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:30:39.980782Z digest=sha256:68909fa3c031fa8a8ce77ebca22569cb5aa37ccddaab0bc73a9f15f09efea216

Observation f305a4a6-4972-4115-b611-ad80f1e9c0ef · inbound

GPLight+: A Genetic Programming Method for Learning Symmetric Traffic Signal Control Policy cites this paper.

GPLight+: A Genetic Programming Method for Learning Symmetric Traffic Signal Control Policy Reward Design with Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:36:53.242484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:36:53.242484Z digest=sha256:3ea4db2aa1da47ace0bf69ac645e37bddd6796958066bfc093465a0d8ca9a3e2

Observation 68ebd632-382e-45fe-8ad4-a5484af427e8 · inbound

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions cites this paper.

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions Reward Design with Language Models

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:00.388060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:00.388060Z digest=sha256:2f4e23fb878d088598e2bd8eeb6a9949232731f28d58b0983c288a9d893a0c71

Observation 0d87d1aa-846d-47c2-8241-a9f836e426a5 · inbound

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation cites this paper.

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation Reward Design with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:39.757357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:39.757357Z digest=sha256:4c09b99c48c49a73d6b84a4f149f3fb45cc871118f016b2ac943138516b2ef42

Observation 53f15e63-da33-4bec-97ac-afdc54c976d8 · inbound

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning cites this paper.

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning Reward Design with Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T16:06:58.488030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:06:58.488030Z digest=sha256:b16fb1ebfba6933b593f93e22faedb2c66b9d73fa6fd2f7c7543a7a76d81dec8

Observation 8d415c79-628b-4620-92da-edc874336346 · inbound

LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning cites this paper.

LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning Reward Design with Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:21:32.889019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:19:30.609974Z digest=sha256:e31b98255b1e7657a4059a54a1413bbe112a5721f5ccd3d81029f6f85f31f40c

Observation 0c805b36-670c-40b6-a7f6-f052dc25baf7 · inbound

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization cites this paper.

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization Reward Design with Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:37:26.093669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:37:26.093669Z digest=sha256:b74f036c090790760c1396ccc0247ee3e5fa97ea2d12b5194f1d0f0ed4030ee2

Observation 587ece49-1ab1-4654-ac28-0f2667bb4dd4 · inbound

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training cites this paper.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Reward Design with Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:42.928860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:42.928860Z digest=sha256:8f9beb742ab6e5e06d3292bbc76e88cee13d7eefe60c895e7bb50c4833c13e8e

Observation e7182a7d-daf9-4bc4-82b3-89d1677e696b · inbound

Debate2Create: Robot Co-design via Multi-Agent LLM Debate cites this paper.

Debate2Create: Robot Co-design via Multi-Agent LLM Debate Reward Design with Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:27:03.942685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:27:03.942685Z digest=sha256:dd5cbd5b4f64f64abd10d408aceeb223965c728f469e93b8cc4b029e77cab5d9

Observation ebcf14d4-251c-4889-8f75-1e163c19b504 · inbound

What Is Preference Optimization Doing, and Why? cites this paper.

What Is Preference Optimization Doing, and Why? Reward Design with Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:40:28.991254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T18:37:48.161545Z digest=sha256:21f8f0e781a40242bac25440c0da8c0ea6c9d232991b5c70f0863fb323db262d

Observation dffd4cd3-bcf8-45cd-b02b-63d131037c6a · inbound

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning cites this paper.

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning Reward Design with Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:58.248439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:35:36.083682Z digest=sha256:ea4b3073e244e8bf6de580b7156c520480108484334583cc59c2093061270547

Observation e2cb6821-a8dc-49f6-adeb-643f1b2a32b6 · inbound

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs cites this paper.

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs Reward Design with Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:42.642392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:09:07.557773Z digest=sha256:9e550578b9c3a1ebe05f411f2f4930a7fac9fab23106bc105bde56156da334a9

Observation 95fafb7d-3014-4d2b-b6dc-572abe964758 · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Reward Design with Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.202959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:14:22.162163Z digest=sha256:0accd4a6e4799a62d1d6c0b1f41974862bdd3f374fd9f0c9034579eb92eca829

Observation dfbba3b9-0fd3-4fa0-86ff-dd484a6ff85b · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Reward Design with Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:58.007955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:10:46.725354Z digest=sha256:531b139c6d44e66be12a5f10382488f68915125419a04a6c6e9a349333a5c53d

Observation a11363b3-61f5-44ea-97ff-02d463e3429f · inbound

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias cites this paper.

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias Reward Design with Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.007336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:51:58.315109Z digest=sha256:0b43fde49787420af7e51f410740063faed617eda95fe2d5a5a33e271506a7d1

Observation 5ee0afc7-d13a-45fd-8fc6-79051e300cb8 · inbound

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models cites this paper.

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models Reward Design with Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.924884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:19:38.587352Z digest=sha256:18dff397a49ecb48f43d7fb9f5a74aedd8d39c8c152f228d7b555ac219812495

Observation 1a53cd80-346a-4d58-93e6-720fad5239bf · inbound

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition cites this paper.

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition Reward Design with Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:33:12.905500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T10:30:48.957568Z digest=sha256:192a8a0e1987386826365ed4a196c09b6058d5764fba05b1ca4077701db04e3c

Observation 19b9b349-5bd5-431a-80e2-fc59e6710367 · inbound

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning cites this paper.

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning Reward Design with Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.227147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:47:16.737433Z digest=sha256:68e0b85f037759b7a86a0c75f994d79f4b505716cc225388f2b2daf1e411eeac

Observation c880b2a2-7d08-4386-8ba3-647b6b3461c4 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Design with Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.288260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:428094ae36789acd554dd93eb9e41cbd963fefa0d2a4264073d927351a408210

Observation e6ae1a8e-4136-4dfc-aa6e-dccf1d2d497b · inbound

ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation cites this paper.

ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation Reward Design with Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.642435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T05:59:13.483829Z digest=sha256:15487f22c8bee0f813961d8545e7693e2eac32214e999f7905a2f751d67e6b53

Observation 29c5a56c-d257-4a05-9423-3de842757ff7 · inbound

TAPAS: Throughput-adaptive Perception for Autonomous Systems cites this paper.

TAPAS: Throughput-adaptive Perception for Autonomous Systems Reward Design with Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:24:05.608019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:24:05.608019Z digest=sha256:36884b33bfa6732e1678a6bee00cbbdd40cc0659da3c37909791dc630b0b10d9

Observation 913bc2cf-6378-4779-8888-177241f91b5a · inbound

Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives cites this paper.

Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives Reward Design with Language Models

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-01T15:27:06.396909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:27:06.396909Z digest=sha256:1873a3fb60d4529a1edc6e11ae80d04a00686370f3b7c6d2b84926bd93c96e72