Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2303.00001.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.302130Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:47:38.286725Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d23a82d5-3c8d-410b-b9e7-3dd456ea4d2c · inbound
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Reward Design with Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eedeb1d1-d01b-488b-bf83-d900cfbef269 · inbound
Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games Reward Design with Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20383376-18c4-4a35-860a-eda6c56fc415 · inbound
Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning Reward Design with Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c8dac2-a844-4e4f-a062-14f1f34ff23c · inbound
LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation Reward Design with Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89208768-7722-478d-a081-5d8dfc6c7f02 · inbound
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3 Reward Design with Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd972443-e59a-40ce-b06d-3303f40eea88 · inbound
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively Reward Design with Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49e6e44-5eb9-430d-b0db-fbe209d244c6 · inbound
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de54f59e-6712-4905-8a83-7fe6714d6b9d · inbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reward Design with Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6cf36c-3d6b-4d3e-a5a5-6284597625a8 · inbound
ACTLLM: Action Consistency Tuned Large Language Model Reward Design with Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add81e89-0e28-4c1e-b6a8-8551fdc962c9 · inbound
Prompt Informed Reinforcement Learning for Visual Coverage Path Planning Reward Design with Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c00c216-2a6c-4e00-bfd4-e78c84b992a8 · inbound
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making Reward Design with Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e0d3f5-5f7a-4d3d-a8d7-a3ec2544814f · inbound
Towards Reliable, Uncertainty-Aware Alignment Reward Design with Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158fe600-ef6b-423d-b6c2-3bc5a8a134f2 · inbound
Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN Reward Design with Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f305a4a6-4972-4115-b611-ad80f1e9c0ef · inbound
GPLight+: A Genetic Programming Method for Learning Symmetric Traffic Signal Control Policy Reward Design with Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ebd632-382e-45fe-8ad4-a5484af427e8 · inbound
Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions Reward Design with Language Models
Reference 232
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d87d1aa-846d-47c2-8241-a9f836e426a5 · inbound
RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation Reward Design with Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53f15e63-da33-4bec-97ac-afdc54c976d8 · inbound
Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning Reward Design with Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d415c79-628b-4620-92da-edc874336346 · inbound
LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning Reward Design with Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c805b36-670c-40b6-a7f6-f052dc25baf7 · inbound
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization Reward Design with Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587ece49-1ab1-4654-ac28-0f2667bb4dd4 · inbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Reward Design with Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7182a7d-daf9-4bc4-82b3-89d1677e696b · inbound
Debate2Create: Robot Co-design via Multi-Agent LLM Debate Reward Design with Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebcf14d4-251c-4889-8f75-1e163c19b504 · inbound
What Is Preference Optimization Doing, and Why? Reward Design with Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dffd4cd3-bcf8-45cd-b02b-63d131037c6a · inbound
From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning Reward Design with Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2cb6821-a8dc-49f6-adeb-643f1b2a32b6 · inbound
PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs Reward Design with Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95fafb7d-3014-4d2b-b6dc-572abe964758 · inbound
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Reward Design with Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfbba3b9-0fd3-4fa0-86ff-dd484a6ff85b · inbound
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Reward Design with Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a11363b3-61f5-44ea-97ff-02d463e3429f · inbound
Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias Reward Design with Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ee0afc7-d13a-45fd-8fc6-79051e300cb8 · inbound
EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models Reward Design with Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a53cd80-346a-4d58-93e6-720fad5239bf · inbound
DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition Reward Design with Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19b9b349-5bd5-431a-80e2-fc59e6710367 · inbound
Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning Reward Design with Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c880b2a2-7d08-4386-8ba3-647b6b3461c4 · inbound
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Design with Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6ae1a8e-4136-4dfc-aa6e-dccf1d2d497b · inbound
ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation Reward Design with Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 29c5a56c-d257-4a05-9423-3de842757ff7 · inbound
TAPAS: Throughput-adaptive Perception for Autonomous Systems Reward Design with Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 913bc2cf-6378-4779-8888-177241f91b5a · inbound
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives Reward Design with Language Models
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.