Pith. sign in

Paper Citation Record · LEDGER

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2504.16078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16078 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:12.272937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T22:46:53.836788Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2cfbac60-fb60-47ee-984a-50b2db981f06 · inbound

Formalizing Learning from Language Feedback with Provable Guarantees cites this paper.

Formalizing Learning from Language Feedback with Provable Guarantees LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:19.638204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:19.638204Z digest=sha256:952b00dec19d0b68dd535fd016a68945b2bd44cf644a451b58883f379d786885

Observation 4c83c026-a8b5-43f4-a14b-5811fa35562b · inbound

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency cites this paper.

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:46:53.839788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T22:42:53.663354Z digest=sha256:f75fb37e3c79ef16dfdc570faac4d401583c759f4c560b67bebb3f8a8fb7af03

Observation e54cfaec-b1ea-4af7-98bb-e3953fd075ed · inbound

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format cites this paper.

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:10.915427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:10.915427Z digest=sha256:676ff6bf013876e527e52d270239669c63e991362e622fe7186f9ccf6e526699

Observation 7f3f4b64-5821-422e-aa76-1762997041e2 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.260995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.260995Z digest=sha256:7e2fb9f924d5de569ff5266319b0620ef18ef0861e712c3bbb8d5a5471040e0f

Observation 0d15c79b-652e-4827-8e3f-80ee1f170659 · inbound

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits cites this paper.

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:51.144607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T18:32:12.396568Z digest=sha256:27defe67f787635a685b496c6dd12dffc872b2f2d6af1b08de1a0ef6e7d28a93

Observation 873fc6ac-4908-48fc-8314-28b34b99ae4a · inbound

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging cites this paper.

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T09:17:00.467000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:17:00.467000Z digest=sha256:4368503c2710376239919b1b6fca16db9c038406ea36820176e3daccfb6493af

Observation c59bc3ed-e23b-4c20-98d8-65a1280885ea · inbound

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle cites this paper.

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:06.208989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:06.208989Z digest=sha256:3ea5e1246c7eda854d72575fb87a16d761483a72fb00a55bfb852847f65b3118

Observation ef0fa664-70fc-401c-a58f-fe810392e29e · inbound

AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization cites this paper.

AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:32:47.083166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:32:47.083166Z digest=sha256:4a3cfe9a3bd2464e5749dbc8974d843e89cb18831545128158ec42b3daf62655

Observation d9ba0ed0-8143-432e-bbba-0fb40675d299 · inbound

EvoRIC: Reinforcement Learning Fine-Tuned LLM-empowered RAN Intelligent Control Toward Autonomous O-RAN cites this paper.

EvoRIC: Reinforcement Learning Fine-Tuned LLM-empowered RAN Intelligent Control Toward Autonomous O-RAN LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:12.272937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:12.272937Z digest=sha256:622f1bbb242b1d0400360cd49055c4bd779b720a4a0fef6d1345dfe992db845e