Pith. sign in

Paper Citation Record · LEDGER

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.16078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16078 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:39:19.638204Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T22:46:53.836788Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2cfbac60-fb60-47ee-984a-50b2db981f06 · inbound

Formalizing Learning from Language Feedback with Provable Guarantees cites this paper.

Formalizing Learning from Language Feedback with Provable Guarantees LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:19.638204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:19.638204Z digest=sha256:0d3088d7dd72108b92d238c7957cb1e3348b540959f100d0b659041afdb749c7

Observation 4c83c026-a8b5-43f4-a14b-5811fa35562b · inbound

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency cites this paper.

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:46:53.839788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:42:53.663354Z digest=sha256:21977488c7be87dec7c34f1ea7dff5800acbc67bbe9c76be030736adae92665d

Observation e54cfaec-b1ea-4af7-98bb-e3953fd075ed · inbound

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format cites this paper.

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:10.915427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:10.915427Z digest=sha256:e482b9c9e02ce88cdf2e763b217a4bb243330d1f8c7f356d00f80d48a5eb20aa

Observation 7f3f4b64-5821-422e-aa76-1762997041e2 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.260995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.260995Z digest=sha256:7a1739a69b707cebf1785d64ff5181238e96744368f54609e4a557edf6faa3c9

Observation 0d15c79b-652e-4827-8e3f-80ee1f170659 · inbound

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits cites this paper.

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:51.144607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T18:32:12.396568Z digest=sha256:5d860460ec62e9ec957f1aa150465ca4ad7da60d1c112c8c405b5dc6164ccb0d

Observation 873fc6ac-4908-48fc-8314-28b34b99ae4a · inbound

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging cites this paper.

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T09:17:00.467000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:17:00.467000Z digest=sha256:4368503c2710376239919b1b6fca16db9c038406ea36820176e3daccfb6493af

Observation c59bc3ed-e23b-4c20-98d8-65a1280885ea · inbound

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle cites this paper.

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:06.208989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:06.208989Z digest=sha256:212192cdc2e3c777680ff912c7af1d42daf9702b4847d81f5452da38ae9dce28

Observation ef0fa664-70fc-401c-a58f-fe810392e29e · inbound

AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization cites this paper.

AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:32:47.083166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:32:47.083166Z digest=sha256:2099820b397213b289155525b2de07ebf5cbdca40998c426a08435c1db344c6e