Pith. sign in

Paper Citation Record · LEDGER

AgentRM: Enhancing Agent Generalization with Reward Modeling

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.18407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18407 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.996079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.639558Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 75960073-c17a-4b22-a1eb-b9bd92624d67 · inbound

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning cites this paper.

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.996079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.996079Z digest=sha256:6926362479ac2c68013a16915412c2b1a35d82ff3dd854906d1bf1d98ddecad9

Observation cf9fdf4f-a4ba-4e57-b9c9-f3686dbf2bf3 · inbound

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents cites this paper.

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:15.862250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:34:15.862250Z digest=sha256:bb8b7fa805b83d664548bfc4ec1ee082c822266af58bfd0d65a4d26294179992

Observation b0bc5883-362f-47e2-b5bd-e51e120a841e · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:22.078905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:22.078905Z digest=sha256:d1c3264582ac185deff563332773505286b2a1c66367e9d2cb159dc4b0bf77b8

Observation cc3c6853-25fa-4ff0-a0d5-6bd34bfbdd3f · inbound

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction cites this paper.

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:40:44.032643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:38:37.818623Z digest=sha256:bcf2ca162a47423c845eb911f6e139c716639c760f23425b9e7dd3de487628af

Observation 7d999d4c-8de5-471b-ae07-425c2ea28736 · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.641467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:90a8ae7bd4315bc09d8cea7f0f1f33b892cd6f739ef44de351ca073f0f68fae3