Pith. sign in

Paper Citation Record · LEDGER

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 18 inbound Pith citation observations for arXiv:2506.11425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11425 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:16:08.504844Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:47:59.075366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:55.557008Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e8a4e8c-3519-4492-9a04-355aaf229ffb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.416963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.416963Z digest=sha256:ada3d9d0ffbebe7dd9a3cc0173d499746013d2f6447dd17b9ff674b2ab9b7fd2

Observation 2be21a07-10c6-4a7e-8a2c-de4d7ffe1fbd · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.792345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:16:08.420827Z digest=sha256:07a0e1f3c0903796628e7399d117e9ac8b31524778aac9bf4a2a24fbc15149ac

Observation 26e9f587-880d-4a73-af48-7edce2138133 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.782440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:16:08.424091Z digest=sha256:ac58ff38b8fad45a089a365333c35aef0cd0840557630d706b8b831cfbfda658

Observation 9ef3b7a6-547a-4159-8202-4cc7b3f3955b · outbound

This paper cites Training language models to follow instructions with human feedback.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training language models to follow instructions with human feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.427531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.427531Z digest=sha256:6718254e7a2847a5cbb38355511955e5fe138993e4374eaf92ad5b43e0a121be

Observation 52a0bac0-0a79-48b3-9954-97d235ffae56 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.431125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.431125Z digest=sha256:16be9bdf6b2b0ad274ee9e9f2aedf7bdaed71963ae78708a6b707a29f600c332

Observation 648ece32-822b-46ac-bbb7-c8210ea004b3 · outbound

This paper cites Program Synthesis with Large Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.434730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.434730Z digest=sha256:b24901ceed1b419012b1f8c910b6776c8b96c5da5e18305c1f6e3e73e379603b

Observation 296b8740-ec45-45a5-ac84-fbdb4c8ad79e · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Measuring Coding Challenge Competence With APPS

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.438286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.438286Z digest=sha256:e617bd12e67301d0592ffdb5d4e304cddbf2dde4bb086a72a389ea8bf3c37a59

Observation d0d8936d-8db1-40fc-b6db-a3e511e0c6b0 · outbound

This paper cites Planning In Natural Language Improves LLM Search For Code Generation.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Planning In Natural Language Improves LLM Search For Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.441362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.441362Z digest=sha256:00efb58a6138149dc3f2013c5d4651d14487ff240ff3996415b3ecf6174c5fd8

Observation 4315de77-c1db-4f60-8992-5305968fe354 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.444325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.444325Z digest=sha256:fab50606bacd2e01de3481803606a07398673cb2af89ef20578d5d16f1796f2d

Observation d9d2ea16-8ced-45e5-b3b9-8b4c7e20e7a4 · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.447400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.447400Z digest=sha256:a133ac047a2749cafd9dfe9bf5b0c483440559be2f62bd6afba048b04eab4b97

Observation ca8a7a1b-1699-4885-a4f5-e1e89984aa7a · outbound

This paper cites STar: Bootstrapping reasoning with reasoning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards STar: Bootstrapping reasoning with reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.450383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.450383Z digest=sha256:dba8522055da4580fa27710669c47187c850a544d597eead3ba6df26da311a40

Observation 3a744ce8-b809-4fce-8d52-db70b358bad4 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Reinforced Self-Training (ReST) for Language Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.453021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.453021Z digest=sha256:ef27be7b9c6270e071018de9b43591c806526346f494ea54ce9bf9f2bf01e3bc

Observation beb20196-bca5-4497-ad80-500001be350b · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Agentless: Demystifying LLM-based Software Engineering Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.456160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.456160Z digest=sha256:b0cae95c9e915ff84171bd2306d0f4b091bfa3fb94597a50083c1923837d7077

Observation 1326b1c1-5622-47b3-b55f-f8146d15cacb · outbound

This paper cites Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.767942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:16:08.459059Z digest=sha256:07277f829313a71de5a21c3c85eb6984a842d17f55b4b3c827fbd64489cde0b6

Observation 0d47d712-a232-45d9-8689-ce8afc3db309 · outbound

This paper cites SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.461930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.461930Z digest=sha256:6984e1a0a39410cef3275f81d70c73b90acbfbdeca099918f75a4261251b2bdc

Observation 23dfffdb-63d4-418f-819d-d61496ea7326 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.464749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.464749Z digest=sha256:76afc0a470bc89caa6236d2c9029fda8b38acc334b1bddc76e09a390401d67aa

Observation 1925b323-2750-4627-87d1-38c89723b7df · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.467524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.467524Z digest=sha256:e0c426a95d16f317c0410add70aa8ec0607b1ce30c780ff961d60e558c312bd6

Observation 1e6cc14a-eafc-44b5-8a91-32cb06d268a9 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.471002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.471002Z digest=sha256:14bb8ecff3fc0f60223fa7befab4c0fa4cfe2d503872e07e2986647329bef898

Observation 28a93968-05ef-4510-b354-23379f10466e · outbound

This paper cites DeepSeek-V3 Technical Report.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.473951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.473951Z digest=sha256:dd6fe6b99ce77394cc3a93d881390779ec694b766a47dc2e585396031b4ee8c3

Observation 9c9ef027-3deb-4379-bb19-94e1e6950ccf · outbound

This paper cites Commit0: Library Generation from Scratch.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Commit0: Library Generation from Scratch

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.477293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.477293Z digest=sha256:3c899653e7fe971dbf3449e00cadd42d7b246f9257b2a0220204312754cbaa6a

Observation ccac6345-da87-4ac5-a2ed-e3325ea68349 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.480495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.480495Z digest=sha256:553947277fffcb18b5671b9ec3dbaa5d591500a28446688a864b7bdef03a386f

Observation 76698c99-b65f-464a-9946-d155b06aff8d · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ReAct: Synergizing Reasoning and Acting in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.483469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.483469Z digest=sha256:d86556783d1763cc3d80d3918f57b96923a3af334f9647ff40b094e958185452

Observation 110172c0-7cf8-4510-8de3-9f3642ec23ee · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.486283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.486283Z digest=sha256:6cd2ef4b46fd9c7f1106a0f03a3e7f09e47b11ffac5ff8757bdbbea4a01803d3

Observation 228809fc-4eca-4d16-96c0-1a1a8348cfe5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.489630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.489630Z digest=sha256:f72e632730b47aca1f3911f9de39c104dbc0b030363969ac961b832598c0c05a

Observation e082b267-d4f3-424e-adf6-f4a1d9132bfd · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.492855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.492855Z digest=sha256:cb079bbca9619618c93fcd2d30d07887fca8b76beffd73dcad8e141333a8c904

Observation c4f109a2-c72f-48c7-b4d2-63befd7472aa · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Mind2Web: Towards a Generalist Agent for the Web

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.495825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.495825Z digest=sha256:918aee87af3af49fe509f3c8cfaed72eea2696dfd35024153266c170ff303dcb

Observation 05cff3f7-a161-4d6c-9da5-72e6cb9872cf · outbound

This paper cites ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.498785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.498785Z digest=sha256:8aafa8d2af0ace2a149b7591374f63aa31aa7b99813f7feef6f0c459474e6327

Observation 83c58c56-16f3-4c58-97a3-5084691e6531 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.501801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.501801Z digest=sha256:02ea424b0fc693c0b063c1cf7f2c3293f16460c3896e4e4f2914c853834492e3

Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.504844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.504844Z digest=sha256:d142c2a3afb74264fb91f8f120affd645c763c1cf22fc2c3a42b9a5a5b14dc99

Pith citing papers

Observation 781e070c-9bf1-4920-951e-e58502d4508d · inbound

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? cites this paper.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.748215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:e60834fb7aeb385ecea6d74b3c3bb5938486c683e17959c2054720528afbeb54

Observation 5b423429-604d-4cd8-81b3-3549e256206a · inbound

SWE-IF: Aligning Code Evaluation with Human Preference cites this paper.

SWE-IF: Aligning Code Evaluation with Human Preference Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:02:07.533516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:02:07.533516Z digest=sha256:7e4de5c60d5025a8939e9b5a4e44cb661427e00e34c419bcb3c13494af58a955

Observation be23bde6-0887-4ca3-8224-7e9394f1323a · inbound

SERA: Soft-Verified Efficient Repository Agents cites this paper.

SERA: Soft-Verified Efficient Repository Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T07:17:13.394276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:17:13.394276Z digest=sha256:4f606a87ff6c45c38462d4d4579e99eead5e5a07cb623efe7bb58407c41f2f21

Observation e5d51157-6f35-401b-8ef8-9bbce2b3ba2b · inbound

Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool cites this paper.

Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:37:45.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:37:45.905230Z digest=sha256:45e1ec63c429a12f0c7c260262de558045ed3907c74437d2a71426489ae69bae

Observation 185f9dde-b76d-43cc-aa4a-48b40b711c9d · inbound

SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents cites this paper.

SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:59.385102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:21:12.961613Z digest=sha256:5ed1a4e0173eee85e918c4dcf21e4551787c7216cb4e649d0fde5b938fe6d511

Observation b3de5738-ffac-4b63-a2f8-55a9dce22fc5 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.121327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:7c739e009f202384ac47b1d63b524ff45d668e4c3d85b808f61b8dbd30e301a0

Observation 740bad67-0768-4658-9cd9-47c50619a692 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:54.783701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:5d0640e5c95ea6dbdca5bde93a42e98e52f33570b02f95abd3e7472bbf6ad337

Observation 7bfd02ae-5fcf-4231-b1ec-fc33107928ff · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:43:45.447084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:8765f1d90ba352dfb7f07cf9f99aed227b7eadbb5a14ce5837043b1bf64b2201

Observation 3a167a8b-7dcc-41da-90bf-ab85eb4d7060 · inbound

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL cites this paper.

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.921429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:49:40.239199Z digest=sha256:6bc833ae0458279dca28ca00a9f4cb26e1a195e839a4125bade127491ef6970f

Observation 076142ad-da01-4de0-a46d-bc4db2f109f8 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.212323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:e52e006b8b03d0c4ab20c0a0cabf385973a4ac181d7b04f5bcec777ed748e91b

Observation 9c15bb12-a2e7-4b9f-8422-7f29e39538a1 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:0016ed6eb9da994d3345a25ace404c8933392d4b9a7054de30ba2dafbff69baf

Observation b55f1ba7-cf7d-4b55-b1b0-60b03a905c3b · inbound

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation cites this paper.

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.320110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:03:54.185019Z digest=sha256:0225c55557a4c0345fb1c88f47694ad5a9218de5b985cb103df93c971c21037a

Observation cdbe7401-d4c8-4938-b932-ec10e1591b6f · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:33.180739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:e46207d276c60b0bb364a144b63ba20a82901445ab6ae887149c068112d7b487

Observation 8f126582-5c6b-4f46-86c0-ef9ee3d9f93d · inbound

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows cites this paper.

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.558939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T20:18:07.134598Z digest=sha256:25490f4b06be1930b36d819dd434aebb4ca21fb89ccfb40cf241fd39012ba9f3

Observation daced041-f2dc-4f63-8d6b-cc3dd8ae344b · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:d37d56efc22cb3f342472159cabce415cc329cb8eacc953802b7a5908b8f4731

Observation 6bd6e6ce-b523-4888-a300-1909ca46bcc4 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:10.280348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:10.280348Z digest=sha256:989467c3cba2bb4958d0155b28c25006ba1fe30584c510421199fd3b6af9c100

Observation 003a6523-95fe-4860-8329-fc4f1bbafd50 · inbound

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning cites this paper.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.759049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.759049Z digest=sha256:17544f19f337f0e08a8d7b139e120fd65f13923ef8b5bc44db9fb3be1d26b39a

Observation a67c7176-23ad-40d9-a1f1-c3825501dc7b · inbound

Self-Evolving Coding Agents cites this paper.

Self-Evolving Coding Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T19:47:59.075366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:47:59.075366Z digest=sha256:2b75c0a5aa15df4738490bbeaac923677afc2bd579b7925192978ca54058c0aa