Pith. sign in

Paper Citation Record · LEDGER

AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2406.04151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04151 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:02:57.241810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T04:04:29.283127Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4192876-0ea8-429f-9da7-a155f72cff9c · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:42:04.322581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:015bca7de14ebb2b2baba81a310ca7c77088c37b1e590281711a7d641d77f698

Observation 1691a1df-8419-4c99-bfe4-080641d630aa · inbound

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision cites this paper.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.241810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.241810Z digest=sha256:bf09009a2a09a246eca702bb9ce027104d71d7556d851907ee18b2f880b3ebb1

Observation 03f50d95-01dc-4e97-b7bc-be4b1cc6ff17 · inbound

AgentRefine: Enhancing Agent Generalization through Refinement Tuning cites this paper.

AgentRefine: Enhancing Agent Generalization through Refinement Tuning AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:24:43.127272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:24:43.127272Z digest=sha256:69a09cf54a86c0df5701caf8abb3e4203590b554ddfb23e273157ff9ebf1f711

Observation abe473b7-3f9b-41da-891e-17fff6b0a661 · inbound

PoAct: Policy and Action Dual-Control Agent for Generalized Applications cites this paper.

PoAct: Policy and Action Dual-Control Agent for Generalized Applications AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:52:59.090718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:52:59.090718Z digest=sha256:9928728bf697a89735be0002b463699938133ad9bad95381137bceb5797fb9a9

Observation 8d393d12-5068-41cb-a44d-1b946843cb2f · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.369509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:9ff45b6863cffc0cda35d759fbe967b1dbe437f566125bafbcab839c77141576

Observation e4e86b9a-3daf-4ab3-b540-ec9c3dac5f4d · inbound

Reasoning Language Models: A Blueprint cites this paper.

Reasoning Language Models: A Blueprint AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-10T18:36:55.243006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:36:55.243006Z digest=sha256:07d89ed83cc79bf42c047aaa875e8244805aa54910c0aeb6c4fd43bf46aea80f

Observation cf3f601e-e429-4d2c-9ab2-ec4c26db51de · inbound

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training cites this paper.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.685354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.685354Z digest=sha256:edf320d6bb7f8e10a85bfaf5b37a83f2fb9e90d02e264fb0ca7b7ef04a48b7f7

Observation eed582d5-f821-472c-ad03-5c6ff4c80787 · inbound

Meta-Prompt Optimization for LLM-Based Sequential Decision Making cites this paper.

Meta-Prompt Optimization for LLM-Based Sequential Decision Making AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T18:00:22.211991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:00:22.211991Z digest=sha256:8879a9860193bd7ee3c6439d34dd0a407b377331cce340d4c76239b4a0c5b7d6

Observation 721aa724-5544-4a75-a6b6-cd19e0770c73 · inbound

Large Language Model-Enhanced Multi-Armed Bandits cites this paper.

Large Language Model-Enhanced Multi-Armed Bandits AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.441317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.441317Z digest=sha256:1e28edac6df775f6d855fe47d230839b10a52a349c3b8a6e1ae624934b7dfa9b

Observation 8f1e490f-ee74-4622-a051-9943c5529764 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 280

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:10.427911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:10.427911Z digest=sha256:bb5e1859bcc00ff3c0859fa1e13230a8f076cb830cfd5990caffc7db71aaddb9

Observation 316f7091-e032-42dd-ac81-1b3f3a273f74 · inbound

Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking cites this paper.

Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:57.158038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:06:57.158038Z digest=sha256:242334b572b298ee80ecec4e80530c08c1eb863fe4d027748204b85d5cbce656

Observation a684049a-f95e-4239-8b00-b51f5d2c69a8 · inbound

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization cites this paper.

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:51.170109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:47:51.170109Z digest=sha256:5c77ae2dac507fffce217b428d81f5948b787d2fe2e9f309a53cf99ff1b1da2a

Observation c81ab498-e158-400f-84ac-8be0fcd0581e · inbound

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems cites this paper.

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:36.698999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:32:36.698999Z digest=sha256:e60255518b6559e12e04a29ad7f72a1cd22815065da82024609d6765ae1d8b5d

Observation 9312c6b6-0e23-4cd2-bc03-6d4bc44f2aa1 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.850713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.850713Z digest=sha256:374c6217bffea97704b909598dd5fbcf1528112af646ae5357ee2e4b971d68f4

Observation e3d9126a-6041-461a-87b0-bc33f2336a47 · inbound

Agent Exchange: Shaping the Future of AI Agent Economics cites this paper.

Agent Exchange: Shaping the Future of AI Agent Economics AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:05:43.921026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:05:43.921026Z digest=sha256:cad376014e16c42c27252da19d68bfc0053c379ddfd3d446f4a916f738436a74

Observation 68c40d27-ad2e-4f13-ad10-91e3b1feebb7 · inbound

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents cites this paper.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.211465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.211465Z digest=sha256:298e597b0ceb750fec02e20b6ed513afa224af09489f0021afff3f768fd595b8

Observation a82720a7-a55c-4045-91ef-b2fbf37e83ef · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:16.660326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:16.660326Z digest=sha256:f7cca081a32fbb7a9d98e798b7b1eaff37776edb5d0c222e243a53e6a4db1ea7

Observation 633d89ad-8345-4656-9680-37d45a7c6df4 · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.538990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:17562508b69bd1e44de40a60a0e1c9fc3a686af6e5231d920e0626b413465837

Observation 861e5531-b915-4017-af3b-a1a535ddcb39 · inbound

Competition and Attraction Improve Model Fusion cites this paper.

Competition and Attraction Improve Model Fusion AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:32:11.923504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:32:11.923504Z digest=sha256:fca63efe56877b330f38b6490d92fcfc03b4ed6458e09733f0e4835f88028f1a

Observation 705747ca-f6e3-4741-b7d0-0560b18a8e74 · inbound

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality cites this paper.

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 196

Resolution
unresolved
no resolver link, observed 2026-08-05T15:38:55.198916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:38:55.198916Z digest=sha256:455f9a46a3ce9777da6ad86a1021d99326944e1fd4f65dd58fb8d6200fa82d61

Observation 9e9ec0d0-78c0-440f-8b6a-f8f10e505721 · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:21.982282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:097aa614aa24c8a14ceada8fd5fcce8789a6837bd7d542a9e45ce589732209f4

Observation b60a2987-9a79-4afe-b54d-839ea9a8336e · inbound

SEAL: Synergistic Co-Evolution of Agents and Learning Environments cites this paper.

SEAL: Synergistic Co-Evolution of Agents and Learning Environments AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.906170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:38:10.466713Z digest=sha256:a8ff0986816993ba8dbd244b1e6cbe4177c2b7c05f718e62adbf26170e3a0d72

Observation 503b0cf1-cd55-4649-ba5a-033f3ca27251 · inbound

Test-Time Deep Thinking to Explore Implicit Rules cites this paper.

Test-Time Deep Thinking to Explore Implicit Rules AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:54:38.202000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T11:52:15.163893Z digest=sha256:92becbc2b060db5e87c7b5849fb9ba60e1e42da23951471c43c1e7bc19c5dba5

Observation 84f171e6-fd0e-40b2-bce1-9a0d91612a4f · inbound

Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration cites this paper.

Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:08.811784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T22:18:15.244033Z digest=sha256:9f01486d8f9ce755931f002ee94f9bf44be6c165771fb558bf30dd47b1daf379

Observation cde52fbb-e5c7-4822-9c73-a604b322cb8c · inbound

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents cites this paper.

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:24.914814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T14:27:50.260308Z digest=sha256:8ef9bed49770d689d40b7bd65117057edf54d4d8c5186034fef8029376e2d159

Observation adea69dc-cc9a-40ab-b884-67c397419fd3 · inbound

A Scalable Approach to Evaluating Moral Sensitivity in LLMs cites this paper.

A Scalable Approach to Evaluating Moral Sensitivity in LLMs AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 176

Resolution
unresolved
no resolver link, observed 2026-07-12T05:44:33.099337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T05:44:33.099337Z digest=sha256:c78a568b9327930d8ea7c58a581585de52b5bfb3bf55e3b5868b1e7b569ea9af

Observation a778e4cb-40b6-4358-a5f8-9d10192347b5 · inbound

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade cites this paper.

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T04:04:29.284362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T03:54:37.954084Z digest=sha256:27dab3506eb5b7337d8b7b3efa4be738f36da63ba85db98fb109877936405184

Observation aa57d38e-4bff-46db-a5f4-b09d9ad51187 · inbound

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade cites this paper.

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T08:18:06.225163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:18:06.225163Z digest=sha256:e9ee138c198bc57a7d15e9f5be2fe8d2fd57cfaee2910fd66c5bef015da2a027

Observation 84e3e9a8-6a83-4214-b5d6-bfda514d015f · inbound

Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems cites this paper.

Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T00:55:17.107245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:55:17.107245Z digest=sha256:bc15660c926ccbbe4c1fa548b428973a2e06c84bfc8728461592a7b5737f9913

Observation dce7f585-c72a-4faf-bd27-b9cb0d912ec5 · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:35.622302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:35.622302Z digest=sha256:23f23a6502753754df373bbb4c75fcf63fa3f07365640137183791a788a2fb6a

Observation e95b5aef-30eb-4297-9537-ef34467951ea · inbound

Agent Security Needs Redefinition through a Holistic Framework cites this paper.

Agent Security Needs Redefinition through a Holistic Framework AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 218

Resolution
unresolved
no resolver link, observed 2026-08-01T06:04:46.361666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:04:46.361666Z digest=sha256:00d9a973f4e630699d2684706721198e435e33723c69733645a33d005c87a198

Observation 8eacd945-76f1-463c-93e5-fdfb59c3d38c · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:18.754752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:18.754752Z digest=sha256:904112db1b8cc2b7e5900b0422c42e218f45a88a0ec28d4e1f0366f56de10549

Observation 8eedf16f-c33f-4993-a960-575a7664928a · inbound

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability cites this paper.

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:58.749336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:54:58.749336Z digest=sha256:14b6617928ac90273c9a69c46791af70ede96e1112afdf45e0ffded675b77e49

Observation 303ee47c-8b6c-4f19-aca6-5f791664b8bc · inbound

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability cites this paper.

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:24:56.479523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:24:56.479523Z digest=sha256:dab7d400587f1a2643fd7303722cc09fa4682223769275a7558549a1b119f78a

Observation 21508525-cb10-4213-94b5-b0911f56fe6b · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.227255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.227255Z digest=sha256:15f041d11eb19456b71104b6124c61fcabed92fffe398fda113ac5318a29b143

Observation def733f9-f144-4f4d-bdad-841d229332bd · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:15.779795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:15.779795Z digest=sha256:fa18cf85686a6f9bf1cb34fa511e1cc39d4b5cda355a5729fd8a363d04b3a59e