Pith. sign in

Paper Citation Record · LEDGER

Agent Learning via Early Experience

As of 6 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 34 inbound Pith citation observations for arXiv:2510.08558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.08558 v3

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:48:06.310505Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:51:46.531155Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T20:00:07.735443Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98353674-d12e-4a05-9b5b-93480057539a · outbound

This paper cites write newline.

Agent Learning via Early Experience write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:56.506509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:56.506509Z digest=sha256:cc09e1800d0b7d4b53ac6a664faf3ef49259ce79185ea984f6d32355e47cd6b0

Observation f2957aec-e8bb-450f-b12c-e803e3d3c071 · outbound

This paper cites Hindsight experience replay.

Agent Learning via Early Experience Hindsight experience replay

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:56.584040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:56.584040Z digest=sha256:111949b1222c013c589af71b771e299af6b2891d72eb11cedea9d9cdc41f3494

Observation 6b25a9a5-87af-4259-8d5e-f356859708ed · outbound

This paper cites Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling.

Agent Learning via Early Experience Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:56.723578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:56.723578Z digest=sha256:a6d5aa56611c53d792af84126018b44b01fd9cca7389687faa7af587ab0672c7

Observation 4c5c6e16-f8aa-480a-8b1f-0177606aa294 · outbound

This paper cites A markovian decision process.

Agent Learning via Early Experience A markovian decision process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:56.900269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:56.900269Z digest=sha256:9dafa9c97bca3843f506bffc8653b05bd0b8aeb61f38ec6a934e0872ed854dbc

Observation 4175b265-f9d7-46db-8d4c-520439f915c2 · outbound

This paper cites OpenAI Gym.

Agent Learning via Early Experience OpenAI Gym

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:57.125475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:57.125475Z digest=sha256:e31cd54977d08fbff927f1639e44804f280ccbc9ccdeb1e8e6c903bd1e73fa30

Observation d74d271c-b12c-48de-8669-261eca641cb8 · outbound

This paper cites Web agents with world models: Learning and leveraging environment dynamics in web navigation.

Agent Learning via Early Experience Web agents with world models: Learning and leveraging environment dynamics in web navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:57.297587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:57.297587Z digest=sha256:5bbefd43615430b1fe76eb2e3fd21c3162520bdd1423f6ee8c53b0a8ee8a0b4a

Observation cfdb7304-bb93-4667-b3a7-0c6d70eae062 · outbound

This paper cites Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun.

Agent Learning via Early Experience Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:57.473817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:57.473817Z digest=sha256:b6b4426e17beb88845b8447c69e0cf7e626ae8c40dba0d5fbcc7b9ee6db81222

Observation 12a48378-77dd-455d-a24a-6354db69e839 · outbound

This paper cites SFT memorizes, RL generalizes: A comparative study of foundation model post-training.

Agent Learning via Early Experience SFT memorizes, RL generalizes: A comparative study of foundation model post-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:57.646059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:57.646059Z digest=sha256:99bd7ddff9929cca6b2d73db53d38352ee68183edb9b692f204e5b1b4221a29d

Observation ad1297fe-b9f7-4f4f-ba0b-5501da60cfaa · outbound

This paper cites TextWorld: A Learning Environment for Text-based Games.

Agent Learning via Early Experience TextWorld: A Learning Environment for Text-based Games

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:57.846696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:57.846696Z digest=sha256:9cfe6bc423da695888637b838fd04abda15f774e53677804c28e7859d64e8d2b

Observation f2653aaf-fc4b-43a2-b890-73a02237e909 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Agent Learning via Early Experience Mind2web: Towards a generalist agent for the web

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.076593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.076593Z digest=sha256:c280b7dbbde1b29a2ebba3dabdae79b71b249f271f3ebc31bb1fbf78bd2bdac7

Observation 0cf5a722-79ac-4f9f-8ef8-669cb43f5b89 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Agent Learning via Early Experience Group-in-Group Policy Optimization for LLM Agent Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.288029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.288029Z digest=sha256:9efc1133662a1646a004b5da256e0131b0e6019c5c4135517553214d19a48745

Observation eff87050-7876-4b69-8fe2-c579e4538ae1 · outbound

This paper cites o rg P. M \.

Agent Learning via Early Experience o rg P. M \

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.389893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.389893Z digest=sha256:41b077e623330ee06e732b5f9c4eb23ea1f07bfbeffe9e137811250666446627

Observation 8eee5330-96f6-42e4-ae48-9245f0c587ee · outbound

This paper cites Multimodal web navigation with instruction-finetuned foundation models.

Agent Learning via Early Experience Multimodal web navigation with instruction-finetuned foundation models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.507889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.507889Z digest=sha256:e33d77c40d5590789edf38e14f97b29775a857cdfffd1788b1cac829012ae687

Observation 9989590c-175f-49ea-992e-207d1eca1144 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Agent Learning via Early Experience Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.616326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.616326Z digest=sha256:f80db8731ca0736795f8536a7060162b7b63533147311ce2115fc11b4bfb8540

Observation ba0c974f-6795-42e0-9cf5-2c226c5b8d20 · outbound

This paper cites Middleware for LLM s: Tools are instrumental for language agents in complex environments.

Agent Learning via Early Experience Middleware for LLM s: Tools are instrumental for language agents in complex environments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.722164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.722164Z digest=sha256:87cad56da1befb15ad8f7d950f304d2db3ca2870166671e3115d465473be8909

Observation cc9ee8f5-43f0-4df3-b64b-4bbfb3ffbc5c · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

Agent Learning via Early Experience Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.803710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.803710Z digest=sha256:0531abeb9be8ecd7e3f92324d493d0461533af3de3af62c60c3c3adb79f7fe08

Observation b71b2ef9-9321-4376-b30e-ba1d1a25f069 · outbound

This paper cites World modelling improves language model agents.

Agent Learning via Early Experience World modelling improves language model agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.908791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.908791Z digest=sha256:1b0112511f329158c9f583ce345a262a959de9c156fbbd64622e891a13af99d8

Observation 3a0b34ad-cecb-4482-bde0-5e88ceebf935 · outbound

This paper cites Recurrent world models facilitate policy evolution.

Agent Learning via Early Experience Recurrent world models facilitate policy evolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.984749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.984749Z digest=sha256:d7629bf9ff0e30f2bd5959add334f11ca5090904f60d03c0ccd2e0599ea66fa3

Observation af8199f1-020f-49a0-ad5c-a33eefcbce67 · outbound

This paper cites Dream to control: Learning behaviors by latent imagination.

Agent Learning via Early Experience Dream to control: Learning behaviors by latent imagination

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.064163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.064163Z digest=sha256:b1367fa0c2ae595b40a7ec0769ac609f496102a152ccf3fad5b5e592f6438097

Observation f7771a86-7a17-41d6-b73d-c51000c17f49 · outbound

This paper cites Mastering atari with discrete world models.

Agent Learning via Early Experience Mastering atari with discrete world models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.153672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.153672Z digest=sha256:b9fa29742e8daf71b7b908f2b889d9c573143fd8ab780601c81c72f51e25362b

Observation 697bf0e5-d2e2-4a84-9146-783fa9004246 · outbound

This paper cites Reasoning with language model is planning with world model.

Agent Learning via Early Experience Reasoning with language model is planning with world model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.271600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.271600Z digest=sha256:be9192ecbd4c0f85bf43e7bb8e563c2f5a70eb59f112a4f14b11dd243623719c

Observation c8ccb03b-ca18-49d7-bd77-0c5a77f7c3b7 · outbound

This paper cites Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps.

Agent Learning via Early Experience Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.356603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.356603Z digest=sha256:b7445016b6ec79cf88062579335e62f34388e5b1a6facecf066237b79bb4a917

Observation d738211a-229d-4dfd-89f8-ebc7e9b917bf · outbound

This paper cites Cogagent: A visual language model for gui agents.

Agent Learning via Early Experience Cogagent: A visual language model for gui agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.442881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.442881Z digest=sha256:1b45b8435ad66c77ec93c93d69f33cf4426a204f16ec18e0ece847374e8da33a

Observation 6e532730-8555-4857-adb8-7336787ffabb · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Agent Learning via Early Experience Lo RA : Low-rank adaptation of large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.563349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.563349Z digest=sha256:10229324474ba54b8c15c9711cc1ccd3b5b32bcb8b6efb1bada56df110053c34

Observation a12b6eaf-3eaa-4498-85a3-211e024ce892 · outbound

This paper cites Large language models can self-improve.

Agent Learning via Early Experience Large language models can self-improve

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.628208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.628208Z digest=sha256:f38de21c88feb043b90cbdadc1c89dd8e4006d47dd4f75a4ac637c0726eb02a7

Observation f9fb40ad-a4c7-4db0-abcc-c85d8e6f6176 · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Agent Learning via Early Experience Large language models cannot self-correct reasoning yet

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.739590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.739590Z digest=sha256:42ca8b81dbc69f533d5f7077a23ceb7943d57f4429c193d53b88cb61557756e3

Observation 6f9831e9-31c0-498d-9d39-748193f910c1 · outbound

This paper cites Imitation learning: A survey of learning methods.

Agent Learning via Early Experience Imitation learning: A survey of learning methods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.850231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.850231Z digest=sha256:8a16aec2497f76ed417937c62eb714e06e6e5380ee755dfc96c202199ba872c5

Observation 2230664e-8cdf-4292-b0e9-c4a75ac757bd · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agent Learning via Early Experience Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:59.970459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:59.970459Z digest=sha256:50f67022ae25b1367b9a5e2275bc27a84fe2cf044e5d0a6e1a51f7b08cbf2262

Observation 6f383fb3-7639-41c5-a13b-9219a6a0639c · outbound

This paper cites Tree search for language model agents.

Agent Learning via Early Experience Tree search for language model agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.093999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.093999Z digest=sha256:0aa34459fc3b5f3f50dbf3deb552a0c7edf7ac8f9d8fb6a6bc49308221d86503

Observation e75f23f8-27ca-4f8e-ada5-ee995680f6ac · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

Agent Learning via Early Experience VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.216972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.216972Z digest=sha256:0ad741a9a03cb2f4ce6755d66f0b1a49f740addd5ee15f287ed6008d4614ad93

Observation 89bcbb31-2d78-4396-96d9-c9a9784ab24e · outbound

This paper cites Visualagentbench: Towards large multimodal models as visual foundation agents.

Agent Learning via Early Experience Visualagentbench: Towards large multimodal models as visual foundation agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.302101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.302101Z digest=sha256:80912934f8bb7adfdb8c4d73939fe1033ccee5a97c5ca4d075db6099fc79d2e1

Observation 8ccca055-feba-4e1c-b71a-0bcc5d883f46 · outbound

This paper cites AAAR-1.0 : Assessing ai's potential to assist research.

Agent Learning via Early Experience AAAR-1.0 : Assessing ai's potential to assist research

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.416385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.416385Z digest=sha256:ed621c1833633daa2bdb846e0c2422c10febaf0b4d189b6018293499cfd1ca4d

Observation d809de26-f766-442d-8de4-5857bee8b53f · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Agent Learning via Early Experience Self-refine: Iterative refinement with self-feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.516077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.516077Z digest=sha256:83d1f08917c63fa72e57321b93a61f4544914920a53fc830fa7a4e35a303edfd

Observation bc199b04-a244-4f96-aef2-b47fa26b334d · outbound

This paper cites Towards Enterprise-Ready Computer Using Generalist Agent.

Agent Learning via Early Experience Towards Enterprise-Ready Computer Using Generalist Agent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.611698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.611698Z digest=sha256:208c9b8f6db192067abac85f492c0a10a22ae2ccb03004e0f6125f03f58ee09e

Observation 1ec18984-96b6-4bdd-8252-bc3c999efc12 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Agent Learning via Early Experience Playing Atari with Deep Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.750900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.750900Z digest=sha256:c75bfc242de39c939da6672e19981d9c5ed97d31b536b45c7095d218dc5c3af0

Observation e9962b77-3bb7-4c3d-9324-d2f28e76cdfd · outbound

This paper cites Manning, Peter Shaw, Mandar Joshi, and Kenton Lee.

Agent Learning via Early Experience Manning, Peter Shaw, Mandar Joshi, and Kenton Lee

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.830638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.830638Z digest=sha256:6c99c5a1c3e81f32e2ccbec8d93fef8b55ff77f4ae5203fb664e28fa5de19c10

Observation b0c7cdd5-2e0c-4548-a7bd-bfd721a933a8 · outbound

This paper cites Hello GPT-4o.

Agent Learning via Early Experience Hello GPT-4o

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.863967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.863967Z digest=sha256:5554fbbf596a118bb627c9f711dec87b398e6cdd8f3b7a99e12e0ed64f5ae05c

Observation 530fde74-325c-4b96-aa89-1e6bf4132b76 · outbound

This paper cites Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents.

Agent Learning via Early Experience Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:00.989334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:00.989334Z digest=sha256:0e54a8a2e4ada1b126a961689d4d8ae72e22a3e5c6302e6718d216da6d0ff7bb

Observation 5d7bfe75-1935-495c-8311-1e34cdf4b344 · outbound

This paper cites Gonzalez.

Agent Learning via Early Experience Gonzalez

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.121050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.121050Z digest=sha256:428fc0b6effa5dcface501f19c16a35b9500b2f8c90e1f69365b42f236757983

Observation 97e52afc-0f4c-4ddc-8fa1-f1cc9882ef00 · outbound

This paper cites Efficient training of artificial neural networks for autonomous navigation.

Agent Learning via Early Experience Efficient training of artificial neural networks for autonomous navigation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.145663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.145663Z digest=sha256:d2d8139f08031334b38b455646f5ba7117b1b0206c6e2cfbb4838115b48b10de

Observation c077b242-53db-460c-8a05-1ae325f6f53d · outbound

This paper cites APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay.

Agent Learning via Early Experience APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.215310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.215310Z digest=sha256:67acd35b1bd28eb2a92ac8b6c87ee5272e4a0578687cd9724d2977b13c829af6

Observation 8ad3072c-a31c-4410-a97a-a3b66aca7efe · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Agent Learning via Early Experience Measuring and narrowing the compositionality gap in language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.373304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.373304Z digest=sha256:50f453611a42ef700fee96e4558d1fb284c6d537c7a1914c9d10c2b09b25306c

Observation 456f2f7e-ae95-44e5-9e90-486f0853b895 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Agent Learning via Early Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.456515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.456515Z digest=sha256:bf0fb8e81b2afdd33f11d102b50019eb5f0eec11c9ac11b16512393392552f5f

Observation ccdb5895-460e-49b8-86f4-161f28e4bca2 · outbound

This paper cites Web RL : Training LLM web agents via self-evolving online curriculum reinforcement learning.

Agent Learning via Early Experience Web RL : Training LLM web agents via self-evolving online curriculum reinforcement learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.572095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.572095Z digest=sha256:6722a8dd8b3a2c5ca630771c645be6688e1986d7c55dd132ed3ca75e370bf783

Observation 4c148590-0c74-4097-b064-dd4f7bd01825 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Agent Learning via Early Experience ToolRL: Reward is All Tool Learning Needs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.703082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.703082Z digest=sha256:8da61a4ad5289edbb5a2e8e86c5661911020380d212581307f052dee31a2bcfb

Observation 0ed98744-c594-4e95-a3ef-e39120147ab7 · outbound

This paper cites Efficient reductions for imitation learning.

Agent Learning via Early Experience Efficient reductions for imitation learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.808110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.808110Z digest=sha256:997009fd9b2eaaa3574f5f890a91fb1447b47a6c4c2f72c28f6eb9f7a0c8938e

Observation 757c5d10-939a-4b0f-a774-c3ba3b52b523 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

Agent Learning via Early Experience A reduction of imitation learning and structured prediction to no-regret online learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.923890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.923890Z digest=sha256:3cb4571bf0cc07f466e259992db024338a6b2a840ac872975134eef11b5e73ba

Observation 6213fdc4-694b-4f78-a163-f97a66536212 · outbound

This paper cites Russell and Peter Norvig.

Agent Learning via Early Experience Russell and Peter Norvig

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.019600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.019600Z digest=sha256:fcd99f2e85654565b99ad0c3c63e91998e1955dc258730e57f8c085ad6f8e2d0

Observation f1508f2a-73e9-4c01-a470-a63425e73673 · outbound

This paper cites Learning from demonstration.

Agent Learning via Early Experience Learning from demonstration

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.138239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.138239Z digest=sha256:bd7d4bd1e5d645365a18ed8bca9b3dece26510fa8f240b9fc7dffc1f58788d43

Observation 3e94de37-0c24-4429-b499-ea7b959930a3 · outbound

This paper cites Mastering atari, go, chess and shogi by planning with a learned model.

Agent Learning via Early Experience Mastering atari, go, chess and shogi by planning with a learned model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.267088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.267088Z digest=sha256:c3540baca591a70af62bd88c45ce82cb217a710b5f52a8dec24ce88d7db2a18d

Observation 5fbde27e-d5d1-45cc-9636-d31c2353c480 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agent Learning via Early Experience DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.362478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.362478Z digest=sha256:74052bae353376d405bef33cb7b7f2eebe3f23322adae48198a306e616aeda87

Observation bdca6f21-046e-45d8-8633-062648eea9ac · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

Agent Learning via Early Experience ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.524137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.524137Z digest=sha256:b345e9c022e555b27edd141892ba618d77c7e7ba7d4b04d143820b5643a0dac7

Observation f42aa539-7284-44f0-a554-901f7a2cda1b · outbound

This paper cites Hybridflow: A flexible and efficient RLHF framework.

Agent Learning via Early Experience Hybridflow: A flexible and efficient RLHF framework

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.675594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.675594Z digest=sha256:552a1ad0b232a9c4293ef65f9cfd7020e4a8b6989ca3499edc4dcdd4960e976b

Observation 1812c5e6-9304-4766-bd46-1c4853408aa5 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Agent Learning via Early Experience Reflexion: language agents with verbal reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:02.837901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:02.837901Z digest=sha256:c02fd8787194f8153b45c0dd417f362a330250daf9805d24b63b823e86cfd2cd

Observation 6fd6c576-536d-4126-bd9e-f9283fa589e1 · outbound

This paper cites \ ALFW \ orld: Aligning text and embodied environments for interactive learning.

Agent Learning via Early Experience \ ALFW \ orld: Aligning text and embodied environments for interactive learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.004397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.004397Z digest=sha256:e9c68cbe3f68fe13a665b9d0354b0d7d21e7b7af017b3afe088882d5211e4bc3

Observation e1feac97-edcc-46e2-a979-39fd7f69520f · outbound

This paper cites Welcome to the era of experience.

Agent Learning via Early Experience Welcome to the era of experience

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.172498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.172498Z digest=sha256:5feeab19598b72490c14a837e3235dc70b96575564f22a2595ee5769c495ba7b

Observation 5b8fd965-3dc7-4c6b-ba68-64252683cef4 · outbound

This paper cites an unresolved cited work.

Agent Learning via Early Experience Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.334789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.334789Z digest=sha256:02ab3548462e8e44c8d57aba3161f560003341033c04720bc5d1ef3de987ba65

Observation 1ba1141b-7b50-48ad-9600-db220049afaf · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Agent Learning via Early Experience Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.340432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.340432Z digest=sha256:9aba2c5b00f7459fe67f887625e21067e9b193e24c35758885dda4646c627316

Observation 3b04eb56-2391-4aa3-92c1-7eb63fc7a2cd · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

Agent Learning via Early Experience Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.400830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.400830Z digest=sha256:6cc8e4b13896a96fb7c482081ac1a533b18006d61508484f1275938c6b9e34d3

Observation 151b51f2-7274-44eb-913e-345770543c78 · outbound

This paper cites Language agents: Foundations, prospects, and risks.

Agent Learning via Early Experience Language agents: Foundations, prospects, and risks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.503938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.503938Z digest=sha256:578b1edad4e44f696d10f29e85beb31b430ba22ecaa083d19acc206e32fe487c

Observation 2e7712b4-add6-48df-bfec-d7f4a1c4338f · outbound

This paper cites Cognitive architectures for language agents.

Agent Learning via Early Experience Cognitive architectures for language agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.585571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.585571Z digest=sha256:322a97c3ebdccc8cb5e7231d1ea9f8461adc131e7c70f378085c8a69d0ed3cff

Observation 24bfd7eb-7037-4ff1-beed-2e826f039d07 · outbound

This paper cites Dyna, an integrated architecture for learning, planning, and reacting.

Agent Learning via Early Experience Dyna, an integrated architecture for learning, planning, and reacting

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.678603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.678603Z digest=sha256:d4d10c8f88546afc28dd8b7b7daeec821e964aedc973a34c985ecc10a26a9da1

Observation ef3492e9-62a0-465f-bfc6-a83d847becb6 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Agent Learning via Early Experience Reinforcement learning: An introduction, volume 1

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.815177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.815177Z digest=sha256:f5aba37d51f6b4af9e1074a5f865690f48248cc8f501bc51c205c278ac3e4c13

Observation 569dcd57-83c3-4d2f-9560-8be2130be822 · outbound

This paper cites Efficient exploration in reinforcement learning.

Agent Learning via Early Experience Efficient exploration in reinforcement learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.913957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.913957Z digest=sha256:0408ddf7595db05e835fd070cd9c6fbe4cad398091b7072f75eeb683a3a373fb

Observation 85ecd5f2-fb70-4812-b4f8-a81b1a283db2 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

Agent Learning via Early Experience Musique: Multihop questions via single-hop question composition

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:03.994776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:03.994776Z digest=sha256:64fb8c6037e0b55bb22c4bc9d5d2ef4b96ec95055047f1caa62fbf214e99b5f3

Observation e850cbe3-f015-4836-9abd-168ba02a1440 · outbound

This paper cites A pp W orld: A controllable world of apps and people for benchmarking interactive coding agents.

Agent Learning via Early Experience A pp W orld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.039176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.039176Z digest=sha256:4ea3a603f9363070bc20e2303befe5083886bd4660f2622ba49b9efaf054a7ca

Observation 43b25c1d-9ebd-4f3c-92f7-3917f6a07cb9 · outbound

This paper cites Investigating the effectiveness of self-critiquing in LLM s solving planning tasks.

Agent Learning via Early Experience Investigating the effectiveness of self-critiquing in LLM s solving planning tasks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.124786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.124786Z digest=sha256:0a3f7a089ea46bd1284e6357bb8a568dcc58c5d452dea566432789a8589ec2d7

Observation a676dd61-4aa2-448b-a334-fd7efb8acd68 · outbound

This paper cites S cience W orld: Is your agent smarter than a 5th grader? In EMNLP, 2022.

Agent Learning via Early Experience S cience W orld: Is your agent smarter than a 5th grader? In EMNLP, 2022

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.197435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.197435Z digest=sha256:81e765aed04cf7a623f677918981ce07194e35610f8f74ca89a62c94bea17c77

Observation 2f2a474c-08e4-4498-84cb-d20a478138f9 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Agent Learning via Early Experience RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.269782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.269782Z digest=sha256:4fe5323d768a1eab01010571ddcb62e2455127d0666040b22a39ba7252e5280c

Observation 98925eb0-8367-487a-b990-c9f09c17d489 · outbound

This paper cites Chi, Quoc V.

Agent Learning via Early Experience Chi, Quoc V

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.282352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.282352Z digest=sha256:77bb980246386d710552e714f33325c4d4bbf2b73928e558f5077648bd7b04cc

Observation 38eac46c-db24-4a9b-94de-e82e5d182901 · outbound

This paper cites Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.

Agent Learning via Early Experience Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.593498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.593498Z digest=sha256:28bff7d6a096c4ad0ad5662945710c89a58370c1db7dd4c07aa22710bd5b52de

Observation 6023e19f-39bf-4353-abe6-2e6596f65974 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Agent Learning via Early Experience Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.792221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.792221Z digest=sha256:998f06a44491cdc89740e3b7cef944eeb61eafb21413416bcd5e386e04d0e8d4

Observation 1411b88a-6749-4946-88a9-3403e3ec155d · outbound

This paper cites Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark.

Agent Learning via Early Experience Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:04.961013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:04.961013Z digest=sha256:8252d93c129ae607e99a33c225a7f4ca556303cd21d31e5472127ed9563d6281

Observation 4db8f092-61da-40b1-9c5f-7c418823dd20 · outbound

This paper cites Agentgym: Evolving large language model-based agents across diverse environments, 2024.

Agent Learning via Early Experience Agentgym: Evolving large language model-based agents across diverse environments, 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.133243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.133243Z digest=sha256:b158a0ce87b13991d2a46d467a0ac5db66d9799209a8670686efff46246f1ae0

Observation 65b0b769-6f38-4193-ac6c-0f190fa8606b · outbound

This paper cites Travelplanner: A benchmark for real-world planning with language agents.

Agent Learning via Early Experience Travelplanner: A benchmark for real-world planning with language agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.227412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.227412Z digest=sha256:8990e0caf92ae83979929e9b61c6245f3a85e8893ac855212353361ccfe1abe5

Observation 48915cd1-c204-4e3b-bbfe-a79248f6fab4 · outbound

This paper cites Revealing the barriers of language agents in planning.

Agent Learning via Early Experience Revealing the barriers of language agents in planning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.371114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.371114Z digest=sha256:6da1672db5620b8a1a7ab78dfb90982d8439bf6039cf225fed8fae4b178d8335

Observation a79b1651-c840-418d-b89d-cf525982f3da · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Agent Learning via Early Experience Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.486278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.486278Z digest=sha256:6c249ee44da0d9080cf3d38b90f6d03a59ef946a73c6941967768fcdd41769a8

Observation 6f3e55cf-1a28-4600-a8f1-b968b8a367fb · outbound

This paper cites An illusion of progress? assessing the current state of web agents.

Agent Learning via Early Experience An illusion of progress? assessing the current state of web agents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.562175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.562175Z digest=sha256:97bc696772652f093bc18097a1daf3410bd5f6d0b22615474f578684a3d9f2dc

Observation 5c2919eb-8c15-4b23-84f5-4da8e101284e · outbound

This paper cites AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents.

Agent Learning via Early Experience AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.644648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.644648Z digest=sha256:4e10c9ec8e14e784aeb3e99808d2c469f783381d87da8fa9ea0bd130822e9523

Observation e34bb7ec-7172-48f6-adf3-1207db33ba34 · outbound

This paper cites an unresolved cited work.

Agent Learning via Early Experience Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.700685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.700685Z digest=sha256:eb0c6b192f6bde070eae9576d150a7042fb4faa1d227bd277a189ed098b07ada

Observation 4409272e-f614-4508-b19e-fa60dc60919a · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Agent Learning via Early Experience Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.785257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.785257Z digest=sha256:db7cc8d0bda113ff35077af1447a2f6356c566fd9a3f9ae4c039f645bbe1f79b

Observation 73b3c26e-60a2-46f3-9889-5aaa862fcbe9 · outbound

This paper cites -bench: A benchmark for tool-agent-user interaction in real-world domains.

Agent Learning via Early Experience -bench: A benchmark for tool-agent-user interaction in real-world domains

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.865689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.865689Z digest=sha256:d175ae26853fb90c474ebd8dab29e606a090102a5b65719ecc908dfa29588795

Observation f266dc3d-a1f2-4fef-b678-82b3a8e7af4b · outbound

This paper cites Dyna-think: Synergizing reasoning, acting, and world model simulation in ai agents.

Agent Learning via Early Experience Dyna-think: Synergizing reasoning, acting, and world model simulation in ai agents

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.921855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.921855Z digest=sha256:644f69fc5e3981a66c074d45992a715981886a7b359cb7333bee6e2cddc16e13

Observation d83ac179-c6d7-4aed-87f2-4fab3e267c79 · outbound

This paper cites an unresolved cited work.

Agent Learning via Early Experience Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:05.984194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:05.984194Z digest=sha256:67cd0acdee9d3e256fea7f2e4f3d104cac0b8c93965c1b1d460ceddd497fb058

Observation 3104a2ea-9b8c-4d5b-8d7b-ff0b2b0dfbba · outbound

This paper cites Breaking the Data Barrier -- Building GUI Agents Through Task Generalization.

Agent Learning via Early Experience Breaking the Data Barrier -- Building GUI Agents Through Task Generalization

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:06.046348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:06.046348Z digest=sha256:cb6e27c55ecf9a0285a490703f833f4ab652829ee8b7d1156094273185bd1d17

Observation d831caac-22ec-405f-923d-a32db62dc671 · outbound

This paper cites Gpt-4v(ision) is a generalist web agent, if grounded.

Agent Learning via Early Experience Gpt-4v(ision) is a generalist web agent, if grounded

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:06.102772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:06.102772Z digest=sha256:51c345a4bd4a346825d56575ec96f909ab91059d5fa24b8d0b8e14136a285850

Observation 8a7caab8-8afd-4fbe-a471-f65565a268cb · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Agent Learning via Early Experience Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:06.154764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:06.154764Z digest=sha256:43d35408d5001e620a6d58aec23c79753acac5ccfb8b6e60565316505242ca42

Observation 351629a1-78ce-472e-8892-59e6e30e08f8 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Agent Learning via Early Experience Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:06.206487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:06.206487Z digest=sha256:22b16da866c2ae11f6a1b07bdf1c9edde19e5d0f7103b81357b7f375215489dc

Observation 76755336-f839-4dc1-9ff6-32b58057e22b · outbound

This paper cites Self-Challenging Language Model Agents.

Agent Learning via Early Experience Self-Challenging Language Model Agents

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:06.257869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:06.257869Z digest=sha256:92b872b7864f4fcad371bd81f9db719159bfb9d5ece1bb470e99eff4409f661e

Observation 3d1bf8f6-3fff-470c-890c-9c7287944595 · outbound

This paper cites Proposer-agent-evaluator ( PAE ): Autonomous skill discovery for foundation model internet agents.

Agent Learning via Early Experience Proposer-agent-evaluator ( PAE ): Autonomous skill discovery for foundation model internet agents

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:06.310505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:06.310505Z digest=sha256:58f304daadb9c75673349385adad0e3f6a07485cece10bc2e79b70f6af813918

Pith citing papers

Observation 005f710f-f0b7-4ba8-93d4-6d9b46aacbdb · inbound

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory cites this paper.

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory Agent Learning via Early Experience

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:27:40.232015Z digest=sha256:ec9f632b55fcbe36fa831eef33d9484dfa35074b75b1113853ed243f7483aae4

Observation cd77b6ba-29ef-4761-a3df-410a882af35c · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL Agent Learning via Early Experience

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:46.531155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:46.531155Z digest=sha256:3b77ad3c3c8fb06867be8cbf5b20da58f10fca9f4fd1a16c2faf242431c49b17

Observation 1f2e425a-fe2c-442f-885e-df90f2f18237 · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics Agent Learning via Early Experience

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:66023e5ff5b92ee1dd33ae8e885ed54793463572390c16d38e9da914090c3cf9

Observation 444a7421-5b68-40d7-9059-f279b1d7d0b0 · inbound

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering cites this paper.

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Agent Learning via Early Experience

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T17:40:14.733882Z digest=sha256:b5f3d2565bdf7a112ef6d51b803f536b10c34ab103317ebe365ecdbc78e5a19b

Observation fe2d8dea-1f92-4ae4-85f9-a56a5df9830a · inbound

Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory cites this paper.

Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory Agent Learning via Early Experience

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:15:13.188244Z digest=sha256:1385a379b33b336fe0dc42d25db94b35752b2bce0a898e17fa1cf8179d6e33f0

Observation 3f1b556f-f2c0-43a1-a866-0f3241e284fe · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Agent Learning via Early Experience

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:6661f4d64a38fada5b9bce6f0ae672f3cf6febae9610c4ad18646305f6ede157

Observation 0dd1f92d-876a-4411-bec9-2be34833247e · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Agent Learning via Early Experience

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:d9358a83a35c2a61be7d2608cb3b01830394842d1879c54cb56e6767d6d7e604

Observation 153fe713-d22d-4d3e-ac86-39b708db7566 · inbound

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models cites this paper.

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models Agent Learning via Early Experience

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T00:50:27.257888Z digest=sha256:1f8d03614926f101b520c44eb003adcff95671aa20c747800befa727503f1308

Observation 3d057f68-acd6-47f9-8821-ce5edd187112 · inbound

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent cites this paper.

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Agent Learning via Early Experience

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:34:22.546508Z digest=sha256:f7abd247c613935609cf1e3c92a39257cd725b32d56aee09133a4927ec92112e

Observation 8a99a819-e3fb-4fea-b322-17548c5d8b1e · inbound

Revisiting the Travel Planning Capabilities of Large Language Models cites this paper.

Revisiting the Travel Planning Capabilities of Large Language Models Agent Learning via Early Experience

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:52:54.690422Z digest=sha256:39b1508b76e225a5145a01749a77a43f7033230e078ff0163c6425b2346eb090

Observation a2ffa417-7f58-4024-b852-e1740b3d2a91 · inbound

From History to State: Constant-Context Skill Learning for LLM Agents cites this paper.

From History to State: Constant-Context Skill Learning for LLM Agents Agent Learning via Early Experience

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T16:50:43.547830Z digest=sha256:d517e95275883756a93ff6fcc4badeee8572213e1b83a3ec72680de02d5ae707

Observation 259159bf-eb49-4132-ae5e-a86900fc5214 · inbound

CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment cites this paper.

CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment Agent Learning via Early Experience

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:16:30.734428Z digest=sha256:9a061ebc8eed50c1ac90d4f248ff107c606ad64148b8d81910a4732f7cff5d79

Observation 506667a1-1d5b-42e5-b27c-4c823f3c70ae · inbound

Learning Agent Routing From Early Experience cites this paper.

Learning Agent Routing From Early Experience Agent Learning via Early Experience

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:15:07.381414Z digest=sha256:e690fedf23adbf7e168cf75b6aa0d565eb9f3747d9dd56ba838badd4ff33b04a

Observation b1e37e50-7213-4dcf-bc7e-a48be15f5573 · inbound

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning cites this paper.

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning Agent Learning via Early Experience

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:48:53.213389Z digest=sha256:608156b8272350a287fb96d3544bb1ef5011563845c92a77933843ad68782c34

Observation 5824f0ba-200a-4502-b3a7-fb0853273a85 · inbound

Differentiable Mixture-of-Agents Incentivizes Swarm Intelligence of Large Language Models cites this paper.

Differentiable Mixture-of-Agents Incentivizes Swarm Intelligence of Large Language Models Agent Learning via Early Experience

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T19:53:04.689519Z digest=sha256:4fea7511c8274350e4186edf15b32d1da152583b1a5cbdf09bd383b9538dabc3

Observation 75a8d3bf-63a4-43c3-94b3-4b219e00150e · inbound

EXG: Self-Evolving Agents with Experience Graphs cites this paper.

EXG: Self-Evolving Agents with Experience Graphs Agent Learning via Early Experience

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T22:17:10.728289Z digest=sha256:7e85d3fea9423b6543f01bc33297581ade28a38b0deae4a9a91f9ad3c3de7d15

Observation 75cae065-4852-417b-95a3-afbc0306ca12 · inbound

ECHO: Terminal Agents Learn World Models for Free cites this paper.

ECHO: Terminal Agents Learn World Models for Free Agent Learning via Early Experience

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T15:04:46.503768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T14:57:03.095107Z digest=sha256:8dbe52421496aa39fbc1a58dc15badaaaf9348fdc4bf48fc67490d6d9fc94907

Observation 6bdf2e05-4eb1-4d78-b2cd-0c719b5ca6bc · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Agent Learning via Early Experience

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:02:33.901234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:82b0868facf34775d51d27f5750e14c7eb0402bc198b6a854bacb01fc904b833

Observation 37c3835c-7587-45f3-b4da-08b33bfee284 · inbound

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents cites this paper.

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents Agent Learning via Early Experience

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:24.862123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:27:50.260308Z digest=sha256:b066039dfe6074f32400cb194a7060d000f1c1efde1090fc4fbff57cffef1f20

Observation 2b94da71-ec1b-4c20-a2f3-7e81eab16ea4 · inbound

Policy and World Modeling Co-Training for Language Agents cites this paper.

Policy and World Modeling Co-Training for Language Agents Agent Learning via Early Experience

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.190939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T15:40:35.531232Z digest=sha256:4d4214866eee2caa3710ec916ccbe53227ab3ec4df68a29a31559b3224323455

Observation a50be362-d665-487c-97ff-bf9f22e4c9cb · inbound

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials cites this paper.

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials Agent Learning via Early Experience

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:28.748387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:33:26.954959Z digest=sha256:ace7e0ad6d0dc81aabb7d7405de2ce0b264d53e2a976c187bf39d893b9b2d0d5

Observation 3a71f265-0b05-4144-be95-d1611ff72898 · inbound

Co-Evolving Skill Generation and Policy Optimization cites this paper.

Co-Evolving Skill Generation and Policy Optimization Agent Learning via Early Experience

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:25.878799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:d7b2cfd258decb177c17e1bcd4c024ad2ad9f4c87e1d314480455b6cc7a97c3a

Observation 9768ab59-e63f-4f93-960e-6e66d8a2b11a · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Agent Learning via Early Experience

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:19:59.997876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:2aa340cf688b20d72a415e43ad002e08d2fd1115a260b25dff3c10c5daa6d4d0

Observation b2d13abf-35c5-4890-858b-8701e787e150 · inbound

Qwen-AgentWorld: Language World Models for General Agents cites this paper.

Qwen-AgentWorld: Language World Models for General Agents Agent Learning via Early Experience

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:59.206662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T23:52:31.403419Z digest=sha256:e999e2a189054e6b26451779237ab13dd6933a589b644e323fec5eef04cf111d

Observation d66174e5-e55e-441c-b5d9-068dc173bb2d · inbound

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making cites this paper.

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making Agent Learning via Early Experience

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:06.177688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T21:11:50.874696Z digest=sha256:a3cd1a4c4da81153bbf025be1dab2a09f05c4c58c22f274a67a8719e79d738d0

Observation 3290972a-195c-4a15-99e8-51fff462aead · inbound

The Interplay of Harness Design and Post-Training in LLM Agents cites this paper.

The Interplay of Harness Design and Post-Training in LLM Agents Agent Learning via Early Experience

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:00:07.736814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T20:57:20.179394Z digest=sha256:a28f582f96f5efde0760bbf949cafded900ca83f08d8b6e9621182b9efbc1ea0

Observation ae2fb3c1-9a00-4b39-8be2-a1a8dcfbdcba · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? Agent Learning via Early Experience

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:49:51.529478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:717e8fa9de4f4af4f57d6c9b45674d3cddaeee5da11dcf0224803486eb3152e8

Observation 197dc697-0f81-465f-92a4-03757ec2c530 · inbound

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning cites this paper.

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning Agent Learning via Early Experience

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:35:58.497615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T01:53:56.067792Z digest=sha256:d2925923775f90406dcca3b37c9d99bcdc9454cef2560bbffab86bc5afbc2665

Observation 89fe5cae-e6f8-4568-bc91-c097dc38e617 · inbound

Self-Evolving World Models for LLM Agent Planning cites this paper.

Self-Evolving World Models for LLM Agent Planning Agent Learning via Early Experience

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:41.504930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T05:48:11.441644Z digest=sha256:a35ae941b30a7fe74ce041d52c194b3b5ab7e8b5cc0fe7b382f6cd63c6a272df

Observation 183b1700-f357-449a-a6ee-def7c82df27d · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agent Learning via Early Experience

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:2d9bec50ec3b56b88b89b58c2bda378bea3194aa8ec6b97acc8601034bdbda11

Observation f65c47cb-620b-48dc-b740-9488350328e6 · inbound

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks cites this paper.

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Agent Learning via Early Experience

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:59.640364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:59.640364Z digest=sha256:8d13fce73fce4bd372b97b95f6f70e63aff3b51e12cbe5ae1a3fcfb9c87ec348

Observation 6154c06e-df20-472c-a8ed-76a6620735ef · inbound

Scaling GUI Agents with Visual State Transitions cites this paper.

Scaling GUI Agents with Visual State Transitions Agent Learning via Early Experience

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T23:04:34.096442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:04:34.096442Z digest=sha256:2ffa955987fdf95a6f0b3d3511d246c2d4acf0f49f193689b9232b4b59232740

Observation ae9ad3dd-80de-4d02-a865-45dbeccee357 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Agent Learning via Early Experience

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.103981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.103981Z digest=sha256:06027961a7a0540139c19bc9bc3832913b87f90b775b81ab39fcea2acbd016fa

Observation 5d95d0b2-43f3-403f-98bf-6a6ec15cb62e · inbound

TAPO: Transition-Aware Policy Optimization for LLM Agents cites this paper.

TAPO: Transition-Aware Policy Optimization for LLM Agents Agent Learning via Early Experience

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.634912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.634912Z digest=sha256:70fa4b89f3cb322ab18d74c74dac93464ef3cd2ea4b2e90d98bb43623280f553