Pith. sign in

Paper Citation Record · LEDGER

WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 66 inbound Pith citation observations for arXiv:2411.02337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.02337 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 66 of 66 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:30:16.492871Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:20:00.041093Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3daaa13-a142-4469-86bb-cb0838180c28 · inbound

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection cites this paper.

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:33:21.204452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:33:21.204452Z digest=sha256:2c52c1fa041b79bfb74f04b343b8577bffb7a9783e1ffc988a647ad6fe8cd4b4

Observation 3d05642c-2028-4729-99d8-165576046229 · inbound

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning cites this paper.

Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T11:25:49.165415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:25:49.165415Z digest=sha256:dfe50ed47b7d99bb242d1e885db2ad1a3b40f375ee572468cc9732e03ef03715

Observation 1a439e27-a856-4426-b5ad-49f735016f73 · inbound

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks cites this paper.

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:32:18.628495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:32:18.491541Z digest=sha256:2248202918795863abbc5db3f4cd9187b8889788c1a24939f8f204b47f159fa5

Observation 8a08d509-760c-4940-bbdb-8350116a3a5f · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.751909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:55d8f585b3a5b7d718f736adeb517e8d1e7cf04a6825379f0f74f4e0f2ab690b

Observation 32a45fee-21df-4e70-b4a3-62c063e1f489 · inbound

WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents cites this paper.

WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:30:16.492871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:30:16.492871Z digest=sha256:6377732c8b63c856e89fb0b1ebc06ccba4354e0c5d80eb080a7c7bbfb7d9750c

Observation f2db07ca-f27c-4984-b4d1-274e67a6632d · inbound

AI Awareness cites this paper.

AI Awareness WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-16T10:19:50.644595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:19:50.644595Z digest=sha256:e78c1cddf2e24d4e3849f5e26441a803a204ad5a6f7faa772053a1399958792d

Observation 5207287a-f3da-46da-ad79-2096f4a1d921 · inbound

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning cites this paper.

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:04:20.014449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:04:20.014449Z digest=sha256:1c677b8bf68ace0d8f82e1fc2480cb0c706beaf19560d7e11ffe7f4ca73e9bd1

Observation 4efd62e6-4159-4d0d-b049-b694982913da · inbound

ProgRM: Build Better GUI Agents with Progress Rewards cites this paper.

ProgRM: Build Better GUI Agents with Progress Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.329762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.329762Z digest=sha256:28892251e19d938504a7b3ab0afa63e71c666ec81b4aca2d34ae163ed104ad9e

Observation c754b48f-8e54-47d5-923c-ac41bbb93d9f · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:02.794899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:02.794899Z digest=sha256:f3a9fc8c793c6b0a487b02dbcd314f03cfedc057541ec6878a119ec2ebfb9a00

Observation 84a7a678-435d-4764-9a4c-23ce173b09bc · inbound

Agent-Environment Alignment via Automated Interface Generation cites this paper.

Agent-Environment Alignment via Automated Interface Generation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:32.880275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:32.880275Z digest=sha256:e4c9fb1ba98e8ae3d0bc4abe700a403876801c6d65457d09d208614555fc48b1

Observation f94e7756-8544-49e9-941d-b4d6cd891db6 · inbound

AgentDNS: A Root Domain Naming System for LLM Agents cites this paper.

AgentDNS: A Root Domain Naming System for LLM Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:11:08.422497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:11:08.422497Z digest=sha256:bf50640d4f8d5b4db4726ac7ed0f89bd4c687191a46a79f243ccb0c436b1a7ee

Observation 767db96d-c03f-4295-8692-071b489e8168 · inbound

ZeroGUI: Automating Online GUI Learning at Zero Human Cost cites this paper.

ZeroGUI: Automating Online GUI Learning at Zero Human Cost WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:24.310946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:24.310946Z digest=sha256:98d974ef931df3f731cbfff906cd1d35e928f1ab4014ed53e4615734df4527a3

Observation 14a271ba-3cbd-410d-abf3-89156620648a · inbound

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation cites this paper.

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:33.721752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:33.721752Z digest=sha256:0fe7703b8150202f2a6e6f3016d7a1ece14953db0c04621e807691a36bbd6029

Observation dcba1132-d734-494e-be0a-fc5a787f1ed5 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060344Z digest=sha256:d8472834f191be8e7fa4ce96e62620373f619ddc054e947e8e9dd9ead5d52e52

Observation 59ccd5d8-9b02-48c3-8bd2-ba65dad8f4ba · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.716794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:afebd59db33904dd8921fbe4c806c689ca6ed6558d230f8af5cbe3c43ba24db4

Observation 7498d1e4-28c7-4c4e-bd30-e41f5108b9a1 · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:21.810260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:21.810260Z digest=sha256:f65ced1447e5663122bc50f3ab078f092eec0dce0c410b0cf0e71f6bcd644f3f

Observation 2e782046-7782-462f-9b1e-06e866525791 · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.978039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.978039Z digest=sha256:372cdac0e331706d846c1e50ca0022b6edc73ed6919a8fdc89b50c3931286b02

Observation 23b686ae-f595-477c-a155-9ea064784556 · inbound

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System cites this paper.

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:04.613379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:04.613379Z digest=sha256:3e7e73a60f7d95598f7baa47c6880d5478ef3561c687310066f8b437f05a38a0

Observation 584b1166-6f2e-4fa4-b163-63d343223acf · inbound

Build the web for agents, not agents for the web cites this paper.

Build the web for agents, not agents for the web WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:00.763378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:17:00.763378Z digest=sha256:30ff52d10779850f696d166d325ad2c791af7b5aa587088d5b27c73edb7b9557

Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · inbound

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards cites this paper.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.504844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.504844Z digest=sha256:6afa13338ed3033ba985b3d83d36915c0a4e0ac5a651f5ca0c4cc9b7b6a447e5

Observation eea4fae6-f494-4e05-9c8d-05c761871989 · inbound

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents cites this paper.

MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:27:37.434374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T00:27:37.360221Z digest=sha256:7ce03569e54f955dfe41ad8d01380e145112636a620ddc97c4eae3a22f683d70

Observation 17d20baa-90fd-4c05-a94d-e0f51339143c · inbound

ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning cites this paper.

ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:28:27.562193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:28:27.562193Z digest=sha256:4f6e6a69255c86c336cc8c74b9e1a70764d4c4ab7aef981d8ca19c5cde9a579c

Observation 4a1f1581-f669-4224-9f39-26e04b903bbc · inbound

WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis cites this paper.

WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:01.164687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:01.164687Z digest=sha256:92a08eb649501331520052d4cddbc461d1e6744260bd4db8aa2065ccd305208c

Observation aa0dba40-ecf1-46e5-9f05-f8ccb1263036 · inbound

MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service cites this paper.

MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:13:26.792085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:13:26.792085Z digest=sha256:269f6d2781ef81cc249bf7fde3cd5940f36ef086ee76bd3d53c9315f3be8d2c4

Observation 25925f65-3209-4701-9598-15f976fa1c29 · inbound

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? cites this paper.

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T18:05:09.345860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:05:09.345860Z digest=sha256:36748488c951e1e6154f38627768cdd4a499d6310b1a8c6ae0df1eb2f5349f39

Observation d5f4170c-1008-4f8d-881f-50d4d395223b · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.489007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:4f31e226660379a134abe1cc1d26f8422efe6dd28fd0270cb464deeda4a16961

Observation 56fc05ce-ba75-47ee-88b3-2906e94b33b0 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:50.893647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:50.893647Z digest=sha256:712bb00b6d78084a2c88f3429815ff444ccf2395d56e1beb537fe254f8b9803d

Observation 3420f575-8354-49a0-a0ac-002bd0281248 · inbound

Cognitive Duality for Adaptive Web Agents cites this paper.

Cognitive Duality for Adaptive Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:36:21.986350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:36:21.986350Z digest=sha256:4a73c6cb8c990653e15a2e5d1519bab22c1a51650cf902a028de430efb528e1f

Observation 8ebe02db-d3f0-423e-8c0f-2cf95809e21e · inbound

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making cites this paper.

EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:24.747926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:24.747926Z digest=sha256:786c1b689e8a4de3f0a1d67a8d699502fcca8b045f126f503e8ee5f272b1e951

Observation a32608b8-1845-430f-b8b0-6d33fa8819dd · inbound

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward cites this paper.

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:21:50.831536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:21:50.831536Z digest=sha256:ae7a63fdd70e307ffe4b7396bbbd4d9d5a700ed109286a4b6e023dbfc2e5c284

Observation ceb3ec28-6bb5-463b-97f4-03a7fcc6ba6e · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.759166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.759166Z digest=sha256:fa66bc66ed2a3f54c045a17b5382555fc5c4f4d953daaae379ef54110ce470e9

Observation 6841f209-0436-4463-aefc-e54ffc9774f9 · inbound

Symbolic Graphics Programming with Large Language Models cites this paper.

Symbolic Graphics Programming with Large Language Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T05:34:10.896680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:34:10.896680Z digest=sha256:adcaf306667c35dc5749cd86559ab9f36debc95716b47b107665c5cf52c31cd7

Observation 404606f8-a388-4e91-bd77-0ac751fe4792 · inbound

A global log for medical AI cites this paper.

A global log for medical AI WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:15.570579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:15.570579Z digest=sha256:353c96535ebde7ba91ff988be72f34efcbe42dc843b8e2ca58c5f4767c0e934f

Observation 456f2f7e-ae95-44e5-9e90-486f0853b895 · inbound

Agent Learning via Early Experience cites this paper.

Agent Learning via Early Experience WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:48:01.456515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:48:01.456515Z digest=sha256:fa69761877e5e049fc637e68426a70af91bd8502717ed7996a6065d10a8b26d1

Observation 3dd41f15-24e5-4f99-a39f-ff294db464cd · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.018864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:45f0bc2f39ef65dfde42739ba924aa9e14c848fea6dcda7b204e03de39cd97a4

Observation d6916505-eaec-47c5-8552-e8d1b831704b · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.011939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T20:55:10.819298Z digest=sha256:da7e40418817f26818d7e25601ca0aebcb7247675e54a2dd5c53c8ac6186a4f0

Observation 52a3c39b-62df-47b7-9026-a1735bc6fb0e · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T21:29:34.747408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:29:34.747408Z digest=sha256:fe1c0a69e322d95534742f75fdc10363a2244314dae8b1483e6ae9a8ccd55b66

Observation 59479678-5c43-4a1e-b726-d3dfcd61f296 · inbound

DynaWeb: Model-Based Reinforcement Learning of Web Agents cites this paper.

DynaWeb: Model-Based Reinforcement Learning of Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:37:41.942922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T09:33:30.444057Z digest=sha256:f4cf57aa5b5a14041eeff302e48898bd401e0cab71df69d1aa2abaadef403a2e

Observation 577a560a-e36c-4b7c-91b0-9fe6805dd2df · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:43.534953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:43.534953Z digest=sha256:e71d7e3a00903e0ca55913b98870ca15aae1321cb5563bd1731ca65dc29a1b56

Observation 8b693a41-6a3b-4de6-9082-f74d0f1123f6 · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.237069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:b19513f34b024aea6fcb45894e796c482fbf4fc05a8346e9aa1e8f2a390e2001

Observation 1b0603ab-9dfe-4f64-bc5d-bb4f5a41243e · inbound

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning cites this paper.

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:56.367208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:45:08.769264Z digest=sha256:ad586296ae1758b2d312ea4775a6ca6f714e1aa2f1001c7448028d05bb6607c8

Observation 79d01310-4cf2-4dc6-a202-fc8fec2f88f1 · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.877886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:cf8e954ec4b19784c0f352d406aa6a331a959492282efa6e5af98fdf40423c53

Observation 666a29e3-17d5-4c68-ad9e-ed837f4bdb64 · inbound

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning cites this paper.

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:24.706366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T19:21:30.956682Z digest=sha256:720ef67b5b42828529a839b90c4d64bc3f9f672006bd81aaabb25ae817297307

Observation f2208569-4677-4a3c-bb48-212c0babdade · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.349693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:ac89213983cf75e9dda80cb9f2de050c18107affc0244bb98571f7416799b1fe

Observation e0674fe6-1e47-49fd-85ed-65a97b37eda1 · inbound

Milestone-Guided Policy Learning for Long-Horizon Language Agents cites this paper.

Milestone-Guided Policy Learning for Long-Horizon Language Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:51:11.409606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T10:50:40.760938Z digest=sha256:650935372ee2b7cca113a85afe1637009540c3760fd9bc1b7db1c276e18c08bc

Observation 058e9db2-225c-470b-902c-fa57a3760c99 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.888142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:d1abf13becc9d78d68a2e662bb0afa20940abf9fdda9d3ea223af5cc1a22d990

Observation 11ce60e9-4099-45ee-8ddb-43cd23bb6da5 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.542352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d754f3a430eb409e6ceb52cea62b5f83451b3c694218a98e7f3aaa9af3e253ee

Observation 61a65632-bb1d-4a68-a58d-f995578abf69 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.149642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.149642Z digest=sha256:dd79ebedc767c404146ae4519c6a4ffb64add70b29486b7ab0ac85111e35424d

Observation 0ed81619-69fc-4fd7-9a72-49183809b3a4 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:26.500402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:80acd44b3d76f654149ff8698f18c340a0d29f8fb016b2c93c3d7bc75e963f28

Observation bd028f99-e4e3-48ac-b386-d6b57c581225 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.620192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:aee77156c7ef8d01dd44dde13d2e7f8452a7e3306b1e8c7980596d55fb491f14

Observation b069b3c8-cb86-418f-94b9-e3117a4981e0 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.549697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:c6920c1721132d18d206aa0aa24750cafb684fba4c53d0111340cfbc62f3a6b7

Observation ed399df6-0832-4121-919d-a59193c3a5bc · inbound

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection cites this paper.

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:09:51.308730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T08:07:59.093519Z digest=sha256:b6869da91c50a1ca54233f679d75e221813e50790245fad00882ec31382fda8d

Observation c62bc5f8-cedb-4ae3-a1b4-4bcd212971a4 · inbound

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection cites this paper.

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:15:01.478092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T18:39:12.161639Z digest=sha256:dcece288cfd3f6092565569e80a8576debb72a2bc9dd282d93f6d92e266844a7

Observation 6e5bcc9a-5671-4435-ad9f-e72c017169f0 · inbound

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate cites this paper.

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:29:34.490134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T04:27:25.041652Z digest=sha256:b9f929d33b17a442570a0dc424d28e67f6a59de0e71befa17100fa340560fe44

Observation 6557edbf-7468-470d-a8f4-45441a6db87a · inbound

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning cites this paper.

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:36.639556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T09:01:22.412685Z digest=sha256:ea87fe31da4b0daa7579bb4f48d56222b0697ea01c48ef94f0bba007760cf6e6

Observation 3d93c6be-e952-4002-9d0c-81f60346f689 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.605741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:82bc642c65e1fe83071b62a55021cd0e58338e86be683d8697e70f1891000d0e

Observation 9821d9a6-fad3-4048-a98a-743e8d7f257d · inbound

Deep Research as Rubric for Reinforcement Learning cites this paper.

Deep Research as Rubric for Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.262998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:16:16.243967Z digest=sha256:39eaa22ea3485d4802063136ffcc3c15c8a524251102161378a4210e1bff628b

Observation abd50648-c9f2-4285-9a0d-9ff79f69f7f6 · inbound

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning cites this paper.

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:31.170792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:22:05.424293Z digest=sha256:2695fb2f22583374b25c244efd05ac1460f4d7d18b25251ac9a374a9bcf56b6b

Observation 9bcde929-0217-494a-ba96-38d72605a508 · inbound

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation cites this paper.

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.241228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T10:53:27.839408Z digest=sha256:d1b76ab4a20e3e541e8f618fbf3cdef369492c118a62a6ab9fc53f92e4e33ed7

Observation 0ce9644d-fea6-4905-a03a-468dc6aaa1ec · inbound

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents cites this paper.

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:59:38.183304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T13:56:51.914966Z digest=sha256:e00cbdd41ad146edf5af1fdeff4f0b60f6e3fc99036e3c4c8788d6236a5adfe1

Observation 8abca0e9-e8ab-472b-b5fc-d427f2488334 · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.042865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:3ab1f3d64f26903666cd092fbef160753ef3701be845a22efc7c423ec051fc20

Observation 18440b04-b9a5-4dea-b32d-f6f4e07bcc55 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:09817f4623715b1127635e9d73247e13fc37be7ec3e93a7b1e4e89069018dc6f

Observation 3ea66aef-dc4e-4898-8121-a7bf36a89eb6 · inbound

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL cites this paper.

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T06:25:23.264527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:25:23.264527Z digest=sha256:8948c49bdcf9142af84bf64846ba69cbcdc03224e8e394cf57dbba7c83f2c3e4

Observation 258eb65e-660f-489a-ac9d-68331475c65b · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.764481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.764481Z digest=sha256:b1ec263071ad849c1c26bfa4ceabbd7600a067f11e7d3f8ff4a8493c1b9beb6d

Observation 2ea62fda-50e5-4f86-b47c-1711cdedaf18 · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.962904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.962904Z digest=sha256:2fed99dd76d5012b4c00e64966d639ab1b42620e4d616fae8bea5b64ad4552ba

Observation 919ded30-bacf-49f0-9ce3-940a28ee6b4a · inbound

Software Engineering for and with GUI Agent cites this paper.

Software Engineering for and with GUI Agent WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-11T20:19:15.589837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:19:15.589837Z digest=sha256:55789b89f29c0876325f215ed1e3fcbaf51f95f01fac3c20ce8197bd8963b178