Pith. sign in

Paper Citation Record · LEDGER

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

As of 4 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2604.18131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18131 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:36:27.381942Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:40:30.985824Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T16:58:43.673293Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact37
  • verified fuzzy23
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9987c547-3a84-4242-903d-9cf38d159151 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:43:35.479641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:c746dfd0412bb8160e1539f114c987291b42200095608a2f2529f0500f62ab8b

Observation c14ae87d-840d-4b97-af62-f824bc72ab98 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration WebWalker: Benchmarking LLMs in Web Traversal

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:10:23.405556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:a4e298fc4c2267e7b3e31724f85b458ae33caae0cb862f768074e9744daa624d

Observation 5927d826-f8ad-40fe-9b36-f282f87c7747 · outbound

This paper cites Qwen3 Technical Report.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Qwen3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.513889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:5505c91fff8dfa7705224f156087415075aaf2b25e2f1863e623fd7c814a01e8

Observation 4ffcccc2-aeab-4d5a-9454-91c30d52c195 · outbound

This paper cites Seed-oss open-source models.https://github.com/ByteDance-Seed/seed-oss.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Seed-oss open-source models.https://github.com/ByteDance-Seed/seed-oss

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.630014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:555ff67fa9b5f03362e7fff878f85daf4b41ab65cceee0cc1bfc2062debfb3ce

Observation dbd63f9e-d63d-4aa8-9140-287f1e5f4eff · outbound

This paper cites Qwen3 technical report.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Qwen3 technical report

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.633823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:19bbbd79f72b2006daa3689e6630a8bad4022c5bae0b600645967ebd81ac1bb1

Observation 956be0d3-309d-422f-ae0c-a889821bf62c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.427351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:ceeab584c5310a5f7d1df15681447ea64a0df54c3c4a8323e215ae49c5aa9c1d

Observation f1a73f88-48e1-40b8-853f-cd1e445be621 · outbound

This paper cites A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.948876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:11c299e4ce7390f5a1ad1e894534c4062f539abce88ce68a31a01e6d5d2ee8da

Observation b639b585-52a5-489f-ad70-58b2abc530b3 · outbound

This paper cites Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:44:08.223014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:7c2a67725fe4adbc9bbe52229156511dad6c981c361a0a678a03d8c88b3cb24a

Observation e117de84-062a-4d2d-9564-f3c49056d95b · outbound

This paper cites Cogito, ergo ludo: An agent that learns to play by reasoning and planning.arXiv preprint arXiv:2509.25052, 2025b.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Cogito, ergo ludo: An agent that learns to play by reasoning and planning.arXiv preprint arXiv:2509.25052, 2025b

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.488892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:2d4a856ee86272d80702467e1fdcd3387a2648c4864cc8e1a1e7d90982a79a7a

Observation 06ceceb9-7280-4be3-bd3a-eb5c6a0c7557 · outbound

This paper cites Self-Supervised Prompt Optimization.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Self-Supervised Prompt Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.492931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:213535fdd97e67eee2450a22c1322094a24c21856e9f030b745dd98020dce3e5

Observation dcce4d15-6ae0-4bf0-9fa1-e067566a0210 · outbound

This paper cites AgentSquare: Automatic LLM Agent Search in Modular Design Space.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration AgentSquare: Automatic LLM Agent Search in Modular Design Space

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.430752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:5f0b2d6d916e602399668b1fa82bfab63cdaafa1deb5374146010bf4ef7c4fa2

Observation cdc7e488-c35f-424b-a262-d0e067529dd5 · outbound

This paper cites LLM-AutoDiff: Auto-Differentiate Any LLM Workflow.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration LLM-AutoDiff: Auto-Differentiate Any LLM Workflow

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:10:23.505093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:b9029f396543273197f6b36c4258c40a9efa28aa2fe4070d93e3b28f02726353

Observation 3ee8dd9d-6854-4cd6-b9f9-2a2dbd5e64d8 · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:42:50.464031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:8e79fc87078219cc8264649e996e02523163c77f6e6a455df0e06231eb96c99d

Observation 662956d9-6f95-4658-95b3-911ec4ab7f15 · outbound

This paper cites MemEvolve: Meta-Evolution of Agent Memory Systems.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration MemEvolve: Meta-Evolution of Agent Memory Systems

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:18:15.377574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:588918282658e6e5b16c3b219f45310086b79b56ff40db7e10a3115a3f4a161b

Observation ceb3aabb-b435-497d-bc62-28752b1c4be3 · outbound

This paper cites Expel: Llm agents are experiential learners.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Expel: Llm agents are experiential learners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.637560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:c24499acb1d8d7e98444ae2332903e67f4bc4040517c16673d7f50f503adcb3d

Observation 01e52578-8665-4e3a-9c1e-48f3b390e376 · outbound

This paper cites Autoguide: Automated generation and selection of state-aware guidelines for large language model agents.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Autoguide: Automated generation and selection of state-aware guidelines for large language model agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.622496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:5d42c925c09583438ee246efc9629bb90a680b059c697d48507aa099e3071453

Observation 79c62f3d-2616-43fc-a395-4b189b0be422 · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration A-MEM: Agentic Memory for LLM Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:47:29.054201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:f51f56d1c67e1ae6a8fac7ac4fddc3222bb0ee46d84496422e8f387b4197942b

Observation 8b112bc3-88b0-4f4b-9966-0327c1cd6161 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:49.561366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:9d979d37d102ade9d75fb565e9d02a09638373fba375bb1c1cc90a45a077e016

Observation 32b0ea56-4b4c-40ba-8354-ee0b55937d18 · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:41:03.163713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:14b011a1dacf95ff74289e497df84dc28463dc780e4d78183e15c958441ac6aa

Observation c0f7519b-80d9-4900-aad4-f6c105090e24 · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:35:55.589248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:79c622f748d463dc8d842ef0330d1ab8eb83c48912ee1e000f73bde6311d1ea5

Observation 8e08c3d4-49bc-47e3-85a1-d889b57b9c05 · outbound

This paper cites From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.497270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:143060099ae375e861b5b5d333036a6bb1973b6ed5f6b0f562f5593bf98b3c51

Observation d9f8a8ca-5035-4450-912d-7044c6f54ea2 · outbound

This paper cites ToolGen: Unified Tool Retrieval and Calling via Generation.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration ToolGen: Unified Tool Retrieval and Calling via Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.375986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:2615333badb900b27f4c9af491b7cac17650954953217586d9397aa5152f1dc7

Observation 0dd1f92d-876a-4411-bec9-2be34833247e · outbound

This paper cites Agent Learning via Early Experience.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Agent Learning via Early Experience

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.268577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:9fb6995ce7128fe7dc2594e1a0e3c2d502d2cd95f7833b4d79a6ab95f6dc7de7

Observation 112e82ec-4649-46d1-bbc9-c81e6ce1c8d7 · outbound

This paper cites WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.463113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:fec484b04b613ce6ca31bb497275f7f90592d0becb9fa8a4a5b10c87dd9b3d4c

Observation 0ca00ec7-4d2a-4e55-9fcd-e12388adae42 · outbound

This paper cites AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.451741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:9f0c76635e514d85920c07aa788161f7f63e128afeacd106daba22f4d9ffe73d

Observation c5e44bd2-d50e-4964-b7d9-a0d1b125defe · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:93768b4e277a0f13fb6a35ad683682af68d8a26e9915d91da5e5b8ac37afc38c

Observation f34256b5-d1c3-4530-a970-0f95b022aaca · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.455390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:cbeeca403365d4d8b4594d80c6e265f242ac6deae93b339cd29547af24eea8ce

Observation 20cf6747-81f8-4cc0-b435-2a0115ac8bf7 · outbound

This paper cites WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.472158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:b9dd3404d7a3c7d2b5b6dc61d655679f0fd59466884832356a66bcd0896d611e

Observation f72d902f-f82e-445b-88b4-4500d7bcc4df · outbound

This paper cites Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.484189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:d36b5ddcd343aa646e277d62940619db3a4871c97025ee6540866b87164b20d6

Observation 4fb95ba2-e076-4cac-beab-8f23ec107d44 · outbound

This paper cites Spice: Self-play in corpus environments improves reasoning.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Spice: Self-play in corpus environments improves reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.381309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:c5dba92a9a9fb6c66357b6a66abca891aa13cf5378b05710d8fb6ed9ba86de34

Observation fb71f621-e6c8-4b55-ae8b-dba0204d34ac · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:17ecf415ac0fc0502b57b4ca3b420579b7595176e5c1c48e80f75d88feedc2c9

Observation c40d748a-1a6a-45f1-b923-dca4511200a7 · outbound

This paper cites Self-Challenging Language Model Agents.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Self-Challenging Language Model Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.528665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:b6b982332f4b9cf2a22302b5da1a654e5b742955eb6cf54b15c4bc08d03b80a1

Observation b784f6f3-78d9-4dde-b611-bb52576ec7e8 · outbound

This paper cites RLSR: Reinforcement Learning from Self Reward.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration RLSR: Reinforcement Learning from Self Reward

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:10:23.385138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:222fc73abadefd4d574d168d15816ec86086e50f52f51694ca1d3d255f176b45

Observation 7c6352e8-f8e1-4f74-bb39-cd9da20a8999 · outbound

This paper cites Dr. Zero: Self-Evolving Search Agents without Training Data.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Dr. Zero: Self-Evolving Search Agents without Training Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:22:03.468610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:1f9fb4ca3b1654027e478e4e73a0af04ff55116be0070523027d2c6c00c28561

Observation 82ab8871-8a53-437e-80f6-67f5ea3fd3c1 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Test-time training with self-supervision for generalization under distribution shifts

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.597526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:89c5f52698cfa14c17cfbaa397c65b66b53629762dd1e01abe63f9eafc7a0955

Observation 939b283f-c44a-4a6f-9de8-66075a7d0af4 · outbound

This paper cites ATLAS: Learning to Optimally Memorize the Context at Test Time.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration ATLAS: Learning to Optimally Memorize the Context at Test Time

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.480004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:2de637e9a1d4a6ec850ab57af68723ba62f4243864f618fb07b9c37de4b18e2d

Observation 874e0c24-2587-4a71-88a0-93625cf750da · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Titans: Learning to Memorize at Test Time

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:15.778565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:5b0f49682ee62f9baf7a9e0959cb2d5da5a0addfed372a5ef01b7565007acd76

Observation 61b4941e-a2eb-4f7d-94a6-c6c9ad753b08 · outbound

This paper cites arXiv preprint arXiv:2512.24695 , year=.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration arXiv preprint arXiv:2512.24695 , year=

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.394732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:8f952885bef972244b4440906fbfd10f4e34feff2df070d9954f3e614b911906

Observation 58c25473-602a-4c5c-aaf7-23436528b467 · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:20:12.625336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:0582161282cd3d4dbab3e58f5d49c3a86d2c751abec684dc75c790a5df5b7924

Observation 04b2e987-b6ec-4e41-95a6-d3c2f25f8008 · outbound

This paper cites With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.370274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:3b531540036eb74ab27f120807147770d3055d6b0093c4dfc6ad509257e52485

Observation 4b1c9e7b-606c-4e76-9206-6aa4c7b3e169 · outbound

This paper cites Test-Time Training with KV Binding Is Secretly Linear Attention.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Test-Time Training with KV Binding Is Secretly Linear Attention

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.390076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:0ce6c5d1aa27348972125bdb5cac7fe5edfc06be0df986a27f697e66a444cf9b

Observation 93921f84-0d4b-4461-9eba-c54a805fa90a · outbound

This paper cites Locas: Your models are principled initializers of locally-supported parametric memories.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Locas: Your models are principled initializers of locally-supported parametric memories

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.537277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:251be2bce1d6aa49ef8024c04135395628ad83a802f754d9b42a7375fd8890cd

Observation 8c8461f9-393c-4608-9713-a5662f24200b · outbound

This paper cites Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.467718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:6a02f0280d320b10b4c34d81fbc20b71b9d558a968d5753e2aedb180e3c469ad

Observation 60c42b99-93bb-4f17-9b2a-f1923f7ff9c3 · outbound

This paper cites Test-Time Learning for Large Language Models.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Test-Time Learning for Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.437950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:b3162bc7507472e24ba697f5f1e5f597deaf14d44e138481a51dafd1f712a13b

Observation 7ea67f49-1638-48a8-a956-4cb00ac93677 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Efficient memory management for large language model serving with pagedattention

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.544228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:47ba63c2162e7a261b53bb7bb58934cdc69adc0c0a3a44724377936a488e9594

Observation bd6beac7-a794-4e22-86a3-63b8b663881a · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.608333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:a43cc98474f9502447ae6356ffb2a1bce6c3fc890bfed480a00beb5312d4c1b3

Observation 2b57d1bc-1841-4ffc-a430-540c486ca2cd · outbound

This paper cites Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.410135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:877c40b8bd21ad98f1b6a5c3289cb8612a760c86b733c60fc552e20261e3ae7f

Observation 2501b603-95ae-4d8c-a412-3444d69b1b1c · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Qwen2.5: A party of foundation models, September 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.615015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:bb27bebb490f6fe68547db919bc411329340508b72a5f6e5c814ec3fefdeb41e

Observation ab100d05-7a66-4adc-8e9e-67de88882312 · outbound

This paper cites gpt-oss-120b and gpt-oss-20b model card.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration gpt-oss-120b and gpt-oss-20b model card

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.590859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:05f73254927fc95629ae30b602c16e356719069fbcd6431e4b860c997367f56c

Observation 393c8e84-5ab1-4fe2-9b43-4cda744dad59 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Kimi K2: Open Agentic Intelligence

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:49:28.234076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:d57fcc4d0a7aae6f5df8f52b68b1240d954398d266d034f682bb8ad6af1031cd

Observation 0d6c6fcb-0e11-416f-98df-b0e2d073817b · outbound

This paper cites Feel free to choose any category that seems underdeveloped or has interesting URLs you haven’t explored yet.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Feel free to choose any category that seems underdeveloped or has interesting URLs you haven’t explored yet

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.577891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:b0cd964c667b417727e155e20ce1e04a6172edd6e32a5ae484051e8f4bb31aad

Observation 66b87b25-1d12-4804-ab2f-8d9b53b7cafd · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.584824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:818b8d9663e670910608293640452a91cbd7bb62ab5b100251c412d07aeff313

Observation be7619ac-8580-4da6-b1fd-761d2aad5bf3 · outbound

This paper cites Please rely on the actual webpage content to inspire your expansion and ensure accuracy.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Please rely on the actual webpage content to inspire your expansion and ensure accuracy

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.570079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:479ea001b2970f12105bbce1cdab209eb26cbedf05109e1d0f5bc28fd4faec13

Observation c6470a4c-acf3-4f37-8d7e-8d2fbf6e2096 · outbound

This paper cites You can expand summaries, add new page entries, or provide deeper insights to make the section richer and more comprehensive.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration You can expand summaries, add new page entries, or provide deeper insights to make the section richer and more comprehensive

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.566494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:d3c027c8d272e9e1039e34c7abc21131a78eee8a96ea33e44c1ce19ebeacdc5a

Observation 34d3c242-89b7-4eff-a399-a2e72e4a6127 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.562047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:88ff9ef72815f732b25a82b62f4a6b5c9de082f03a50a6ba58d8f7b1f6b57df8

Observation 176fa36e-3c89-4ec3-915c-d787f0176794 · outbound

This paper cites Continue this exploration process until your guidebook reaches at least{min_token}tokens.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Continue this exploration process until your guidebook reaches at least{min_token}tokens

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.573942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:e0dae862d0e0045b8a6f3ffa0758b170a28e51f754b48b01f02fb08137b4e63d

Observation f16c0a29-8118-4c8b-bdba-d9bc90b9e5ad · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.625858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:589f71e79888a3db91b9ace5f991869f4ca06845ac5695335cb0681955e9060c

Observation a2a194c1-23f9-40d9-b412-d44d94ac72f3 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.648018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:9a2f57fe5e837e47d380ec8df0894a0c1794db7fa9cdfae967638f61d6b6d010

Observation 96babc0c-5e11-4181-a351-15ffde95257f · outbound

This paper cites Repeat until all categories are done.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Repeat until all categories are done

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.662219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:f32d2607f831232d723829ec221d74a58e23646208ff667ef85de41aa2cd4396

Observation b3f85665-9d0f-41db-831c-d867fdfaebf0 · outbound

This paper cites If it exceeds {token_limit}, compress verbose sections withrewrite_category_section().

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration If it exceeds {token_limit}, compress verbose sections withrewrite_category_section()

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.539974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:db289ed83dadc236603655a8757312f44b6d52aa84704b46876a6799a8546e18

Observation 52f7dd9d-b83f-4472-8836-d4534bf03178 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.651497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:4be379f6b428d63e76f1755f37e5cf8851872ffbc18ed80ae307f8457781800d

Observation 71e57e94-1c54-4ed9-91a7-bdc25a6f3674 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.547728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:3187f3192fcd9cccc3fa4df9ff129ef9b8c26502f66af1441aea418c69f322e1

Observation b235c5d9-f83c-4bae-b306-0d675ff11594 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.640872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:191947b77d72041adf7cc45e107fce7a1aeac8398fb40025b419c60af7986300

Observation fd8eb59d-4cd6-4b22-ad41-272272998251 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.604719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:c7104e270e3c82593bba586e64994f0351ce264de6f0959b1bf2ed5f6332843d

Observation 79ae878d-7290-402c-9509-c52442384ef0 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.558664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:4a3ddc9f9426123a990e65a3160315e10f50439cd256afd72bcc22c91959b804

Observation 3f573d24-0b0d-4c1d-b78a-8410e19e9e02 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.536029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:303592672e779a022ece996eafd188be42d7ce161559d1cb799cfd25e5ac46a7

Observation 351bdc7b-7cf4-44d2-a676-ce1ad58446b3 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.555065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:7b12cdcb05918c598b94361717c042605c6913998d68303608ee3ecc7ee4b221

Observation ccdfe041-528d-4f90-8a68-ae1cb642b913 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.528844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:15688fb39b2a7a5d21e433f589e833bc0fc118fbe263087603f3d2d86a2d8da0

Observation bbfefa89-6bcb-4a11-bf45-8a4c089eff42 · outbound

This paper cites Evaluation rules.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Evaluation rules

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.581469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:33fc3f01095a4a8c43bbd91b6d49cd8f4f8f09d6dee67d4717ec6abbd3f8d71c

Observation d0c9c27e-c400-4fbc-9f6d-fe525f7fe2f4 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.587623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:ea95ac155bd54aa50351d87653a8a8fcb55fff2970d961ef8b90313f1ed8b64c

Observation 375cfe5f-405c-4b2f-af46-98c3ef0d5366 · outbound

This paper cites Do NOT assume missing information.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Do NOT assume missing information

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.532431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:528393b76a999a792276bfa74b7aa600cccfac7432e52dff44fd148a4e10b391

Observation bf39d9e4-135c-4457-923e-5d74c286901f · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.593973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:ad618edd6d90cae795ee711f55e249e5a7c9f534088c04df77a801ccef2630b0

Observation 3b898a50-bb3d-47c3-9db2-79c983262460 · outbound

This paper cites Missing any part leads to NOT SUCCESS.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Missing any part leads to NOT SUCCESS

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.618740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:9045aeac345e2cf561f22df762a0443b55274ce079fddd9bdac67dd03c1d2dc5

Observation ecbb0a0c-afc7-4239-9801-c92a676f3cd7 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.611288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:994644f8db8ee1a0918fc0b44bf67a7089351507e2b14ad0581a87e9d1174417

Observation 438933b1-cc71-4fe6-a051-bd07cef1ddd4 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.601411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:944aff7e499553948eeaec1cb06149669a5fadc5297b502a973143e8978ed7b2

Observation d7a22fde-addd-4b6c-bd37-e334106b7f81 · outbound

This paper cites Instructions:You should briefly explain your reasoning before giving the final verdict.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Instructions:You should briefly explain your reasoning before giving the final verdict

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.551653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:51fe2813c6a601ea74e935e3579d2027ea3a682815556ebb291c07e9fe549bf7

Observation 658b7f15-dee3-415a-94e9-d44da434fca7 · outbound

This paper cites an unresolved cited work.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-22T03:56:02.520800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:161ad3c073ad6a35612d49e7930e9592d879c589cde27dd0086ba0513bf4f749

Observation 649bee9c-9171-41e4-9714-66b2ec88bc0c · outbound

This paper cites content summary.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration content summary

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.525336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:f6f87e0859e4dcdf123be463991c6fbfa8d2a10deee2939b932d88868aaccca0

Observation 1d32b665-8878-4ae6-9145-717d6996dc51 · outbound

This paper cites content summary.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration content summary

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.644185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:71a14a321aecccd49a909004806b7f428361d13e8952c1a9d08f083d9f3b4b61

Observation 41de6431-8acb-43d0-a534-924bed26e4d0 · outbound

This paper cites About Us - Team Introduction.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration About Us - Team Introduction

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.654833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:841f720bb4ec0e8e6e818e5d64510da1eef62afbb615795cddaa11634278d32b

Observation 2300fff7-852e-4bad-a06e-8613b565da7c · outbound

This paper cites In this case, you must decide to return to the [Main Page URL] to start looking for new clues from scratch.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration In this case, you must decide to return to the [Main Page URL] to start looking for new clues from scratch

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T03:56:02.658669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:71226c5969efbaf5997cd9f51a64ff1e899e9ebdec68bb295d3316f5aeaabfed

Pith citing papers

Observation 8de6d7fb-a4b1-429e-8b2b-ea0c51c37110 · inbound

From Question Answering to Task Completion: A Survey on Agent System and Harness Design cites this paper.

From Question Answering to Task Completion: A Survey on Agent System and Harness Design Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

Reference 163

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.674751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T04:40:30.985824Z digest=sha256:f1e8d4252bd82e8940e864dc963023f1c18b0c72eed0110c6ec68706d10c1c33