Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:39.559055Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 0 inbound Pith citation observations for arXiv:2608.00155.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:39.559055Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 102 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 67b611d2-1277-424a-96f0-bd6abc117c0c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15c3b5f3-0cc4-407d-97dd-53e6f149a015 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774e17c1-81ba-4c06-8184-c3942a218a0c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d0b54b-6b3b-4420-ba41-a3bb96d97d7c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic context engineering: Evolving contexts for self-improving language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b30595c-b4bf-4321-82f1-c1ac004ba677 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d1cbd62-4880-457e-9d8e-76e4c51ce49d · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A-mem: Agentic memory for llm agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4a68ac-830d-4563-8af2-bc2c42882100 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9997acb5-415b-4035-aced-f5f12b482392 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae5e5f3-dae5-48a1-9f88-e580831d52be · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669b2ad1-81a2-4075-bd74-1379c5d0b903 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Memento-skills: Let agents design agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff5ba51-e69a-4fe0-958c-0ca5c67e030e · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8650ca55-dd31-4e2b-b6c2-2e8233c3b8e2 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Appworld: A controllable world of apps and people for benchmarking interactive coding agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0313bafd-50e0-4c88-95b5-24fa3f54c936 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gonzalez
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 192ac56d-3eb2-4e4d-966f-b596268400d9 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Swe-bench: Can language models resolve real-world github issues? In Proc
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22fc5a9-a231-4699-b27a-21e74d18ae93 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Humanity's Last Exam
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ed4434-2e42-4aac-a302-79bb5630afb5 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ab7ed43-9bc6-413b-8b5e-082f746f58f4 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Stream- bench: Towards benchmarking continuous improvement of language agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19de0ee3-c2cd-4dac-84f9-1710b859a36c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8cc89d3-74d9-495a-a8d0-64d85f8c11ec · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenAI GPT-5 System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba34156-2f11-4e2f-be16-4be50f196bea · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gemini 3.1 Pro model card, February 2026
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ba58dd-a90c-4480-9b54-3e6fb70502a1 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Introducing Claude Opus 4.7, April 2026
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e4e000b-6c7e-4e9c-b814-b4951cdeda99 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c9edd8-577d-471a-81e4-899bf8f7293c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training with self-supervision for generalization under distribution shifts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef7ec98-9596-452d-b766-b5a37395abac · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec2d7e4-606f-45df-bcbf-08412cdd5a1a · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7128a647-c5e2-4fc7-bee4-44afb379adaa · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gradient episodic memory for continual learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b29ec8-639f-4f57-9e71-ec2d6c5b41ed · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Scaling llm test-time compute optimally can be more effective than scaling model parameters
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea24b3aa-c55c-429d-8f36-4f6fd65c80e7 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training on nearest neighbors for large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b76d30-5c03-4755-a817-87bedd5539bf · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Efficiently learning at test-time: Active fine-tuning of llms
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8353d4e8-7813-4986-b32e-8b3fd29c1545 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The surprising effectiveness of test-time training for few-shot learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aaae17b-f559-45b8-b4aa-042734c11429 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? In-place test-time training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78bb23fd-0340-46e7-a979-5c9cb6c60f57 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time adaptation for llm agents via environment interaction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c6d34b6-56e8-42fd-be9c-40b104d6fb5b · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time learning for large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c98bfe-fc30-4da1-b482-26138dfd8ce1 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttrl: Test-time reinforcement learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 560c579e-c5e5-4904-893b-432431fa2cb1 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729ca982-0b5d-48a6-bc45-3aa1aa448f4f · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to discover at test time
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dbd5315-bd55-49ce-a095-a6213b44dc44 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a992a21-8df4-4314-91b8-8d6506facb9c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? What if consensus lies? selective-complementary reinforcement learning at test time
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb42fec-7ea7-47d8-add2-4a5c603aa6c1 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76d5b23-7867-48c6-ba76-c7b3adbf5fd3 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01a108e-71ff-4206-a9bc-190f1a4e571e · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-Time Learning with an Evolving Library
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9d85b7-402a-49fc-aec2-7efb0b250981 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2262f2-6162-4fac-abde-a3da7f3f3208 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fb1c9a-50dd-4827-84ac-1aa5afaa48dd · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic plan caching: Test-time memory for fast and cost-efficient llm agents
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed2c9bfc-f745-4cf6-b322-a65549161db3 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18107843-204b-4eb5-b983-e6c2028238e2 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09b02a8-d8b8-453a-b705-88ffabf8a7c3 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5ea4cc-4e98-4228-9cbd-eb3796899c71 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Panini: Continual learning in token space via structured memory
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6a2359-872e-42b0-bcc8-368eb0a6533b · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df12cecf-038a-4dc1-8a2b-c5be36a6d665 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec8e04f-7212-49fe-a4ed-8b154e68aa71 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb473b65-f10e-480b-bb9b-04a615958777 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xskill: Continual learning from experience and skills in multimodal agents
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da43f976-c6ce-463c-9799-361ce994f6aa · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Online Experiential Learning for Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc78d07-c649-491c-9b06-7196fe1fb0e7 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42845b4-b90e-46aa-b1b5-2dbf86837931 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 227b17f5-290a-48c0-94bd-9da4a8674c98 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6271b35-7585-4cff-ac8d-17152f2132e8 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fcce961-c4c5-4a04-8cba-613fc43b0b43 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOS: Learning Skill Curation for Self-Evolving Agents
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d83db8-2aa5-4635-b235-d2f7d73e311e · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? EvoSkill: Automated Skill Discovery for Multi-Agent Systems
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a4cd23-7cf0-42d9-bfbc-c364071b1d2d · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenSkill: Open-World Self-Evolution for LLM Agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3406787a-4e71-4b33-9ab7-b8a54d9deec6 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8206b5-8e87-4d2a-82ea-3561f5a75200 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Meta-Harness: End-to-End Optimization of Model Harnesses
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4deb2ae-ff8d-44e6-a47e-4bf4271f2963 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd65faba-7aaf-4848-ae01-57b277ee676c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Selaur: Self evolving llm agent via uncertainty-aware rewards
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78298233-2985-415e-a37a-9f4e5254b1ca · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-Improving Language Models with Bidirectional Evolutionary Search
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b75e4e-19e0-460c-86e8-ec6acc6fd764 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e9eb84-99d0-47dc-beda-b6654d619d6a · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4c3a23-e484-4c81-82a6-2206d90e5e79 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c2eb88-de5c-4ea8-b3e8-4ef04ac48b49 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-evolving multi-agent collaboration networks for software development
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb119d41-1db2-491c-b863-2c7cd2db0aad · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SEW: Self-Evolving Agentic Workflows for Automated Code Generation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a438e6fe-bc99-4020-aadc-cd4a1d5f5693 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ecd6a5f-8ebc-4011-8b8c-1248324a388d · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? AlphaEvolve: A coding agent for scientific and algorithmic discovery
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a19d7cd-4a82-425a-8a93-a521749a1200 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotest: Evolutionary test-time learning for self-improving agentic systems
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3b0b20-44fa-43b8-a733-cfb5ed07a0be · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76f30c6-aa60-4f0a-8914-2712ff2475be · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73fce81a-b1bf-4152-9555-ee4a82465c9c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Your agent may misevolve: Emergent risks in self-evolving llm agents
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b8e116-6b94-4a8e-8c9e-eafca8f43605 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2935ea0f-764a-4a2a-8341-5d42cd473aad · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 188f85ae-edff-4a3f-aa09-91fc9df16cc3 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webvoyager: Building an end-to-end web agent with large multimodal models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf78fd0c-0c32-48c7-86ad-fc0033c497ef · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592bab4f-14f5-43de-8234-c160ebd7d284 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f05cad0a-6ce5-48e5-82e9-5d21a17a8a02 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c2e747-2463-4d6e-b47b-e9c0e5f07032 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc7197c-1adf-4548-8a30-f846e95381f2 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81771bf5-edef-4fed-94ad-808de9210a98 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e25965-c2c8-4b22-b36c-349ab6a533c4 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47a9d7b-74e2-4fe3-8fbd-963edfbf6074 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gaia: a benchmark for general ai assistants
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b512849c-9191-4ddc-a38b-e8badd0f1281 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0dc2de-6651-456f-a1aa-1c902b2c9a9f · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f015896-0a35-463c-a30a-23394db146bc · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0c0772-e8a4-4962-8322-b8fd75fcc637 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent workflow memory
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62aa1b8c-60ad-45aa-9b9b-2d5cc21f14e3 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? none identified
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3377049d-f4bc-4c76-8fb7-557fa03b9c9e · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? reasoning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a5d1a1-0e80-419f-9eef-e66ae93408dc · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Order from most to least important
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10004a5-26a9-4d39-95d1-277af5e60038 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90dc9e6e-7d6c-4c87-8624-ff0085958f7c · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Status:
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53aba8b1-c42c-45e0-9c21-a9a2afb989b4 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09cd7c8b-48af-4bba-ade3-477b10cbe395 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a09f38-aa8c-413d-a4b4-396c781e0e26 · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6984840c-21ed-47f3-a1db-0713b084276d · outbound
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.