Pith. sign in

Paper Citation Record · LEDGER

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

As of 7 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 0 inbound Pith citation observations for arXiv:2608.00155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00155 v1

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:39.559055Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 102 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67b611d2-1277-424a-96f0-bd6abc117c0c · outbound

This paper cites A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.678535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.678535Z digest=sha256:e5a0a0c8cd961e1aa20f91e8477f499631372f4bfa1cfa9b449ed71570b9c62c

Observation 15c3b5f3-0cc4-407d-97dd-53e6f149a015 · outbound

This paper cites Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.740362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.740362Z digest=sha256:d378875413953c41ddf2439e8b79e0af448a0b8bfd8e178988055260a8c51c92

Observation 774e17c1-81ba-4c06-8184-c3942a218a0c · outbound

This paper cites A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.844983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.844983Z digest=sha256:70f570b6271af1b64cc6c39d2d59a0d31ec1354bfbd850e82719619695e0c5e3

Observation f6d0b54b-6b3b-4420-ba41-a3bb96d97d7c · outbound

This paper cites Agentic context engineering: Evolving contexts for self-improving language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic context engineering: Evolving contexts for self-improving language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.945741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.945741Z digest=sha256:075c40093d15c577fa76a38f37cc394bbd1633eda9888fcf6b3770c7d1bd1d11

Observation 4b30595c-b4bf-4321-82f1-c1ac004ba677 · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.047941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.047941Z digest=sha256:bed5ae8a3e157a6ffceb598a364e3b44d806872e7cb10390c9c0f205e001bf0d

Observation 4d1cbd62-4880-457e-9d8e-76e4c51ce49d · outbound

This paper cites A-mem: Agentic memory for llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A-mem: Agentic memory for llm agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.168633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.168633Z digest=sha256:839df3d29a7cd14d465b6fe4e60eaaff2a5aec799f5513b6dba6d40cc975dbc0

Observation 5b4a68ac-830d-4563-8af2-bc2c42882100 · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.238737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.238737Z digest=sha256:9bee18f4d0b81dcced7bfba761a9f2d74541c67ec4ddded5e960ef49f46b5267

Observation 9997acb5-415b-4035-aced-f5f12b482392 · outbound

This paper cites Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.339997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.339997Z digest=sha256:ca40bee9667af670864313941de3d76317fc67d8aed3425c2ffb21b2b287e1d3

Observation 2ae5e5f3-dae5-48a1-9f88-e580831d52be · outbound

This paper cites Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.448963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.448963Z digest=sha256:1f53bc6d1de919ea3e48fddfabbb0163ba6a2e8fdbe1a26d1fa36af31b4a3049

Observation 669b2ad1-81a2-4075-bd74-1379c5d0b903 · outbound

This paper cites Memento-skills: Let agents design agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Memento-skills: Let agents design agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.553680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.553680Z digest=sha256:7859cb0adc972ead1174553d1bd6eb76c186adcef5956aeacba2628da78ebc9d

Observation aff5ba51-e69a-4fe0-958c-0ca5c67e030e · outbound

This paper cites Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.613003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.613003Z digest=sha256:f7fd7c2ea35ca9be95b5c682322c08c9a623386bb6cd1461da0243b03497410e

Observation 8650ca55-dd31-4e2b-b6c2-2e8233c3b8e2 · outbound

This paper cites Appworld: A controllable world of apps and people for benchmarking interactive coding agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Appworld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.709243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.709243Z digest=sha256:966888053c78650a776cbe5db6b8b75792c4716bbbc9a97bda483703cf5a4e58

Observation 0313bafd-50e0-4c88-95b5-24fa3f54c936 · outbound

This paper cites Gonzalez.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gonzalez

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.826496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.826496Z digest=sha256:0e30b7484dd1c45e8533742f10daa8a9a4f93be315d473d46f6d4f8265fc5089

Observation 192ac56d-3eb2-4e4d-966f-b596268400d9 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In Proc.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Swe-bench: Can language models resolve real-world github issues? In Proc

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.918060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.918060Z digest=sha256:b45c0d72200b94312e3ca12836034c43a26ad551be13e640f72ba5a06efe3b03

Observation c22fc5a9-a231-4699-b27a-21e74d18ae93 · outbound

This paper cites Humanity's Last Exam.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Humanity's Last Exam

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.996592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.996592Z digest=sha256:f9595c5fccadcc26baf4f97b0b49fabfe8987c3d00cc3f7ef858e9c4a694dee7

Observation 45ed4434-2e42-4aac-a302-79bb5630afb5 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.087301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.087301Z digest=sha256:9982f874bc390d6e99a5f3275b611861c44a9a0a98fa389872a4c6590ef543e4

Observation 0ab7ed43-9bc6-413b-8b5e-082f746f58f4 · outbound

This paper cites Stream- bench: Towards benchmarking continuous improvement of language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Stream- bench: Towards benchmarking continuous improvement of language agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.156197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.156197Z digest=sha256:db6aa72e5d099afd418ced2ec26827a73b10ef9cdad55cc750102b94a7e5c105

Observation 19de0ee3-c2cd-4dac-84f9-1710b859a36c · outbound

This paper cites Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.201860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.201860Z digest=sha256:beb0e6d8ceb3af034a3a79f817f3776f107457efee3ed9f0cb606b1099bf9dee

Observation b8cc89d3-74d9-495a-a8d0-64d85f8c11ec · outbound

This paper cites OpenAI GPT-5 System Card.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenAI GPT-5 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.267699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.267699Z digest=sha256:0828892db97177e42455733411db645526ace77606d0136f1a5f9e96ca238ba4

Observation 5ba34156-2f11-4e2f-be16-4be50f196bea · outbound

This paper cites Gemini 3.1 Pro model card, February 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gemini 3.1 Pro model card, February 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.359163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.359163Z digest=sha256:f3b9fc69cc814959853dcbe67b5189d85208dfba46738c1a4d9e62417fa2d45d

Observation b9ba58dd-a90c-4480-9b54-3e6fb70502a1 · outbound

This paper cites Introducing Claude Opus 4.7, April 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Introducing Claude Opus 4.7, April 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.441241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.441241Z digest=sha256:a3c7d22b3fe7052d3f6c6d77a26a86f8043fe108fd0bdb49c235b3fa0331fcb9

Observation 1e4e000b-6c7e-4e9c-b814-b4951cdeda99 · outbound

This paper cites Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.517774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.517774Z digest=sha256:765840b2e14622144dca85f95648a5e15d287154f5081aaa2efdf07b2cb063a6

Observation a9c9edd8-577d-471a-81e4-899bf8f7293c · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training with self-supervision for generalization under distribution shifts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.557621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.557621Z digest=sha256:dbf4a7385efd128456b850712a24630e5941cb81b7df912117218a82661cb21a

Observation 1ef7ec98-9596-452d-b766-b5a37395abac · outbound

This paper cites Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.616665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.616665Z digest=sha256:03caf0d1c0f99f07e96f227ce0a246dce263775b6d48437e1a0684d46adaf767

Observation 4ec2d7e4-606f-45df-bcbf-08412cdd5a1a · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.706158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.706158Z digest=sha256:8bbd1cdcb4bb9597b2f7f56bea05498764611eb35d381ed0d2ee9c70f41ba62e

Observation 7128a647-c5e2-4fc7-bee4-44afb379adaa · outbound

This paper cites Gradient episodic memory for continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gradient episodic memory for continual learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.785239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.785239Z digest=sha256:77e3fddcc24a72dd7ad6b5b8989dd680d990ec303744fcfec3cec7534fd11543

Observation 13b29ec8-639f-4f57-9e71-ec2d6c5b41ed · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Scaling llm test-time compute optimally can be more effective than scaling model parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.834609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.834609Z digest=sha256:99547abf816c038f292bccb43627bf19881c8d522445e73829bd946dab22b674

Observation ea24b3aa-c55c-429d-8f36-4f6fd65c80e7 · outbound

This paper cites Test-time training on nearest neighbors for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training on nearest neighbors for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.907311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.907311Z digest=sha256:8b250540d7e7877eefbb4516722dc2f8aaf293a4156c770d6f5eb3316490b8f0

Observation 69b76d30-5c03-4755-a817-87bedd5539bf · outbound

This paper cites Efficiently learning at test-time: Active fine-tuning of llms.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Efficiently learning at test-time: Active fine-tuning of llms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.947220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.947220Z digest=sha256:a602e705128f8391d33cc6554e67196c4635ca52d6f2a573d5ebc88080e9c1cb

Observation 8353d4e8-7813-4986-b32e-8b3fd29c1545 · outbound

This paper cites The surprising effectiveness of test-time training for few-shot learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The surprising effectiveness of test-time training for few-shot learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.990733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.990733Z digest=sha256:8f4360cb009f1c7e4c4720b9754e634330f2b61f2e63fc1116775f1b72257f4e

Observation 1aaae17b-f559-45b8-b4aa-042734c11429 · outbound

This paper cites In-place test-time training.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? In-place test-time training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.052856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.052856Z digest=sha256:2be1444f6f3176f3bfe3a4dbe9bdb4f7f1c85089fede59a50564b844f0121d17

Observation 78bb23fd-0340-46e7-a979-5c9cb6c60f57 · outbound

This paper cites Test-time adaptation for llm agents via environment interaction.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time adaptation for llm agents via environment interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.147161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.147161Z digest=sha256:528c94395533bf2ec7bc71af12d5860abbad78d94bddc3fa74933d7f3ec20d35

Observation 8c6d34b6-56e8-42fd-be9c-40b104d6fb5b · outbound

This paper cites Test-time learning for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time learning for large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.266388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.266388Z digest=sha256:ec85b5c2ec3ef892d96b4340c5bf6304722038f6cbb139e07fdc91639feffcff

Observation 31c98bfe-fc30-4da1-b482-26138dfd8ce1 · outbound

This paper cites Ttrl: Test-time reinforcement learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttrl: Test-time reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.347224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.347224Z digest=sha256:eb49a80c3a66b4317c8b77616b19dd745991bde99dce8f94a359b4cf281a13e2

Observation 560c579e-c5e5-4904-893b-432431fa2cb1 · outbound

This paper cites Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.407176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.407176Z digest=sha256:39e931f0830d687ef57d43de1dfc5141d58d4c738960fe625401a3ff0247e7af

Observation 729ca982-0b5d-48a6-bc45-3aa1aa448f4f · outbound

This paper cites Learning to discover at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to discover at test time

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.486960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.486960Z digest=sha256:25e95fc8639b1a288c9f044dca512962277db8948a2c0c58a04ba6ca6f656b63

Observation 0dbd5315-bd55-49ce-a095-a6213b44dc44 · outbound

This paper cites Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.584703Z digest=sha256:8f3eb8893dfea065b83c3674a732c1df73bbb6fdb3cc78778bfb421c463919f4

Observation 8a992a21-8df4-4314-91b8-8d6506facb9c · outbound

This paper cites What if consensus lies? selective-complementary reinforcement learning at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? What if consensus lies? selective-complementary reinforcement learning at test time

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.672842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.672842Z digest=sha256:312acec6e7d9fb6fe4abf8cb27a96a8a8bdbe305d9c0f3b24246be8f7825ff1a

Observation 0cb42fec-7ea7-47d8-add2-4a5c603aa6c1 · outbound

This paper cites Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.783687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.783687Z digest=sha256:9302795eba18b83ac13d7c38fed56f54e09e9973dab4dbe8e59af483b697a4c1

Observation d76d5b23-7867-48c6-ba76-c7b3adbf5fd3 · outbound

This paper cites Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.899505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.899505Z digest=sha256:e774000fb7296979ff06e44fd08d23eb76fb146a0306ed15cb7f2e642656fa7c

Observation b01a108e-71ff-4206-a9bc-190f1a4e571e · outbound

This paper cites Test-Time Learning with an Evolving Library.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-Time Learning with an Evolving Library

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.009288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.009288Z digest=sha256:509297f61a6baf2696d25b714f490b36dbe7c5eea2432b632d9d24895e3c30d5

Observation eb9d85b7-402a-49fc-aec2-7efb0b250981 · outbound

This paper cites Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.114246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.114246Z digest=sha256:573f34b319a060df7243d08a1142eb1554cc7917de11b21c83a6a1cdb1d3a9b5

Observation 4b2262f2-6162-4fac-abde-a3da7f3f3208 · outbound

This paper cites Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.226463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.226463Z digest=sha256:cb40c445c7c4bdcb7be58ceac47cb7a36c6e378569b16379c99d0a48757aaf50

Observation 26fb1c9a-50dd-4827-84ac-1aa5afaa48dd · outbound

This paper cites Agentic plan caching: Test-time memory for fast and cost-efficient llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic plan caching: Test-time memory for fast and cost-efficient llm agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.337320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.337320Z digest=sha256:463872e66bf6aed8d11d60fb3845a7f2969449118d4d177028e27a4ba29debff

Observation ed2c9bfc-f745-4cf6-b322-a65549161db3 · outbound

This paper cites TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.442077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.442077Z digest=sha256:71ad229d3e7ea6447963d85e8a0efa6ea7c99f440e9e98083198620419bcc84e

Observation 18107843-204b-4eb5-b983-e6c2028238e2 · outbound

This paper cites Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.540281Z digest=sha256:1649659914f704894b0f47ff8cee55fb22f2541507679183a86dda28cd905c9b

Observation b09b02a8-d8b8-453a-b705-88ffabf8a7c3 · outbound

This paper cites Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.652568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.652568Z digest=sha256:03808b107907733511e7006ac0ba4229a325b45ee29a8876c9bbe76724df6b35

Observation 9f5ea4cc-4e98-4228-9cbd-eb3796899c71 · outbound

This paper cites Panini: Continual learning in token space via structured memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Panini: Continual learning in token space via structured memory

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.761418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.761418Z digest=sha256:c361533d4db0b8aa07bacc9a3859df366e1bbcf2df950024e6c000633ec47622

Observation 5a6a2359-872e-42b0-bcc8-368eb0a6533b · outbound

This paper cites Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.866944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.866944Z digest=sha256:a3db01d4c44e4d67d577c3c6ab3f24d39c77d9ef3f5c79390b449377a4e2ada5

Observation df12cecf-038a-4dc1-8a2b-c5be36a6d665 · outbound

This paper cites Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.939316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.939316Z digest=sha256:d044ef6ae9fed621f8928479161c8c22a33de62631fdd470c1e99906f9093a08

Observation fec8e04f-7212-49fe-a4ed-8b154e68aa71 · outbound

This paper cites Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.017911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.017911Z digest=sha256:f0a8ba75a368da665f6cd2c43719a8345def167c80d67cf8b454dd339da8e86c

Observation fb473b65-f10e-480b-bb9b-04a615958777 · outbound

This paper cites Xskill: Continual learning from experience and skills in multimodal agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xskill: Continual learning from experience and skills in multimodal agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.087726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.087726Z digest=sha256:7239cebfb7697c338a5953dece70d7a5155bdf7c49c40532ad8e466afbc6efc1

Observation da43f976-c6ce-463c-9799-361ce994f6aa · outbound

This paper cites Online Experiential Learning for Language Models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Online Experiential Learning for Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.159727Z digest=sha256:73a9bc5727e35db174f55f03c0d837206051a5071537386bd9195b43f348e718

Observation ddc78d07-c649-491c-9b06-7196fe1fb0e7 · outbound

This paper cites Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.227678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.227678Z digest=sha256:0ed138387dea6173c7ed86e8f6bef57bd040eb4d3f606e87118893d5814fbfcd

Observation d42845b4-b90e-46aa-b1b5-2dbf86837931 · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.336093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.336093Z digest=sha256:fcdeee5c1099dc039915d4e08cd62781a7bde3a4a1173a67fa64850055d7c766

Observation 227b17f5-290a-48c0-94bd-9da4a8674c98 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.459460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.459460Z digest=sha256:c496a04287707011e0fc468fec2201fea81830424bdb3fccabe85d279d112603

Observation e6271b35-7585-4cff-ac8d-17152f2132e8 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.579166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.579166Z digest=sha256:f0f49d438181b57c0dcc701e45918ea99eaf710bbd69e5cba9a524be8b92a4ce

Observation 1fcce961-c4c5-4a04-8cba-613fc43b0b43 · outbound

This paper cites SkillOS: Learning Skill Curation for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOS: Learning Skill Curation for Self-Evolving Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.690973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.690973Z digest=sha256:f4a2d11e454a2895fd29ffdf72f7216ce702c8fc4ea49dd9c62184bbb695b0f3

Observation 60d83db8-2aa5-4635-b235-d2f7d73e311e · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.780877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.780877Z digest=sha256:c1da33d46c56e7922337b6c6c488d70036144a06f03ad30432df8d2e73bcbdff

Observation 77a4cd23-7cf0-42d9-bfbc-c364071b1d2d · outbound

This paper cites OpenSkill: Open-World Self-Evolution for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenSkill: Open-World Self-Evolution for LLM Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.900154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.900154Z digest=sha256:7ab883e6152bc60dc7f8386a3ab70ea9135631514b9f7bc214e14c98cb4ab1af

Observation 3406787a-4e71-4b33-9ab7-b8a54d9deec6 · outbound

This paper cites SkillOpt: Executive Strategy for Self-Evolving Agent Skills.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.013114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.013114Z digest=sha256:5b19027ba32b41af85b3790762c8bcc8cd5fc7a8c0d9840eb31426933e9bbc60

Observation 9f8206b5-8e87-4d2a-82ea-3561f5a75200 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.133344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.133344Z digest=sha256:b9fce741056d78be83833897af46986f5cc66c2eee6e1527088e8b48faa54b2f

Observation f4deb2ae-ff8d-44e6-a47e-4bf4271f2963 · outbound

This paper cites Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.247731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.247731Z digest=sha256:2dfa011dfcbfd107ef27bb46a2d86a597b0559cdafea9e16702cf1c989858d4b

Observation fd65faba-7aaf-4848-ae01-57b277ee676c · outbound

This paper cites Selaur: Self evolving llm agent via uncertainty-aware rewards.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Selaur: Self evolving llm agent via uncertainty-aware rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.328391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.328391Z digest=sha256:0eb1e9f7858cec55f7d9b9ea198ec3e42ed460abeb4a2ad2cb1155e3aed014d1

Observation 78298233-2985-415e-a37a-9f4e5254b1ca · outbound

This paper cites Self-Improving Language Models with Bidirectional Evolutionary Search.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-Improving Language Models with Bidirectional Evolutionary Search

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.464200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.464200Z digest=sha256:8d9adce8c1bb63045c1661a2b35d0a9979df5be4bf947a41871e877f861d4495

Observation 42b75e4e-19e0-460c-86e8-ec6acc6fd764 · outbound

This paper cites Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.602393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.602393Z digest=sha256:7544e4eb8895639dbcc885f829838f5f12ea31c19f2250fb2c0313cb054f8d40

Observation c2e9eb84-99d0-47dc-beda-b6654d619d6a · outbound

This paper cites Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.737392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.737392Z digest=sha256:621807e30c828834a94c312a2ebb86fec1bd639efe53cb261528ff680b47ad3a

Observation 8d4c3a23-e484-4c81-82a6-2206d90e5e79 · outbound

This paper cites Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.901016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.901016Z digest=sha256:1ad78c5b6e48292df609d424ca64ee9e7c248065dbed8f7aca46743dff1341d8

Observation 81c2eb88-de5c-4ea8-b3e8-4ef04ac48b49 · outbound

This paper cites Self-evolving multi-agent collaboration networks for software development.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-evolving multi-agent collaboration networks for software development

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.006298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.006298Z digest=sha256:1c3c252855fec9c7a171d1e848db306c5626486fe8fac7c4e33d8290ed030c39

Observation fb119d41-1db2-491c-b863-2c7cd2db0aad · outbound

This paper cites SEW: Self-Evolving Agentic Workflows for Automated Code Generation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SEW: Self-Evolving Agentic Workflows for Automated Code Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.085301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.085301Z digest=sha256:55d25d5b3bd02f84c84f15d528985bb3fc11124c3d4dc864e27a496e15edd520

Observation a438e6fe-bc99-4020-aadc-cd4a1d5f5693 · outbound

This paper cites Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.179134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.179134Z digest=sha256:c8dbbbe51fcc93a6cfed7fd60097fd20cc41664ef8637053fb69dad12939814c

Observation 3ecd6a5f-8ebc-4011-8b8c-1248324a388d · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.248349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.248349Z digest=sha256:6dac40bf279600542a531c91c42d907417b89ba025c2d9962795edff47b05698

Observation 6a19d7cd-4a82-425a-8a93-a521749a1200 · outbound

This paper cites Evotest: Evolutionary test-time learning for self-improving agentic systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotest: Evolutionary test-time learning for self-improving agentic systems

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.318781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.318781Z digest=sha256:fbd0ab67ea9707be9b0f1bb14e389ef6d42945aab8032eb289da90b0017580fa

Observation 3a3b0b20-44fa-43b8-a733-cfb5ed07a0be · outbound

This paper cites Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.425048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.425048Z digest=sha256:2bccffdfeed3f076a6d83f1b27fc696192ef95317fb346a7378457a397b4410e

Observation b76f30c6-aa60-4f0a-8914-2712ff2475be · outbound

This paper cites Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.498006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.498006Z digest=sha256:e15499434913f552a72fbeb8274a475abe4c903b7eba1ed45f18e1e0cbfd2dc7

Observation 73fce81a-b1bf-4152-9555-ee4a82465c9c · outbound

This paper cites Your agent may misevolve: Emergent risks in self-evolving llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Your agent may misevolve: Emergent risks in self-evolving llm agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.591245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.591245Z digest=sha256:50328ee964707a3fe8e3e8b17d6cf0ea474cfc986db9026185fbd38b799f69a8

Observation d6b8e116-6b94-4a8e-8c9e-eafca8f43605 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.698493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.698493Z digest=sha256:47deeba8e88f8a72bda8635a416fb8a7045b2f862b9da488331de3b97267fdc5

Observation 2935ea0f-764a-4a2a-8341-5d42cd473aad · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.811240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.811240Z digest=sha256:101b13807a33189035961a0a0a28cdd0a878d202d99af44088028cc25460e1ab

Observation 188f85ae-edff-4a3f-aa09-91fc9df16cc3 · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webvoyager: Building an end-to-end web agent with large multimodal models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.916387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.916387Z digest=sha256:36bc9d2fa4d9381263938645c050fdceddf36b3b1d022843c291ca04ed05729e

Observation bf78fd0c-0c32-48c7-86ad-fc0033c497ef · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.996496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.996496Z digest=sha256:5ee7660232dc3ea189f744d923e2e94f58e5598662a8e7a02219cbc7e4d9cce5

Observation 592bab4f-14f5-43de-8234-c160ebd7d284 · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.075464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.075464Z digest=sha256:af9aa437ad8b4d5c986e98f47870843cf5c73c78cfe04ebdeddbca8adbbc2338

Observation f05cad0a-6ce5-48e5-82e9-5d21a17a8a02 · outbound

This paper cites The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.158997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.158997Z digest=sha256:9b35e138781d2a77cef9517a2f2d6addca2d24937be3be24cbc917c43903bfb0

Observation 02c2e747-2463-4d6e-b47b-e9c0e5f07032 · outbound

This paper cites Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.214251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.214251Z digest=sha256:2731c0c3ac6151b50675b55e8148d8e24b077318b8c379b6a8c03369ba7ba3a6

Observation 4cc7197c-1adf-4548-8a30-f846e95381f2 · outbound

This paper cites Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.287770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.287770Z digest=sha256:af56bdf76384d6b5e03212412f1058ccd28d8b892511d682855897f8821d3b44

Observation 81771bf5-edef-4fed-94ad-808de9210a98 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.372036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.372036Z digest=sha256:5a6c02b104d697499a0c58b93b7c864f05c29c5efaaf78d3c4a6bde8807d45bc

Observation b6e25965-c2c8-4b22-b36c-349ab6a533c4 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.455621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.455621Z digest=sha256:ed103a95a839b71dc8e524f1ddec90ed64b49ff79a0c8f75d84c23aaf995b6e7

Observation b47a9d7b-74e2-4fe3-8fbd-963edfbf6074 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gaia: a benchmark for general ai assistants

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.536812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.536812Z digest=sha256:eb306128ec441ba89c1e0bc8d160945ace0e9908256c27d369f8f7db1aecd3d0

Observation b512849c-9191-4ddc-a38b-e8badd0f1281 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.625544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.625544Z digest=sha256:c6a2e116af72141017561883154f89a30cbf2fcf5ee648fc3c5a44f676ade171

Observation 6f0dc2de-6651-456f-a1aa-1c902b2c9a9f · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.705418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.705418Z digest=sha256:0c0b39229b2185b2d03a21190d8ade907ba896d4080a29df4005530e21e7e62d

Observation 9f015896-0a35-463c-a30a-23394db146bc · outbound

This paper cites Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.779498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.779498Z digest=sha256:37e2944a70c1ee069aaa828b4118697de674ed468e1d011a51b319417079f408

Observation ba0c0772-e8a4-4962-8322-b8fd75fcc637 · outbound

This paper cites Agent workflow memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent workflow memory

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.857487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.857487Z digest=sha256:d376f4f7b1815c06c9c74974e1490cce56baaa453ae10a24d030a1c2d43226a0

Observation 62aa1b8c-60ad-45aa-9b9b-2d5cc21f14e3 · outbound

This paper cites none identified.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? none identified

Reference 92

Resolution
malformed identifier
no resolver link, observed 2026-08-04T01:12:38.939920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.939920Z digest=sha256:53efc89ecdaf7d376615fcbac7d7c9e276467e1822a6bb17eb01146a06c8e565

Observation 3377049d-f4bc-4c76-8fb7-557fa03b9c9e · outbound

This paper cites reasoning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? reasoning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.018297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.018297Z digest=sha256:f3d313a821c3913722d84f47113c8b2d4aeb2eadfd534b230314578915c6fce5

Observation e2a5d1a1-0e80-419f-9eef-e66ae93408dc · outbound

This paper cites Order from most to least important.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Order from most to least important

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.092541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.092541Z digest=sha256:d50211b2ae00ce19fcb261c7fe43273facc7921ce76213cc9327145c2b053c7c

Observation f10004a5-26a9-4d39-95d1-277af5e60038 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.153063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.153063Z digest=sha256:03b9c62abcc0b64117665e732a3f2a8a82fe7a09cc799e7ce2753f507794254a

Observation 90dc9e6e-7d6c-4c87-8624-ff0085958f7c · outbound

This paper cites Status:.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Status:

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.230523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.230523Z digest=sha256:0dde2f21a9f7fdeee077221a53a78558441dc6bc075c429b3ecbe2641ee17244

Observation 53aba8b1-c42c-45e0-9c21-a9a2afb989b4 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.310212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.310212Z digest=sha256:bde3791d045dcbe46963e3194aa81afd8d644e64fef3ae0cc59cff2b463bb3e0

Observation 09cd7c8b-48af-4bba-ade3-477b10cbe395 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.419533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.419533Z digest=sha256:32aecaa55835d4c62e8b13f332ab7e71e6f197e193925f11026fdeefbf172e3a

Observation 00a09f38-aa8c-413d-a4b4-396c781e0e26 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.476352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.476352Z digest=sha256:9475ab64f6b02ee79059325041ae49ad5c3a57a943b83fb2d522b1535e03533a

Observation 6984840c-21ed-47f3-a1db-0713b084276d · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.559055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.559055Z digest=sha256:910ba9e174d7c893195ebaca1000b2be866258cc67c90825c60782ead9f074eb

Pith citing papers

No inbound Pith citation observations are available.