Pith. sign in

Paper Citation Record · LEDGER

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

As of 21 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 0 inbound Pith citation observations for arXiv:2608.00155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00155 v1

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:39.559055Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 102 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67b611d2-1277-424a-96f0-bd6abc117c0c · outbound

This paper cites A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.678535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.678535Z digest=sha256:52d6640d1d6b27a99c05d600d37966e56e744f2e7bf6b71e60e4b111ba6de34a

Observation 15c3b5f3-0cc4-407d-97dd-53e6f149a015 · outbound

This paper cites Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.740362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.740362Z digest=sha256:9fb26443c71c1ec6565af4780ec04ca564d8d28b40c0c37ed2ff916f09f76fd0

Observation 774e17c1-81ba-4c06-8184-c3942a218a0c · outbound

This paper cites A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.844983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.844983Z digest=sha256:de8b02ed68e9dc3a3494a46196356d588eb2a03c9e80a221d45956f11b290bce

Observation f6d0b54b-6b3b-4420-ba41-a3bb96d97d7c · outbound

This paper cites Agentic context engineering: Evolving contexts for self-improving language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic context engineering: Evolving contexts for self-improving language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.945741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.945741Z digest=sha256:cee519145ec023e191c609b257bbb793d0ba4bac6523c309e1b8d7a8768dd3e7

Observation 4b30595c-b4bf-4321-82f1-c1ac004ba677 · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.047941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.047941Z digest=sha256:974d93f4d1abe9e21c86a9646f9665279433ad9375cf09c594e8522c49cf2a07

Observation 4d1cbd62-4880-457e-9d8e-76e4c51ce49d · outbound

This paper cites A-mem: Agentic memory for llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A-mem: Agentic memory for llm agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.168633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.168633Z digest=sha256:8b5e9a912a1f3ab49045ff09f34e6b1b822ab95d9781d056d5e5a3bbfe2f2d86

Observation 5b4a68ac-830d-4563-8af2-bc2c42882100 · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.238737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.238737Z digest=sha256:79dc1e6a3135e9060374f18f4227812992be26e72d71d9a5c82bc19f1a6c3f7b

Observation 9997acb5-415b-4035-aced-f5f12b482392 · outbound

This paper cites Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.339997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.339997Z digest=sha256:0263e4acb9df08bc8fb2ed5cac4943e4cd6dc349367a706659209feb538890d5

Observation 2ae5e5f3-dae5-48a1-9f88-e580831d52be · outbound

This paper cites Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.448963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.448963Z digest=sha256:f32a46af0190c1e289dc27abc68b7f5793795aa0b8a30d621fa31675f01faa59

Observation 669b2ad1-81a2-4075-bd74-1379c5d0b903 · outbound

This paper cites Memento-skills: Let agents design agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Memento-skills: Let agents design agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.553680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.553680Z digest=sha256:26386522b4c27227d1228665d5909b9871faa830f5a1eda3ea2e9fac38dbcee8

Observation aff5ba51-e69a-4fe0-958c-0ca5c67e030e · outbound

This paper cites Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.613003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.613003Z digest=sha256:232e7c216b161c18445446591ce0003b71f51ac05d0208a93ed6da65ebcad3b2

Observation 8650ca55-dd31-4e2b-b6c2-2e8233c3b8e2 · outbound

This paper cites Appworld: A controllable world of apps and people for benchmarking interactive coding agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Appworld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.709243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.709243Z digest=sha256:354989745afc12a5f9adefeb94b726e347ae64129afb6c2d74a2d5f916fe8335

Observation 0313bafd-50e0-4c88-95b5-24fa3f54c936 · outbound

This paper cites Gonzalez.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gonzalez

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.826496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.826496Z digest=sha256:27a75ce23bb9413e597180dcfa50dead39072e04fc4b5894e327a4c8eaf7511e

Observation 192ac56d-3eb2-4e4d-966f-b596268400d9 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In Proc.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Swe-bench: Can language models resolve real-world github issues? In Proc

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.918060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.918060Z digest=sha256:693420df6056d9a2bbdbd077abd96294f77e75008315ec82788e2a743633b3e2

Observation c22fc5a9-a231-4699-b27a-21e74d18ae93 · outbound

This paper cites Humanity's Last Exam.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Humanity's Last Exam

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.996592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.996592Z digest=sha256:9cf9c3179191d2447f0b42ddbd908a6d1757eedca486ec4064673a8f5870ba6d

Observation 45ed4434-2e42-4aac-a302-79bb5630afb5 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.087301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.087301Z digest=sha256:8703981514c2355c60565c808786a93854ef880cce2420a90de3978fc8c18c29

Observation 0ab7ed43-9bc6-413b-8b5e-082f746f58f4 · outbound

This paper cites Stream- bench: Towards benchmarking continuous improvement of language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Stream- bench: Towards benchmarking continuous improvement of language agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.156197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.156197Z digest=sha256:d36b48d45b0f0c6fe4fee5f40b646a509a5ebc9c60fa37dbf043099c3605b80a

Observation 19de0ee3-c2cd-4dac-84f9-1710b859a36c · outbound

This paper cites Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.201860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.201860Z digest=sha256:5b91eee6d09dae2a8753041b1755da47fbf52142ffc967ad46496bf2d935fead

Observation b8cc89d3-74d9-495a-a8d0-64d85f8c11ec · outbound

This paper cites OpenAI GPT-5 System Card.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenAI GPT-5 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.267699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.267699Z digest=sha256:1883fabd4d9f526a25aeae6c40e3e290097898c347412c7b8048d911019a1ac4

Observation 5ba34156-2f11-4e2f-be16-4be50f196bea · outbound

This paper cites Gemini 3.1 Pro model card, February 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gemini 3.1 Pro model card, February 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.359163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.359163Z digest=sha256:d7166f3199ef1e6aa3591fab063f45f995ca8c5f94661d873747504e1cd13f59

Observation b9ba58dd-a90c-4480-9b54-3e6fb70502a1 · outbound

This paper cites Introducing Claude Opus 4.7, April 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Introducing Claude Opus 4.7, April 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.441241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.441241Z digest=sha256:cf0d000d745798fc2017c599a35d7fc6d11d42f38446fcce1a1c75dd631e1755

Observation 1e4e000b-6c7e-4e9c-b814-b4951cdeda99 · outbound

This paper cites Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.517774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.517774Z digest=sha256:be3f6741871ff33c3ff66d173d8992835d2011aee6a547d3da122514a2d9d59e

Observation a9c9edd8-577d-471a-81e4-899bf8f7293c · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training with self-supervision for generalization under distribution shifts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.557621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.557621Z digest=sha256:2cd3358cab2ce40a3136f5cec811c70e978c26830bafc7a76860ff0235e85034

Observation 1ef7ec98-9596-452d-b766-b5a37395abac · outbound

This paper cites Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.616665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.616665Z digest=sha256:702885f95968d8737c191c04f99e17b4f9f1d95d3136b827dfee16c972d67736

Observation 4ec2d7e4-606f-45df-bcbf-08412cdd5a1a · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.706158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.706158Z digest=sha256:7589e32a7757d4d2164efb590df7229b5e21ca6f4c978596cdea68af85f17c97

Observation 7128a647-c5e2-4fc7-bee4-44afb379adaa · outbound

This paper cites Gradient episodic memory for continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gradient episodic memory for continual learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.785239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.785239Z digest=sha256:b219fbeadd92d8d8b4ce6e0b34f4b197ebcd661742ab4337362ae4fcc5e8a3d8

Observation 13b29ec8-639f-4f57-9e71-ec2d6c5b41ed · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Scaling llm test-time compute optimally can be more effective than scaling model parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.834609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.834609Z digest=sha256:953cb671f387e5f1f6ffb9bb1d5a60e8ccefd6718c8c15eae00f627c13fe9f69

Observation ea24b3aa-c55c-429d-8f36-4f6fd65c80e7 · outbound

This paper cites Test-time training on nearest neighbors for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training on nearest neighbors for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.907311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.907311Z digest=sha256:3b81ea01d040d38d2930e0056ee74134ef2447550018afcffcd05fec92c47c57

Observation 69b76d30-5c03-4755-a817-87bedd5539bf · outbound

This paper cites Efficiently learning at test-time: Active fine-tuning of llms.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Efficiently learning at test-time: Active fine-tuning of llms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.947220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.947220Z digest=sha256:bea5004445ffc1a6f1ba8ce7e91b1c588e623f9456acd38cfed393a5996172c5

Observation 8353d4e8-7813-4986-b32e-8b3fd29c1545 · outbound

This paper cites The surprising effectiveness of test-time training for few-shot learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The surprising effectiveness of test-time training for few-shot learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.990733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.990733Z digest=sha256:fb88ac9e6704553c9a0725fb1661bc68353c54f2b105d312b673a8f7859cc4cb

Observation 1aaae17b-f559-45b8-b4aa-042734c11429 · outbound

This paper cites In-place test-time training.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? In-place test-time training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.052856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.052856Z digest=sha256:f364a3f0325552aeb8be5a3b0d6b4d5fa16591ee370e116fcbcf8b40b0f12d56

Observation 78bb23fd-0340-46e7-a979-5c9cb6c60f57 · outbound

This paper cites Test-time adaptation for llm agents via environment interaction.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time adaptation for llm agents via environment interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.147161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.147161Z digest=sha256:2a54d4efe090784443a355535af941c12a241dd4b10fde907628a80d9d813147

Observation 8c6d34b6-56e8-42fd-be9c-40b104d6fb5b · outbound

This paper cites Test-time learning for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time learning for large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.266388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.266388Z digest=sha256:194f76434e9504facf97712acc3854fe9a54a9b7841b10999142be2f3455912d

Observation 31c98bfe-fc30-4da1-b482-26138dfd8ce1 · outbound

This paper cites Ttrl: Test-time reinforcement learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttrl: Test-time reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.347224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.347224Z digest=sha256:e47fd1ff4452969eea836a6a40b1d90ae3fa0f0e97172cc3cf514f186e86ebc1

Observation 560c579e-c5e5-4904-893b-432431fa2cb1 · outbound

This paper cites Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.407176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.407176Z digest=sha256:c316265704a30e5b7fa7e0877397ec0be652461dc690bc4ba9092c50af368d1c

Observation 729ca982-0b5d-48a6-bc45-3aa1aa448f4f · outbound

This paper cites Learning to discover at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to discover at test time

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.486960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.486960Z digest=sha256:9984a3a49b5f4320e0297153fc608d4efd70ba4ebcb9f8f6617a800972578d2a

Observation 0dbd5315-bd55-49ce-a095-a6213b44dc44 · outbound

This paper cites Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.584703Z digest=sha256:6e2f87c203e309d2bb1d902d58a6d483318dd3fb33e812c157d6cc3190127338

Observation 8a992a21-8df4-4314-91b8-8d6506facb9c · outbound

This paper cites What if consensus lies? selective-complementary reinforcement learning at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? What if consensus lies? selective-complementary reinforcement learning at test time

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.672842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.672842Z digest=sha256:672bc378397cc7be9c2db3d6f31945bfb58a4d014aa878a97fa9b5a2da6f8c15

Observation 0cb42fec-7ea7-47d8-add2-4a5c603aa6c1 · outbound

This paper cites Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.783687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.783687Z digest=sha256:2df8dd549f825ddf242427e3c26ec1a4aba3d4c35ff01c5332bba7b0d32ebec4

Observation d76d5b23-7867-48c6-ba76-c7b3adbf5fd3 · outbound

This paper cites Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.899505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.899505Z digest=sha256:391a546a07a28441f8104c573e69e6172a55e52ecb99d5b781f4fdf2798c3288

Observation b01a108e-71ff-4206-a9bc-190f1a4e571e · outbound

This paper cites Test-Time Learning with an Evolving Library.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-Time Learning with an Evolving Library

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.009288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.009288Z digest=sha256:314fbdff75c88dc1cbe5a29d2cbbd2be6fda6d272fd7e23f50b59cac2b44c608

Observation eb9d85b7-402a-49fc-aec2-7efb0b250981 · outbound

This paper cites Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.114246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.114246Z digest=sha256:290319c16d28c7ad4a2fb169962bb8d54955c94da187bbc510526094ac1b28f9

Observation 4b2262f2-6162-4fac-abde-a3da7f3f3208 · outbound

This paper cites Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.226463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.226463Z digest=sha256:d3abb5c5d188217a9f0f47b2cb9a39225901d961c39e2f7047e93ba829fdac89

Observation 26fb1c9a-50dd-4827-84ac-1aa5afaa48dd · outbound

This paper cites Agentic plan caching: Test-time memory for fast and cost-efficient llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic plan caching: Test-time memory for fast and cost-efficient llm agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.337320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.337320Z digest=sha256:ec09ed20143168f6b3aac4dc391f814d26868ba1f0e291a84c0d64d2b6f6bf81

Observation ed2c9bfc-f745-4cf6-b322-a65549161db3 · outbound

This paper cites TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.442077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.442077Z digest=sha256:8d116d2120fa2c177a3400da3ab1fdbc02be93c6eeea76485505527f11457a11

Observation 18107843-204b-4eb5-b983-e6c2028238e2 · outbound

This paper cites Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.540281Z digest=sha256:8db0eff00ea6cf881754c21480207b8c6427c7b5fb4dfebd549afaff5d4a2fb6

Observation b09b02a8-d8b8-453a-b705-88ffabf8a7c3 · outbound

This paper cites Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.652568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.652568Z digest=sha256:44714c9864f709a3b23f2190bb6f601938fe56227445069d3b4c5b058aa55943

Observation 9f5ea4cc-4e98-4228-9cbd-eb3796899c71 · outbound

This paper cites Panini: Continual learning in token space via structured memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Panini: Continual learning in token space via structured memory

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.761418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.761418Z digest=sha256:ae9b42ef1107cfc9799c6b0afe5dbd4d023199b6470281f1d53a9dfec2ca8951

Observation 5a6a2359-872e-42b0-bcc8-368eb0a6533b · outbound

This paper cites Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.866944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.866944Z digest=sha256:98212da5329da6808189a11bef742a854d6535dc31ba6387cbe373277a637239

Observation df12cecf-038a-4dc1-8a2b-c5be36a6d665 · outbound

This paper cites Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.939316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.939316Z digest=sha256:e0cff01d81b0e4f310df19b83dc219d58ca9538b61f8334f2222a50cab9cf01c

Observation fec8e04f-7212-49fe-a4ed-8b154e68aa71 · outbound

This paper cites Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.017911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.017911Z digest=sha256:63ff82eb7b98efd0dcc688dacf377dc3b86d4196b7c4553ac1875e371e95ee70

Observation fb473b65-f10e-480b-bb9b-04a615958777 · outbound

This paper cites Xskill: Continual learning from experience and skills in multimodal agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xskill: Continual learning from experience and skills in multimodal agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.087726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.087726Z digest=sha256:47c9d5fa231732f2846a72187f18f47647ab23107004fa394652c26d68f76650

Observation da43f976-c6ce-463c-9799-361ce994f6aa · outbound

This paper cites Online Experiential Learning for Language Models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Online Experiential Learning for Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.159727Z digest=sha256:0b95f603415625a0bb83a5ee3a78b0b54018e0a17decb0fe0d9f2d6145c050f7

Observation ddc78d07-c649-491c-9b06-7196fe1fb0e7 · outbound

This paper cites Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.227678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.227678Z digest=sha256:12656f04ad3aec0896cfb020e48caf2e4644fef4745cf8e33867aaeb69de16d1

Observation d42845b4-b90e-46aa-b1b5-2dbf86837931 · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.336093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.336093Z digest=sha256:b735419cf534f91efe965acd3c155ea71728cc866a0fd0890c540d7dec44618b

Observation 227b17f5-290a-48c0-94bd-9da4a8674c98 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.459460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.459460Z digest=sha256:989bcfa5dbb165a120285969be78a9fc438081f3e9b1693295ab4cf39cfe5c46

Observation e6271b35-7585-4cff-ac8d-17152f2132e8 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.579166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.579166Z digest=sha256:bce5cbfc343422eba58c7ddcc997333309fae657f0e64701a4f82f98389d1ac4

Observation 1fcce961-c4c5-4a04-8cba-613fc43b0b43 · outbound

This paper cites SkillOS: Learning Skill Curation for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOS: Learning Skill Curation for Self-Evolving Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.690973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.690973Z digest=sha256:cabc467de9036db5b9a453f9e7518b57d33a641a687f159cf70dacee5d372462

Observation 60d83db8-2aa5-4635-b235-d2f7d73e311e · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.780877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.780877Z digest=sha256:2db6c63b5b240b5b8ab017feed49c6e64b9e5ef1513319575c644faedce3b592

Observation 77a4cd23-7cf0-42d9-bfbc-c364071b1d2d · outbound

This paper cites OpenSkill: Open-World Self-Evolution for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenSkill: Open-World Self-Evolution for LLM Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.900154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.900154Z digest=sha256:120d47c499a48dedadd5f058b3644eca868cf903d461cb24c50cd4cca6226896

Observation 3406787a-4e71-4b33-9ab7-b8a54d9deec6 · outbound

This paper cites SkillOpt: Executive Strategy for Self-Evolving Agent Skills.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.013114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.013114Z digest=sha256:70dafc624facefe24b6e7d4766e7465f2a490f538f1ccfeba6a8b482dc8acd08

Observation 9f8206b5-8e87-4d2a-82ea-3561f5a75200 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.133344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.133344Z digest=sha256:600c24855759c24f925bb1d621180e8b948ba17443481080931b23a231fc8081

Observation f4deb2ae-ff8d-44e6-a47e-4bf4271f2963 · outbound

This paper cites Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.247731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.247731Z digest=sha256:eb02eb23f793a0dfdd42631c4a54eb973e991959e166c50ea8bc95f5ea62f06d

Observation fd65faba-7aaf-4848-ae01-57b277ee676c · outbound

This paper cites Selaur: Self evolving llm agent via uncertainty-aware rewards.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Selaur: Self evolving llm agent via uncertainty-aware rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.328391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.328391Z digest=sha256:8a9ad09dab81d77253548af4527d1e38b2cd1a2999cabd6b2cc6382e75d47e83

Observation 78298233-2985-415e-a37a-9f4e5254b1ca · outbound

This paper cites Self-Improving Language Models with Bidirectional Evolutionary Search.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-Improving Language Models with Bidirectional Evolutionary Search

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.464200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.464200Z digest=sha256:0b34b84f758911509b42ccec19779d3f8b60237d86da649624ea1886928ee0af

Observation 42b75e4e-19e0-460c-86e8-ec6acc6fd764 · outbound

This paper cites Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.602393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.602393Z digest=sha256:2c5f88e78e96146de246a1479105378b656d25d294cd2636ed51c61e9b871a13

Observation c2e9eb84-99d0-47dc-beda-b6654d619d6a · outbound

This paper cites Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.737392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.737392Z digest=sha256:8d20f2759f37f6a78009ecea514f4ff72cc809da5b53e6363004bb449ef2bc69

Observation 8d4c3a23-e484-4c81-82a6-2206d90e5e79 · outbound

This paper cites Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.901016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.901016Z digest=sha256:04fdce304cfbb07a40a08f5767aa4d44bf9b49f595f0332328596ae30eb746e8

Observation 81c2eb88-de5c-4ea8-b3e8-4ef04ac48b49 · outbound

This paper cites Self-evolving multi-agent collaboration networks for software development.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-evolving multi-agent collaboration networks for software development

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.006298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.006298Z digest=sha256:7bc34bfbf5f79a480219b60f30ea9bd76febf086cd7c319d2df3a51605bd1aa8

Observation fb119d41-1db2-491c-b863-2c7cd2db0aad · outbound

This paper cites SEW: Self-Evolving Agentic Workflows for Automated Code Generation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SEW: Self-Evolving Agentic Workflows for Automated Code Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.085301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.085301Z digest=sha256:2285287666b77b089127f9ddc2828189a7ae7258b01029b2053e32978260ddf5

Observation a438e6fe-bc99-4020-aadc-cd4a1d5f5693 · outbound

This paper cites Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.179134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.179134Z digest=sha256:390bce9ec9cc30b272d08425194bd1bbfb6ba32fd899843f96b3aa473c4d90f8

Observation 3ecd6a5f-8ebc-4011-8b8c-1248324a388d · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.248349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.248349Z digest=sha256:d63dcec1938024d9ca2c735618e7e3a45e89a352ff97026fd1d76337b66e0419

Observation 6a19d7cd-4a82-425a-8a93-a521749a1200 · outbound

This paper cites Evotest: Evolutionary test-time learning for self-improving agentic systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotest: Evolutionary test-time learning for self-improving agentic systems

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.318781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.318781Z digest=sha256:a7d8b4fc01f8f5b0e1db7f7bdb91d1d4cecd73c3dd0692b89855bf32e2bd68fa

Observation 3a3b0b20-44fa-43b8-a733-cfb5ed07a0be · outbound

This paper cites Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.425048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.425048Z digest=sha256:e360022150bd7d277104a38c0c7b1ee02af1051499603a48f94a773e6a5383f6

Observation b76f30c6-aa60-4f0a-8914-2712ff2475be · outbound

This paper cites Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.498006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.498006Z digest=sha256:a6db9cd993951d745756b2d2f83d685c3d28f2ede91484569a0b282b707de4c8

Observation 73fce81a-b1bf-4152-9555-ee4a82465c9c · outbound

This paper cites Your agent may misevolve: Emergent risks in self-evolving llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Your agent may misevolve: Emergent risks in self-evolving llm agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.591245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.591245Z digest=sha256:ea1a1187bb9b97f4ba41bfdb72e913bc68a96ab0d055910fa6148cf29b0140e9

Observation d6b8e116-6b94-4a8e-8c9e-eafca8f43605 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.698493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.698493Z digest=sha256:91221146f39b6bc6e0f4d94310f7d72f49dc2c82b72187dadc5362e280f62618

Observation 2935ea0f-764a-4a2a-8341-5d42cd473aad · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.811240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.811240Z digest=sha256:254f048cc71a53d2644fdce7acdfa48a987d1d54b0d49d12ae08a4938324f11b

Observation 188f85ae-edff-4a3f-aa09-91fc9df16cc3 · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webvoyager: Building an end-to-end web agent with large multimodal models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.916387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.916387Z digest=sha256:82ac6608975aa5d2fcf49dcde2cce13b4b80e84e8efe4a24b5ed8609568da2e3

Observation bf78fd0c-0c32-48c7-86ad-fc0033c497ef · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.996496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.996496Z digest=sha256:e536e6c2b39e45407ebb3aa9dd944c945f5cd4764046103f19b871b64aa19da9

Observation 592bab4f-14f5-43de-8234-c160ebd7d284 · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.075464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.075464Z digest=sha256:4b37b307395731a32f2f7a9614474c2f29f5b7479e79ba6b49df46e65be9e1cc

Observation f05cad0a-6ce5-48e5-82e9-5d21a17a8a02 · outbound

This paper cites The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.158997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.158997Z digest=sha256:c616d2e84ee4d313d9b1c759e09623d9275d2123e53fe4c269cc76fce09bf64d

Observation 02c2e747-2463-4d6e-b47b-e9c0e5f07032 · outbound

This paper cites Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.214251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.214251Z digest=sha256:90c553cf4a0f4d19b49dccd856a3bd7a3b74f3e0883d186ecb674a4558a77828

Observation 4cc7197c-1adf-4548-8a30-f846e95381f2 · outbound

This paper cites Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.287770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.287770Z digest=sha256:d6706de698ed1590532cc5d428072db7129cd264422f191624d1e481ead34aca

Observation 81771bf5-edef-4fed-94ad-808de9210a98 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.372036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.372036Z digest=sha256:4acd1eededa0f795276e7eaad2628f67010343623cb4bf102026d9c0338c16e0

Observation b6e25965-c2c8-4b22-b36c-349ab6a533c4 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.455621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.455621Z digest=sha256:8a0ca6fc6a227ff09d9e5ad1143849f1d5f22cd962b2d0c55e054239b47b65b6

Observation b47a9d7b-74e2-4fe3-8fbd-963edfbf6074 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gaia: a benchmark for general ai assistants

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.536812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.536812Z digest=sha256:83326bad9d11e856c3aea77bc2e4ce5ab23e5bf5ccde8ff507d1fdb4c1257d30

Observation b512849c-9191-4ddc-a38b-e8badd0f1281 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.625544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.625544Z digest=sha256:e29186bcde0526b2e8b8684ac14d37b5a9792ddb752e7b40226c2503e242f20b

Observation 6f0dc2de-6651-456f-a1aa-1c902b2c9a9f · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.705418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.705418Z digest=sha256:1b5a373d785693edcd13b6229cc2e7447a9492e3ccd4946be3b77dd2c8f8843d

Observation 9f015896-0a35-463c-a30a-23394db146bc · outbound

This paper cites Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.779498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.779498Z digest=sha256:537abc59a1ab861597be1198c87060b15fefdbce282a545fc33f46b099fcd974

Observation ba0c0772-e8a4-4962-8322-b8fd75fcc637 · outbound

This paper cites Agent workflow memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent workflow memory

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.857487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.857487Z digest=sha256:9395b20351549950a5e411ee7a75115de4a41c36d583528954f8b72cb99872ad

Observation 62aa1b8c-60ad-45aa-9b9b-2d5cc21f14e3 · outbound

This paper cites none identified.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? none identified

Reference 92

Resolution
malformed identifier
no resolver link, observed 2026-08-04T01:12:38.939920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.939920Z digest=sha256:be791c18e76dd6fe236de67c26c3d759100809a2261b5b0111309ce3ad93b35a

Observation 3377049d-f4bc-4c76-8fb7-557fa03b9c9e · outbound

This paper cites reasoning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? reasoning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.018297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.018297Z digest=sha256:877a92b5bb523dd2b3982fea9c617d1f22b6f1a0de02663b7e950c5f1a9fed1e

Observation e2a5d1a1-0e80-419f-9eef-e66ae93408dc · outbound

This paper cites Order from most to least important.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Order from most to least important

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.092541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.092541Z digest=sha256:842c322a1d60d79475ccd96557eccb50162d624bea5d01a0e90929955dc91005

Observation f10004a5-26a9-4d39-95d1-277af5e60038 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.153063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.153063Z digest=sha256:3b8b1ed4a5fb079b49ccd6d4b8c1884508ed37231f637c06ae58dc5cb7023eca

Observation 90dc9e6e-7d6c-4c87-8624-ff0085958f7c · outbound

This paper cites Status:.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Status:

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.230523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.230523Z digest=sha256:4dcb855cd47d4226605fb0cfe61f32cc85bcd8d8aed0e844cf5947b70c0afe50

Observation 53aba8b1-c42c-45e0-9c21-a9a2afb989b4 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.310212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.310212Z digest=sha256:3e2ed4d1cdbcf56b0d90420396496388d1c3ace241f1b44b05259e4dae2336d7

Observation 09cd7c8b-48af-4bba-ade3-477b10cbe395 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.419533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.419533Z digest=sha256:5b130ba83c9cc3ab7ed4e202d6c5faa306a2d01c14fdfb0074675a01aaa39d71

Observation 00a09f38-aa8c-413d-a4b4-396c781e0e26 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.476352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.476352Z digest=sha256:51c430d462b76c97a225112f3453a448f9f52756ed651390d06491ad02475d42

Observation 6984840c-21ed-47f3-a1db-0713b084276d · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.559055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.559055Z digest=sha256:5f8c8eb62fc367b3f9364ff63e3b7dad4f46046f60f876358bc43542748943d1

Pith citing papers

No inbound Pith citation observations are available.