Pith. sign in

REVIEW 1 minor 1 cited by

Engagement Process models actions and observations as decoupled event streams over time rather than paired at fixed steps.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 22:38 UTC pith:WIZMBBQI

load-bearing objection EP decouples actions and observations into separate time streams to handle latency and persistent effects in POMDPs, but the claim that full decision-theoretic structure carries over rests on an unshown argument.

arxiv 2605.11484 v2 pith:WIZMBBQI submitted 2026-05-12 cs.AI

Engagement Process: Rethinking the Temporal Interface of Action and Observation

classification cs.AI
keywords engagement processtemporal interfacePOMDPaction observation decouplingdeliberation latencypersistent actionsmulti-rate coordinationdelayed feedback
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proposes Engagement Process as a formalism for agent-environment interaction that makes time explicit by treating actions and observations as separate streams unfolding along a timeline. Standard approaches pair them at synchronized decision steps, which hides issues like deliberation latency, delayed feedback, and actions that continue without new inputs. By decoupling the streams while retaining POMDP decision structure, the interface allows agents to organize behavior around explicit time, coordinate subsystems at different rates, and compose interactions. Toy, LLM-agent, and learning experiments show this reveals temporal patterns invisible to step-based models and lets policies adjust when time carries costs.

Core claim

Engagement Process (EP) is an interaction formalism that inherits the decision-theoretic structure of POMDPs while representing actions and observations as decoupled event streams along time, rather than updates paired at fixed decision steps. This captures single-agent timing issues such as deliberation latency, delayed feedback, and persistent actions, while supporting richer agent-side organization, multi-rate coordination, and compositional interaction among subsystems. Across experiments, EP exposes temporal behaviors hidden by step-based interfaces and enables policies to adapt under explicit time costs.

What carries the argument

Engagement Process (EP), the interaction formalism that decouples actions and observations into independent event streams along an explicit time axis.

Load-bearing premise

Decoupling actions and observations into separate time streams preserves the decision-theoretic properties of POMDPs without introducing inconsistencies that prevent policy optimization.

What would settle it

A concrete case where an EP policy produces value functions or optimal behavior that cannot be recovered from the equivalent synchronized POMDP without loss of correctness or optimality.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Single-agent timing issues such as deliberation latency, delayed feedback, and persistent actions become directly representable.
  • Agent subsystems can coordinate at independent rates instead of sharing a common step clock.
  • Policies adapt explicitly to time costs rather than treating time as an implicit side effect of step count.
  • Compositional interaction among agent modules becomes possible through shared but unsynchronized time streams.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same decoupling may simplify modeling of physical robots where actuator persistence and sensor delays are inherent.
  • LLM agents using tool calls with variable latency could organize internal reasoning around the same explicit time streams.
  • Policy search methods might need new update rules that propagate values across asynchronous event streams.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 1 minor

Summary. The paper proposes the Engagement Process (EP), an interaction formalism that represents actions and observations as decoupled event streams along time rather than paired updates at fixed decision steps. It claims that EP inherits the decision-theoretic structure of POMDPs, thereby capturing timing phenomena such as deliberation latency, delayed feedback, and persistent actions while enabling richer agent organization, multi-rate coordination, and compositional subsystem interaction. The claims are asserted to be supported by experiments across toy environments, LLM-agent settings, and learning tasks that reveal temporal behaviors obscured by conventional step-based interfaces.

Significance. If the inheritance of POMDP structure without inconsistencies or loss of policy-optimization properties can be formally established and the experiments provide verifiable evidence of new capabilities, the work would offer a meaningful conceptual advance for modeling asynchronous and variable-time-scale interactions in AI agents. The explicit treatment of time costs and event streams addresses a practical gap in current interfaces.

minor comments (1)
  1. [Abstract] The abstract states that experiments across toy, LLM-agent, and learning settings support the claims but supplies no quantitative results, methods, or verification details, preventing assessment of empirical strength.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their summary of the manuscript and for recognizing the potential significance of explicitly modeling time in the action-observation interface. We note that the report lists no specific major comments following the 'MAJOR COMMENTS:' heading, so we provide no point-by-point responses below. We are happy to address any additional questions the referee may have.

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper proposes Engagement Process (EP) as a formalism that inherits POMDP decision-theoretic structure while decoupling actions and observations into event streams. No equations, derivations, fitted parameters, or self-citations appear in the abstract or summary that would reduce any claimed result to an input by construction. The inheritance claim is presented as a modeling choice rather than a derived theorem that loops back on itself. This is a standard non-finding for a conceptual interface proposal whose central contribution is definitional rather than predictive.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only review; no information available on free parameters, axioms, or invented entities.

pith-pipeline@v0.9.1-grok · 5682 in / 1009 out tokens · 34942 ms · 2026-06-30T22:38:17.349328+00:00 · methodology

0 comments
read the original abstract

Task completion in digital and physical environments increasingly involves complex temporal interaction, where actions and observations unfold over different time scales rather than align with fixed observation--action steps. To model such interactions, we propose \emph{Engagement Process} (EP), an interaction formalism that inherits the decision-theoretic structure of POMDPs while making time explicit in the action--observation interface. EP represents actions and observations as decoupled event streams along time, rather than updates paired at fixed decision steps. This interface captures single-agent timing issues such as deliberation latency, delayed feedback, and persistent actions, while supporting richer agent-side organization, multi-rate coordination, and compositional interaction among subsystems. Across toy, LLM-agent, and learning experiments, EP exposes temporal behaviors hidden by step-based interfaces and enables policies to adapt under explicit time costs.

Figures

Figures reproduced from arXiv: 2605.11484 by Jiahao Zhang, Jialian Li, Jiaming Song, Jie Chen, Junhong Liu, Weiran Guo, Xutao Wang, Yuchen Cao.

Figure 1
Figure 1. Figure 1: Comparison of interaction interfaces. POMDPs pair observations and actions at fixed decision [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: LLM-based experiments. Tasks can be interpreted as a triage and scheduling problem over a [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Urgency-conditioned deliberation-mode distributions in the single-task setting. The EP-trained [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Urgency-conditioned deliberation-mode distributions in the sequential-task setting. EP learns a [PITH_FULL_IMAGE:figures/full_fig_p019_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: EP interrupting an in-progress checkpoint handling. At tick [PITH_FULL_IMAGE:figures/full_fig_p021_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Loop unable to interrupt an in-progress checkpoint handling. At tick [PITH_FULL_IMAGE:figures/full_fig_p022_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Representative episode from the resume_pressure family with three dishes, three tutor problems, and one stove slot. The upper lanes show generated tutor segments for Q1–Q3, while the lower lanes show cooking signals, valid response windows, and finish actions. 34 [PITH_FULL_IMAGE:figures/full_fig_p034_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction

    cs.AI 2026-07 conditional novelty 6.0

    An 8B LLM post-trained with SFT, RL, embodied-expert training, and model merging reaches high in-domain embodied-task success with very short responses.

Reference graph

Works this paper leans on

46 extracted references · 46 canonical work pages · cited by 1 Pith paper · 9 internal anchors

  1. [1]

    GPT-4 Technical Report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  2. [2]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025

  3. [3]

    Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances.arXiv preprint arXiv:2204.01691, 2022

  4. [4]

    ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

    Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, et al. pi0.7: a steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026

  5. [5]

    Openclaw: Personal ai assistant

    OpenClaw. Openclaw: Personal ai assistant. https://github.com/openclaw/openclaw, 2026. Open-source agent framework

  6. [6]

    Claude code: Anthropic’s agentic coding system

    Anthropic. Claude code: Anthropic’s agentic coding system. https://www.anthropic.com/ product/claude-code, 2025. Product page

  7. [7]

    Principles of metareasoning.Artificial intelligence, 49(1-3):361–395, 1991

    Stuart Russell and Eric Wefald. Principles of metareasoning.Artificial intelligence, 49(1-3):361–395, 1991

  8. [8]

    Using anytime algorithms in intelligent systems.AI magazine, 17(3):73–73, 1996

    Shlomo Zilberstein. Using anytime algorithms in intelligent systems.AI magazine, 17(3):73–73, 1996

  9. [9]

    Metareasoning: Theoretical and methodological developments, 2025

    Linden J Ball and Beth H Richardson. Metareasoning: Theoretical and methodological developments, 2025

  10. [10]

    Sutton, Doina Precup, and Satinder Singh

    Richard S. Sutton, Doina Precup, and Satinder Singh. Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning.Artificial Intelligence, 112(1–2):181–211, 1999. doi: 10.1016/S0004-3702(99)00052-1

  11. [11]

    Hierarchical reinforcement learning: A survey and open research challenges.Machine Learning and Knowledge Extraction, 4(1):172–221, 2022

    Matthias Hutsebaut-Buysse, Kevin Mets, and Steven Latré. Hierarchical reinforcement learning: A survey and open research challenges.Machine Learning and Knowledge Extraction, 4(1):172–221, 2022

  12. [12]

    Reinforcement learning methods for continuous-time markov decision problems.Advances in neural information processing systems, 7, 1994

    Steven Bradtke and Michael Duff. Reinforcement learning methods for continuous-time markov decision problems.Advances in neural information processing systems, 7, 1994

  13. [13]

    An introduction to event- triggered and self-triggered control

    Wilhelmus PMH Heemels, Karl Henrik Johansson, and Paulo Tabuada. An introduction to event- triggered and self-triggered control. In2012 ieee 51st ieee conference on decision and control (cdc), pages 3270–3285. IEEE, 2012

  14. [14]

    An overview of recent advances in event-triggered control.Science China Information Sciences, 68(6): 161201, 2025

    Xian-Ming Zhang, Qing-Long Han, Xiaohua Ge, Derui Ding, Boda Ning, and Bao-Lin Zhang. An overview of recent advances in event-triggered control.Science China Information Sciences, 68(6): 161201, 2025. 11

  15. [15]

    Revisiting active perception.Autonomous Robots, 42(2):177–196, 2018

    Ruzena Bajcsy, Yiannis Aloimonos, and John K Tsotsos. Revisiting active perception.Autonomous Robots, 42(2):177–196, 2018

  16. [16]

    A survey on active simultaneous localization and mapping: State of the art and new frontiers.IEEE Transactions on Robotics, 39(3):1686–1705, 2023

    Julio A Placed, Jared Strader, Henry Carrillo, Nikolay Atanasov, Vadim Indelman, Luca Carlone, and José A Castellanos. A survey on active simultaneous localization and mapping: State of the art and new frontiers.IEEE Transactions on Robotics, 39(3):1686–1705, 2023

  17. [17]

    Handling delay in real-time reinforcement learning

    Ivan Anokin, Rishav Rishav, Matthew Riemer, Stephen Chung, Irina Rish, and Samira Ebrahimi Kahou. Handling delay in real-time reinforcement learning. InInternational Conference on Learning Representations, 2025

  18. [18]

    Asynchronous tool usage for real-time agents.arXiv preprint arXiv:2410.21620, 2024

    Antonio A Ginart, Naveen Kodali, Jason Lee, Caiming Xiong, Silvio Savarese, and John Emmons. Asynchronous tool usage for real-time agents.arXiv preprint arXiv:2410.21620, 2024

  19. [19]

    Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming

    Martin L. Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley, New York, 1994. ISBN 9780471619772

  20. [20]

    Littman and Anthony R

    Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. Planning and acting in partially observable stochastic domains.Artificial Intelligence, 101(1–2):99–134, 1998. doi: 10.1016/S0004-3702(98)00023-X

  21. [21]

    Howard.Dynamic Probabilistic Systems, Volume II: Semi-Markov and Decision Processes

    Ronald A. Howard.Dynamic Probabilistic Systems, Volume II: Semi-Markov and Decision Processes. Wiley, New York, 1971

  22. [22]

    Hierarchical reinforcement learning with the maxq value function decomposi- tion.Journal of artificial intelligence research, 13:227–303, 2000

    Thomas G Dietterich. Hierarchical reinforcement learning with the maxq value function decomposi- tion.Journal of artificial intelligence research, 13:227–303, 2000

  23. [23]

    Bradtke and Michael O

    Steven J. Bradtke and Michael O. Duff. Reinforcement learning methods for continuous- time markov decision problems. In G. Tesauro, D. Touretzky, and T. Leen, editors, Advances in Neural Information Processing Systems, volume 7, pages 393–400. MIT Press, 1994. URL https://proceedings.neurips.cc/paper_files/paper/1994/file/ 07871915a8107172b3b5dc15a6574ad3...

  24. [24]

    POMDPs in continuous time and dis- crete spaces

    Bastian Alt, Matthias Schultheis, and Heinz Koeppl. POMDPs in continuous time and dis- crete spaces. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 13151–13162. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/ file/992...

  25. [25]

    Event-triggered real-time scheduling of stabilizing control tasks.IEEE Transactions on Automatic Control, 52(9):1680–1685, 2007

    Paulo Tabuada. Event-triggered real-time scheduling of stabilizing control tasks.IEEE Transactions on Automatic Control, 52(9):1680–1685, 2007. doi: 10.1109/TAC.2007.904277

  26. [26]

    Sanfelice, and Andrew R

    Rafal Goebel, Ricardo G. Sanfelice, and Andrew R. Teel.Hybrid Dynamical Systems: Modeling, Stability, and Robustness. Princeton University Press, Princeton, 2012. doi: 10.23943/princeton/ 9780691153896.001.0001

  27. [27]

    ReAct: Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InThe Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=WE_ vluYUL-X

  28. [28]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Processing Systems, volume 36, pages 6...

  29. [29]

    Reflexion: Language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. In A. Oh, T. Nau- mann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neu- ral Information Processing Systems, volume 36, pages 8634–8652. Curran Associates, Inc., 2023. URL https://proce...

  30. [30]

    A full-duplex speech dialogue scheme based on large language model

    Peng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan, Wei Xia, and Yuanjun Xiong. A full-duplex speech dialogue scheme based on large language model. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Processing Systems, volume 37, pages 13372–13403. Curran Asso- ciates, Inc., 2024. URL h...

  31. [31]

    Language model can listen while speaking

    Ziyang Ma, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen, Yuping Wang, Yuxuan Wang, and Xie Chen. Language model can listen while speaking. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 24831–24839, 2025

  32. [32]

    Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

    Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, and Hung yi Lee. Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities.arXiv preprint arXiv:2503.04721, 2025

  33. [33]

    P., Juhasz, A., Pohl, A., et al

    Gengyuan Zhang, Tanveer Hannan, Hermine Kleiner, Beste Aydemir, Xinyu Xie, Jian Lan, Thomas Seidl, V olker Tresp, and Jindong Gu. A ViLA: Asynchronous vision-language agent for streaming multimodal data interaction.arXiv preprint arXiv:2506.18472, 2025. doi: 10.48550/arXiv.2506. 18472

  34. [34]

    Robotouille: An asynchronous planning benchmark for LLM agents.arXiv preprint arXiv:2502.05227, 2025

    Gonzalo Gonzalez-Pumariega, Leong Su Yean, Neha Sunkara, and Sanjiban Choudhury. Robotouille: An asynchronous planning benchmark for LLM agents.arXiv preprint arXiv:2502.05227, 2025. ReAct (GPT-4o): 47% sync, 11% async

  35. [35]

    From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models

    Junlong Tong, Zilong Wang, YuJie Ren, Peiran Yin, Hao Wu, Wei Zhang, and Xiaoyu Shen. From static inference to dynamic interaction: A survey of streaming large language models.arXiv preprint arXiv:2603.04592, 2026. Taxonomy: output-streaming, sequential-streaming, concurrent-streaming

  36. [36]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Assoc...

  37. [37]

    Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters.arXiv preprint arXiv:2408.03314, 2024

  38. [38]

    Le, Christopher Ré, and Azalia Mirhoseini

    Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V . Le, Christopher Ré, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling,

  39. [39]

    URLhttps://arxiv.org/abs/2407.21787

  40. [40]

    Qwen3 Technical Report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

  41. [41]

    Deepmath-103k

    Zhangwei He. Deepmath-103k. https://huggingface.co/datasets/zwhe99/ DeepMath-103K, 2025. Hugging Face dataset. 13

  42. [42]

    Real-time reasoning agents in evolving environments.arXiv preprint arXiv:2511.04898, 2025

    Yule Wen, Yixin Ye, Yanzhe Zhang, Diyi Yang, and Hao Zhu. Real-time reasoning agents in evolving environments.arXiv preprint arXiv:2511.04898, 2025. Introduces Real-Time Reasoning Gym and AgileThinker

  43. [43]

    arXiv preprint arXiv:2506.07223 , year=

    Yangqing Zheng, Shunqi Mao, Dingxin Zhang, and Weidong Cai. LLM-enhanced rapid-reflex async-reflect embodied agent for real-time decision-making in dynamically changing environments. arXiv preprint arXiv:2506.07223, 2025. Proposes TCM and RRARA; evaluated on HAZARD benchmark

  44. [44]

    From reactive to active sensing: A survey on information gathering in decision-theoretic planning.ACM Computing Surveys, 55(13s):280:1–280:22, 2023

    Tiago Veiga and Jennifer Renoux. From reactive to active sensing: A survey on information gathering in decision-theoretic planning.ACM Computing Surveys, 55(13s):280:1–280:22, 2023. doi: 10.1145/3583068

  45. [45]

    OpenThoughts: Data Recipes for Reasoning Models

    Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, et al. Openthoughts: Data recipes for reasoning models.arXiv preprint arXiv:2506.04178, 2025

  46. [46]

    active option

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework. InProceedings of the Twentieth European Conference on Computer Systems, pages 1279–1297, 2025. 14 A Extended Related Work Streaming, full-duplex, and asynchronous agent systems.Recent agent...