Pith. sign in

REVIEW 2 major objections 1 cited by

DeltaBox achieves 14 ms checkpoint and 5 ms rollback for AI agent sandboxes by recording only changes between similar states.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 16:12 UTC pith:YLLJHBWD

load-bearing objection DeltaBox gives a concrete OS path to ms-level C/R for agent sandboxes by layering filesystems and incremental process dumps, but the numbers rest on unquantified similarity between checkpoints. the 2 major comments →

arxiv 2605.22781 v2 pith:YLLJHBWD submitted 2026-05-21 cs.OS cs.AI

DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback

classification cs.OS cs.AI
keywords AI agentscheckpoint rollbacksandboxOS abstractionchange-based C/RLLM agentsstate explorationDeltaState
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

AI agents performing tree search or reinforcement learning must repeatedly save and restore complete sandbox states that include files and process memory. Full duplication each time produces latencies of hundreds of milliseconds that limit how many states can be explored inside a fixed time budget. The paper observes that successive checkpoints are usually very similar and therefore proposes to copy only the differences. It realizes this idea with two new OS mechanisms that together deliver the reported millisecond latencies on SWE-bench and RL workloads, allowing agents to visit substantially more nodes without changing the underlying agent logic.

Core claim

DeltaBox supplies an OS-level DeltaState abstraction for transactional change-based checkpoint and rollback. DeltaFS stores file states in layers, freezes the current writable layer and inserts a fresh one at each checkpoint so that all further writes become copy-on-write, while rollback reduces to a layer switch. DeltaCR records only incremental changes to process state and accelerates rollback by forking directly from a frozen template process instead of replaying conventional pipelines. These two mechanisms together produce the measured 14 ms checkpoint and 5 ms rollback times.

What carries the argument

DeltaState, an OS abstraction for change-based transactional checkpoint/rollback implemented via layered filesystems (DeltaFS) and incremental process dumps with template forking (DeltaCR).

Load-bearing premise

Consecutive checkpoints in AI agent executions share enough similarity that recording only the differences remains faster and sufficient compared with full-state copies.

What would settle it

Run the same agent workloads on a scenario where each step alters a large fraction of sandbox state and measure whether DeltaBox checkpoint or rollback times exceed those of full-copy baselines.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Agents can explore substantially more nodes inside any fixed time budget on tasks such as SWE-bench.
  • High-frequency state exploration becomes practical for deeper tree search and larger fan-out in LLM agents.
  • Reinforcement learning loops that rely on sandbox resets avoid the previous C/R bottleneck.
  • Sandboxed agent platforms gain a concrete mechanism for frequent, low-cost state versioning without full duplication.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same change-tracking approach could apply to other domains that need rapid state versioning, such as interactive debugging or scientific simulations.
  • If real agent executions produce greater state divergence than the evaluated workloads, the latency advantage would shrink.
  • Native OS support for delta-style checkpoint primitives might become a standard facility once demonstrated in agent settings.
  • Integration with existing container runtimes would allow the technique to reach cloud-scale agent deployments.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper claims that AI agents require frequent full-state checkpoint/rollback (C/R) for exploration tasks like tree search and RL, but existing mechanisms incur hundreds of ms to seconds of latency due to full duplication. Observing that consecutive agent checkpoints are highly similar, it introduces the DeltaState abstraction and two mechanisms—DeltaFS (layered filesystem with dynamic freezing and CoW on new writable layers) and DeltaCR (incremental process dumps with direct fork from frozen templates)—to realize change-based transactional C/R. The resulting DeltaBox sandbox is evaluated on SWE-bench and RL micro-benchmarks, reporting 14 ms checkpoint and 5 ms rollback latencies that enable substantially more nodes to be explored under fixed time budgets.

Significance. If the performance numbers and the underlying similarity assumption hold under realistic agent workloads, the work would remove a major latency bottleneck for high-frequency state exploration in LLM agents, potentially enabling deeper search and larger fan-outs in test-time compute and RL settings.

major comments (2)
  1. [Abstract] Abstract: The central performance claim (14 ms checkpoint / 5 ms rollback) rests on the assertion that 'subsequent checkpoints in AI agents are highly similar,' yet the abstract supplies no quantitative evidence—delta sizes, change ratios, memory churn statistics, or frequency of full-state fallback—on SWE-bench or RL workloads. Without these data the ms-level result cannot be assessed for generality; low similarity would cause the mechanisms to revert to near-full-state costs.
  2. [Evaluation (implied)] The manuscript does not report how often DeltaFS or DeltaCR fall back to full duplication, nor the measured sizes of the incremental deltas under the evaluated workloads; these quantities are load-bearing for the claim that change-based C/R suffices for millisecond latency.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. The comments correctly identify that quantitative evidence of state similarity is needed to substantiate the millisecond-level claims. We will revise the manuscript to include these metrics.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The central performance claim (14 ms checkpoint / 5 ms rollback) rests on the assertion that 'subsequent checkpoints in AI agents are highly similar,' yet the abstract supplies no quantitative evidence—delta sizes, change ratios, memory churn statistics, or frequency of full-state fallback—on SWE-bench or RL workloads. Without these data the ms-level result cannot be assessed for generality; low similarity would cause the mechanisms to revert to near-full-state costs.

    Authors: We agree that the abstract would be strengthened by including supporting quantitative evidence. In the revised version we will add a concise statement summarizing the measured state similarity (average delta size relative to full state and observed change ratios) from the SWE-bench and RL workloads. This will allow readers to evaluate the generality of the approach directly from the abstract. revision: yes

  2. Referee: [Evaluation (implied)] The manuscript does not report how often DeltaFS or DeltaCR fall back to full duplication, nor the measured sizes of the incremental deltas under the evaluated workloads; these quantities are load-bearing for the claim that change-based C/R suffices for millisecond latency.

    Authors: We acknowledge the omission. Although the reported end-to-end latencies implicitly rely on high similarity, the manuscript does not explicitly tabulate fallback frequency to full duplication or per-workload delta sizes. We will add a dedicated subsection and table in the evaluation section that reports these quantities for both DeltaFS and DeltaCR across all benchmarks. If additional instrumentation is required, we will collect the data in the revision. revision: yes

Circularity Check

0 steps flagged

No circularity; claims rest on empirical measurements of new mechanisms

full rationale

The paper states an observation ('subsequent checkpoints in AI agents are highly similar') as the premise for DeltaFS and DeltaCR, then reports measured latencies (14 ms checkpoint, 5 ms rollback) from evaluations on SWE-bench and RL benchmarks. No equations, fitted parameters, self-citations, or uniqueness theorems are invoked to derive the performance numbers; they are presented as direct experimental outcomes. The central mechanisms are implemented OS abstractions whose correctness is validated externally rather than reduced to prior self-referential inputs.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 3 invented entities

The central performance claim rests on the domain assumption that consecutive agent states are highly similar and on the introduction of two new OS abstractions whose implementation details are not supplied in the abstract.

axioms (1)
  • domain assumption Subsequent checkpoints in AI agents are highly similar
    Stated as the key insight enabling change-based rather than full duplication.
invented entities (3)
  • DeltaState no independent evidence
    purpose: Change-based transactional checkpoint/rollback abstraction
    New OS-level abstraction proposed to realize the similarity insight.
  • DeltaFS no independent evidence
    purpose: Layered filesystem for change-based file C/R
    New mechanism that freezes writable layers and uses copy-on-write.
  • DeltaCR no independent evidence
    purpose: Incremental process-state C/R with direct fork from template
    New mechanism for process state that bypasses traditional pipelines.

pith-pipeline@v0.9.1-grok · 5843 in / 1252 out tokens · 40397 ms · 2026-06-30T16:12:48.287299+00:00 · methodology

0 comments
read the original abstract

LLM-powered AI agents require high-frequency state exploration (e.g., test-time tree search and reinforcement learning), relying on rapid checkpoint and rollback (C/R) of the complete sandbox state, including files and process state (e.g., memory, contexts, etc.). Existing mechanisms duplicate the entire state, causing hundreds of milliseconds to seconds of latency per C/R, which severely bottlenecks deep search and large-scale fan-outs. This paper observes that subsequent checkpoints in AI agents are highly similar. Therefore, instead of full duplication, a sandbox should only duplicate the changes between consecutive checkpoints (Key Insight). However, it is non-trivial to realize the idea, mainly due to the missing OS supports. This paper proposes a new OS-level abstraction, DeltaState, to enable the change-based transactional C/R for AI agents with two co-designed OS mechanisms. First, DeltaFS enables change-based filesystem C/R by organizing the file states into layers and dynamically freezing the writable layer and inserting a new one during checkpoint, reducing file updates to copy-on-write, and making rollback a simple layer switch. Second, DeltaCR enables change-based process state C/R using incremental dumps, and accelerates rollback by bypassing traditional pipelines to directly fork() from a frozen template process. We then present DeltaBox, a novel agent sandbox achieving millisecond level C/R through the two new mechanisms. Evaluations on SWE-bench and RL micro-benchmarks show DeltaBox completes checkpoint and rollback in millisecond-level latency (14ms and 5ms, respectively), empowering agents to explore substantially more nodes under fixed time budgets.

Figures

Figures reproduced from arXiv: 2605.22781 by Baochuan Yang, Dong Du, Haibo Chen, Jingkai He, Shiqi Liu, Si Yu, Yubin Xia, Yunpeng Dong, Yuze Hou, Zhonghu Xu.

Figure 1
Figure 1. Figure 1: Pass rate on SWE-bench Verified. (a) Linear ReAct vs. MCTS across three coding models. (b) Base vs. RL-trained across three open-weight model families. tree search and RL workloads. We propose the key insight of change-based DeltaState management. • We design DeltaFS, a runtime-reconfigurable overlayfs extension enabling unmount-free layer switching and lazy file descriptor redirection. • We design DeltaCR… view at source ↗
Figure 4
Figure 4. Figure 4: Fig.4.1 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 2
Figure 2. Figure 2: Design Overview. DeltaBox utilizes diff-based checkpoint/restore (i.e., deltaCheckpoint and deltaRestore) to enable millisecond-level checkpoint/rollback. An agent application can fully run inside a sandbox, or utilize a sandbox to execute a set of tools and use the C/R capabilities of DeltaBox to ensure prompt rollback when necessary. Layer 4: Search Strategy (Linear / BoN / MCTS) Layer 2: DeltaFS (overla… view at source ↗
Figure 3
Figure 3. Figure 3: The DeltaBox architecture. The StateManager coordinates DeltaFS (Layer 2, filesystem state) and DeltaCR (Layer 3, process state) to maintain consistent (filesystem, mem￾ory) state pairs at each search tree node. Base storage (Layer 1) provides the real storage functionalities. It would be better to adopt XFS (with reflink) to achieve block-level CoW and elimi￾nate write amplification. 1. The search strateg… view at source ↗
Figure 4
Figure 4. Figure 4: DeltaFS architecture. (a) Traditional overlayfs will prepare an upper layer file system which is writable for apps, and maintain a lower layer file system which is read-only and basically includes everything in the sandbox image. (b) DeltaFS extends the idea to support dynamic overlay, i.e., when an agent finishes a step of task and needs to make a checkpoint, instead of duplicating all files, DeltaFS inse… view at source ↗
Figure 5
Figure 5. Figure 5: DeltaCR architecture. Checkpointing creates both a CRIU image chain and a frozen template. Restores use the template fast path on hit, or the CRIU chain on miss; NPD keeps external I/O off the agent path. Dual-path checkpoint. At every checkpoint, DeltaCR simultaneously performs an asynchronous CRIU incre￾mental dump and a template-creating fork(). The CRIU dump provides a durable image for crash recovery … view at source ↗
Figure 4
Figure 4. Figure 4: Fig.4.1 [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Per-event blocking-time CDF, pooled across the 9 trajectory replays underlying Table. 2 (3 workloads × 3 reps): (a) checkpoint, (b) restore. DeltaBox’s distribution is shifted 1– 3 orders of magnitude left of every coupled backend; the gap holds into the tail (no event-class crossover at any percentile). flask sympy django astropy matplotlib 0 1 2 3 4 Time / LLM-only floor (×) 1.05× 1.06× 1.04× 1.03× 1.04×… view at source ↗
Figure 7
Figure 7. Figure 7: End-to-end time for a 100-iteration MCTS trajec￾tory (Qwen3-Coder-30B) on five SWE-bench Verified instances, normalized per-instance to the LLM-only floor (1.0× = pure LLM RTT sum). FC-Diff+dm and CHV+dm pair each VMM with dm-snapshot for filesystem coupling; FC-Diff+dm’s chain merge follows only the MCTS ancestor path. See §6.2.1 for the CubeSandbox approximation rationale. at the largest fan-out (p99=14.… view at source ↗
Figure 8
Figure 8. Figure 8: RL training fan-out characterisation, 5 substrates on the same GPU cluster. (a) Substrate 1:𝑁 fan-out latency on a single GPU (144 MB synthetic-anonymous parent template; 𝐾=10 prefork actions × 5 MB mid-state; 5 reps). The parent here is larger than DeltaBox’s realistic ∼15 MB agent.py measured in Table.3, so this panel stress-tests the substrate primitives rather than DeltaBox’s production parent RSS. (b)… view at source ↗
Figure 9
Figure 9. Figure 9: Per-event ckpt/restore latency vs. per-event [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Per-event CoW fault absorption vs. post-restore idle window (50 swe-search MCTS restore events). Shaded band: p1–p99 idle range from 2,311 LLM-driven restore events. Reflink-aware copy-up shares extents with the lower layer; only the 4 KB blocks the partial-write actually dirties con￾tribute to duplicated bytes, so the reflink curve sits below the no-reflink lines but tracks them in slope. The benefit gro… view at source ↗
Figure 13
Figure 13. Figure 13: End-of-trajectory CRIU dump storage on 9 SWE￾bench instances (3 per archetype, averaged), replayed through real criu dump on a 5 MB Python process matched to the Mode A footprint. Comparison: reachability-aware GC (§5.2.1) versus retaining every checkpoint. GC effectiveness. In a stress test with moderate branching factor and depth, the GC mechanism reclaims snapshot stor￾age at each prune event. GC runs … view at source ↗
Figure 11
Figure 11. Figure 11: Per-edit copy-up bytes (a) and physical I/O bytes (b) vs. edited-file size (log–log); real SWE-bench agent edits across three filesystem configurations, per-bin medians. Shaded band marks the typical agent edit range. ext4 and XFS-without￾reflink coincide on (a): copy-up benefit comes entirely from re￾flink, not XFS. 10 −1 10 0 10 1 10 2 Per-event checkpoint latency (ms) 0 100 200 ckpt events Standard (n=… view at source ↗
Figure 12
Figure 12. Figure 12: Per-event checkpoint latency on 1,689 ckpt events from 87 MCTS runs across 9 SWE-bench Verified repositories. Adaptive: pure-read cmds (LW, blue; 𝑛=1047) skip the dump; FS-mutating cmds (std, orange; 𝑛=642) take the full incremental dump. Standard (gray dashed): same events forced through the std path; 62.0% of events route to the LW peak. Lightweight skip ratio. Across 87 production MCTS runs spanning 9 … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities

    cs.CR 2026-07 accept novelty 6.5

    Execution-security research for AI coding agents is fragmented across 17 mechanism categories with five unaddressed cross-cutting gaps, including missing head-to-head isolation-vs-capability evaluation and untested re...

Reference graph

Works this paper leans on

63 extracted references · 63 canonical work pages · cited by 1 Pith paper · 6 internal anchors

  1. [1]

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In In- ternational Conference on Learning Representations , B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (Eds.), Vol. 2024. 54107–54157. https://proceedings.i...

  2. [2]

    Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A Re- alistic Web Environment for Building Autonomous Agents. InIn- ternational Conference on Learning Representations , B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. S...

  3. [3]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629 [cs.CL] https://arxiv. org/abs/2210.03629

  4. [4]

    E2B. 2024. E2B: The Enterprise AI Agent Cloud.https://e2b.dev

  5. [5]

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar

  6. [6]

    In In- ternational Conference on Learning Representations , Y

    Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning. In In- ternational Conference on Learning Representations , Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.), Vol. 2025. 10131– 10165. https://proceedings.iclr.cc/paper_files/paper/2025/file/ 1b623663fd9b874366f3ce019fdfdd44-Paper-Conference.pdf

  7. [7]

    Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2024. Language agent tree search unifies rea- soning, acting, and planning in language models. InProceedings of the 41st International Conference on Machine Learning (Vienna, Aus- tria)(ICML’24). JMLR.org, Article 2572, 23 pages

  8. [8]

    Cheng Zhang, Erhu Feng, Xi Zhao, Yisheng Zhao, Wangbo Gong, Jiahui Sun, Dong Du, Zhichao Hua, Yubin Xia, and Haibo Chen

  9. [9]

    arXiv:2509.00531 [cs.MA] https://arxiv.org/abs/2509.00531

    MobiAgent: A Systematic Framework for Customizable Mobile Agents. arXiv:2509.00531 [cs.MA] https://arxiv.org/abs/2509.00531

  10. [10]

    The OpenClaw Project. 2026. openclaw/openclaw: Your own personal AI assistant. Any OS. Any Platform. The lobster way.https://github. com/openclaw/openclaw

  11. [11]

    OpenAI. 2024. OpenAI o1 System Card. arXiv: 2412.16720 [cs.AI] https://arxiv.org/abs/2412.16720

  12. [12]

    DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capabil- ity in LLMs via Reinforcement Learning. arXiv: 2501.12948 [cs.AI] https://arxiv.org/abs/2501.12948

  13. [13]

    Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

    Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini. 2024. Large Lan- guage Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv:2407.21787 [cs.LG] https://arxiv.org/abs/2407.21787

  14. [14]

    Jingkai He, Tianjian Li, Erhu Feng, Dong Du, Qian Liu, Tao Liu, Yu- bin Xia, and Haibo Chen. 2026. History Doesn’t Repeat Itself but Roll- outs Rhyme: Accelerating Reinforcement Learning with RhymeRL. In Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (USA) (ASPLOS ’26...

  15. [15]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathemat- ical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] https://arxiv.org/abs/2402.03300

  16. [16]

    Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, juncai liu, LingJun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Ru Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Yuxuan Song, Xiangpeng Wei, Hao...

  17. [17]

    Lixiang Ao, George Porter, and Geoffrey M. Voelker. 2022. FaaSnap: FaaS made fast using snapshot-based VMs. InProceedings of the Sev- enteenth European Conference on Computer Systems (Rennes, France) (EuroSys ’22). Association for Computing Machinery, New York, NY, USA, 730–746. https://doi.org/10.1145/3492321.3524270

  18. [18]

    Dong Du, Tianyi Yu, Yubin Xia, Binyu Zang, Guanglu Yan, Cheng- gang Qin, Qixuan Wu, and Haibo Chen. 2020. Catalyzer: Sub- millisecond Startup for Serverless Computing with Initialization-less Booting. InProceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (Lausanne, Switzerland)(...

  19. [19]

    Dmitrii Ustiugov, Plamen Petrov, Marios Kogias, Edouard Bugnion, and Boris Grot. 2021. Benchmarking, analysis, and optimization of serverless function snapshots. InProceedings of the 26th ACM Inter- national Conference on Architectural Support for Programming Lan- guages and Operating Systems (Virtual, USA)(ASPLOS ’21). Associa- tion for Computing Machine...

  20. [20]

    Xiaohu Chai, Tianyu Zhou, Keyang Hu, Jianfeng Tan, Tiwei Bie, Anqi Shen, Dawei Shen, Qi Xing, Shun Song, Tongkai Yang, Le Gao, Feng Yu, Zhengyu He, Dong Du, Yubin Xia, Kang Chen, and Yu Chen. 2025. Fork in the road: reflections and optimizations for cold start latency in production serverless systems. InProceedings of the 19th USENIX Conference on Operati...

  21. [21]

    Lazar Cvetković, François Costa, Mihajlo Djokic, Michal Friedman, and Ana Klimovic. 2024. Dirigent: Lightweight Serverless Orches- tration. InProceedings of the ACM SIGOPS 30th Symposium on Op- erating Systems Principles (Austin, TX, USA)(SOSP ’24). Association for Computing Machinery, New York, NY, USA, 369–384. https: //doi.org/10.1145/3694715.3695966

  22. [22]

    Dong Du, Qingyuan Liu, Xueqiang Jiang, Yubin Xia, Binyu Zang, and Haibo Chen. 2022. Serverless computing on heterogeneous comput- ers. In Proceedings of the 27th ACM International Conference on Archi- tectural Support for Programming Languages and Operating Systems (Lausanne, Switzerland)(ASPLOS ’22) . Association for Computing Machinery, New York, NY, US...

  23. [23]

    Zijun Li, Jiagan Cheng, Quan Chen, Eryu Guan, Zizheng Bian, Yi Tao, Bin Zha, Qiang Wang, Weidong Han, and Minyi Guo. 2022. RunD: A Lightweight Secure Container Runtime for High-density Deploy- ment and High-concurrency Startup in Serverless Computing. In2022 USENIX Annual Technical Conference (USENIX ATC 22) . USENIX As- sociation, Carlsbad, CA, 53–68.htt...

  24. [24]

    Hanfei Yu, Rohan Basu Roy, Christian Fontenot, Devesh Tiwari, Jian Li, Hong Zhang, Hao Wang, and Seung-Jong Park. 2024. Rainbow- Cake: Mitigating Cold-starts in Serverless with Layer-wise Container Caching and Sharing. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1...

  25. [25]

    Jialiang Huang, MingXing Zhang, Teng Ma, Zheng Liu, Sixing Lin, Kang Chen, Jinlei Jiang, Xia Liao, Yingdi Shan, Ning Zhang, Mengt- ing Lu, Tao Ma, Haifeng Gong, and YongWei Wu. 2024. TrEnv: Trans- parently Share Serverless Execution Environments Across Different Functions and Nodes. InProceedings of the ACM SIGOPS 30th Sym- posium on Operating Systems Pri...

  26. [26]

    Frans Kaashoek

    Ariel Szekely, Adam Belay, Robert Morris, and M. Frans Kaashoek

  27. [27]

    In Proceedings of the ACM SIGOPS 30th Symposium on Operating Sys- tems Principles (Austin, TX, USA) (SOSP ’24)

    Unifying serverless and microservice workloads with SigmaOS. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Sys- tems Principles (Austin, TX, USA) (SOSP ’24) . Association for Com- puting Machinery, New York, NY, USA, 385–402. https://doi.org/10. 1145/3694715.3695947

  28. [28]

    E2B. 2026. E2B Sandbox persistence. https://e2b.dev/docs/sandbox/ persistence

  29. [29]

    The CRIU Project. 2011. CRIU: Checkpoint/Restore In Userspace. https://criu.org

  30. [30]

    Alexandru Agache, Marc Brooker, Andreea Florescu, Alexandra Ior- dache, Anthony Liguori, Rolf Neugebauer, Phil Piwonka, and Diana- Maria Popa. 2020. Firecracker: lightweight virtualization for server- less applications. InProceedings of the 17th Usenix Conference on Net- worked Systems Design and Implementation (Santa Clara, CA, USA) (NSDI’20). USENIX Ass...

  31. [31]

    DeepSeek-AI. 2026. DeepSeek-V4 Technical Report . Technical Re- port. DeepSeek-AI. https://huggingface.co/deepseek-ai/DeepSeek- V4-Pro/blob/main/DeepSeek_V4.pdf

  32. [32]

    Tencent Cloud. 2026. CubeSandbox. https://github.com/ TencentCloud/CubeSandbox

  33. [33]

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Bench- marking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. InAdvances in Neural Infor...

  34. [34]

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. 2024. AgentBench: Evaluating LLMs as Agents. InIn- ternational Conference on Learning Repre...

  35. [35]

    John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. SWE-agent: Agent- Computer Interfaces Enable Automated Software Engineering. In Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Assoc...

  36. [36]

    Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Daniel Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig. 2025. OpenHands: An Open Platform for AI So...

  37. [37]

    Paul Gauthier. 2023. Aider: AI Pair Programming in Your Terminal. https://aider.chat/

  38. [38]

    Edward Oakes, Leon Yang, Dennis Zhou, Kevin Houck, Tyler Harter, Andrea Arpaci-Dusseau, and Remzi Arpaci-Dusseau. 2018. SOCK: 18 Rapid Task Provisioning with Serverless-Optimized Containers. In 2018 USENIX Annual Technical Conference (ATC) . USENIX Associa- tion, 57–70. https://www.usenix.org/conference/atc18/presentation/ oakes

  39. [39]

    James Cadden, Thomas Unger, Yara Awad, Han Dong, Orran Krieger, and Jonathan Appavoo. 2020. SEUSS: skip redundant paths to make serverless fast. In Proceedings of the Fifteenth European Conference on Computer Systems (Heraklion, Greece)(EuroSys ’20). Association for Computing Machinery, New York, NY, USA, Article 32, 15 pages. https://doi.org/10.1145/3342...

  40. [40]

    Nikita Lazarev, Varun Gohil, James Tsai, Andy Anderson, Bhushan Chitlur, Zhiru Zhang, and Christina Delimitrou. 2024. Sabre: hardware-accelerated snapshot compression for serverless Mi- croVMs. In Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation (Santa Clara, CA, USA)(OSDI’24). USENIX Association, USA, Article 1, 18 pages

  41. [41]

    LangChain, Inc. 2024. LangGraph: Building Stateful, Multi-Actor Ap- plications with LLMs.https://github.com/langchain-ai/langgraph

  42. [42]

    Agent Lightning: Train ANY AI Agents with Reinforcement Learning

    Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K. Qiu, and Yuqing Yang. 2025. Agent Lightning: Train ANY AI Agents with Reinforcement Learning. arXiv:2508.03680 [cs.AI] https://arxiv.org/abs/2508.03680

  43. [43]

    Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, and William Wang. 2025. SWE-Search: Enhancing Soft- ware Agents with Monte Carlo Tree Search and Iterative Re- finement. In International Conference on Learning Representations , Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.), Vol. 2025. 64485–64515. https://proceedings.iclr.cc/p...

  44. [44]

    LangChain, Inc. 2022. LangChain: Building Applications with LLMs through Composability.https://github.com/langchain-ai/langchain

  45. [45]

    Gonzalez, Hao Zhang, and Ion Sto- ica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Sto- ica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles (SOSP ’23)

  46. [46]

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. Hybrid- Flow: A Flexible and Efficient RLHF Framework. InProceedings of the Twentieth European Conference on Computer Systems (EuroSys ’25)

  47. [47]

    Jian Hu, Xibin Wu, Zilin Zhu, Xianyu, Weixun Wang, Dehao Zhang, and Yu Cao. 2024. OpenRLHF: An Easy-to-use, Scalable and High- performance RLHF Framework. arXiv: 2405.11143 [cs.AI] https:// arxiv.org/abs/2405.11143

  48. [48]

    Daytona. 2024. Daytona.https://daytona.io

  49. [49]

    ZeroBoot. 2026. ZeroBoot: Sub-millisecond VM Sandboxes for AI Agents via Copy-on-Write Forking.https://github.com/zerobootdev/ zeroboot

  50. [50]

    Boyang Yan. 2025. Fault-Tolerant Sandboxing for AI Coding Agents: A Transactional Approach to Safe Autonomous Execution. arXiv:2512.12806 [cs.AI] https://arxiv.org/abs/2512.12806

  51. [51]

    Jialiang Huang, Teng Ma, Zheng Liu, Sixing Lin, Kang Chen, Jinlei Jiang, Xia Liao, Yingdi Shan, Yongwei Wu, Ning Zhang, Mengting Lu, Tao Ma, Haifeng Gong, and Mingxing Zhang. 2026. TrEnv-X: Trans- parently Share Serverless Execution Environments Across Different Functions and Nodes.ACM Transactions on Computer Systems(March 2026). https://doi.org/10.1145/3805475

  52. [52]

    Ben Holmes, Baltasar Dinis, Lana Honcharuk, Joshua Fried, and Adam Belay. 2025. Taming Serverless Cold Starts Through OS Co- Design. arXiv: 2509.14292 [cs.OS] https://arxiv.org/abs/2509.14292

  53. [53]

    Yanning Yang, Dong Du, Haitao Song, and Yubin Xia. 2024. On- demand and Parallel Checkpoint/Restore for GPU Applications. In Proceedings of the 2024 ACM Symposium on Cloud Computing (Red- mond, W A, USA) (SoCC ’24) . Association for Computing Machin- ery, New York, NY, USA, 415–433. https://doi.org/10.1145/3698038. 3698510

  54. [54]

    Tullmann, J

    P. Tullmann, J. Lepreau, B. Ford, and M. Hibler. 1996. User-level check- pointing through exportable kernel state. InProceedings of the Fifth In- ternational Workshop on Object-Orientation in Operation Systems . 85–

  55. [55]

    https://doi.org/10.1109/IWOOOS.1996.557874

  56. [56]

    Dirk Vogt, Armando Miraglia, Georgios Portokalidis, Herbert Bos, Andy Tanenbaum, and Cristiano Giuffrida. 2015. Speculative Mem- ory Checkpointing. In Proceedings of the 16th Annual Middleware Conference (Vancouver, BC, Canada) (Middleware ’15) . Association for Computing Machinery, New York, NY, USA, 197–209. https: //doi.org/10.1145/2814576.2814802

  57. [57]

    Dearle and D

    A. Dearle and D. Hulse. 1995. On page-based optimistic process check- pointing. InProceedings of International Workshop on Object Orien- tation in Operating Systems . 24–32. https://doi.org/10.1109/IWOOS. 1995.470583

  58. [58]

    Emil Tsalapatis, Ryan Hancock, Tavian Barnes, and Ali José Mashti- zadeh. 2021. The Aurora Single Level Store Operating System. InPro- ceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (Virtual Event, Germany)(SOSP ’21). Association for Com- puting Machinery, New York, NY, USA, 788–803. https://doi.org/10. 1145/3477132.3483563

  59. [59]

    Plank, Micah Beck, Gerry Kingsley, and Kai Li

    James S. Plank, Micah Beck, Gerry Kingsley, and Kai Li. 1995. Libckpt: transparent checkpointing under Unix. In Proceedings of the USENIX 1995 Technical Conference Proceedings (New Orleans, Louisiana)(TCON’95). USENIX Association, USA, 18

  60. [60]

    Jason Ansel, Kapil Arya, and Gene Cooperman. 2009. DMTCP: Trans- parent checkpointing for cluster computations and the desktop. In Proceedings of the 2009 IEEE International Symposium on Parallel and Distributed Processing (IPDPS ’09). IEEE Computer Society, USA, 1–12. https://doi.org/10.1109/IPDPS.2009.5161063

  61. [61]

    The Btrfs Project. 2009. Btrfs Documentation. https://btrfs. readthedocs.io

  62. [62]

    OpenZFS. 2013. OpenZFS Documentation.https://openzfs.github.io/ openzfs-docs/

  63. [63]

    Linux Kernel Project. 2018. EROFS: Enhanced Read-Only File System. https://erofs.docs.kernel.org. 19