{"work":{"id":"0624be05-1d97-4fd6-8300-b04b8a3ab04b","openalex_id":"https://openalex.org/W4319451738","doi":"10.48550/arxiv.2302.01973","arxiv_id":"2601.11868","raw_key":null,"title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","authors":null,"authors_text":"Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich","year":2026,"venue":"cs.SE","abstract":"AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .","external_url":"https://arxiv.org/abs/2601.11868","cited_by_count":0,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2601.11868","created_at":"2026-05-09T04:30:11.467156+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","render_title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces"},"hub":{"state":{"work_id":"0624be05-1d97-4fd6-8300-b04b8a3ab04b","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":136,"external_cited_by_count":0,"distinct_field_count":12,"first_pith_cited_at":"2026-02-02T16:17:38+00:00","last_pith_cited_at":"2026-07-09T09:56:50+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T11:29:32.754028+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":16},{"context_role":"dataset","n":5},{"context_role":"baseline","n":2},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":16},{"context_polarity":"use_dataset","n":3},{"context_polarity":"baseline","n":2},{"context_polarity":"unclear","n":2},{"context_polarity":"support","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","claims":[{"claim_text":"AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\\% on the bench","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"evaluation pipelines have not internalized an adversarial mindset, and that proactive auditing could help close the security gap for the fast-paced benchmarking space. 1 Introduction The progress of AI is mostly tracked by a wide range of benchmarks. Hundreds of new benchmarks have been released in the past two years, spanning software engineering [25, 15, 13], web naviga- tion [61], desktop computing [55], general AI assistance [35], terminal operations [34], enterprise workflows [50], and tool","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[20] Yuan Liang, Ruobin Zhong, Haoming Xu, Chen Jiang, Yi Zhong, Runnan Fang, Jia-Chen Gu, Shumin Deng, Yunzhi Yao, Mengru Wang, et al. Skillnet: Create, evaluate, and connect ai skills. arXiv preprint arXiv:2603.04448, 2026. [21] Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. Agentbench: Evaluating llms as agents.arXiv preprint arXiv:2308.03688, 2023. [22] Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Har","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Jesse Hu, Christopher Michael Rytting, Ryan Marten, Yixin Wang, Alex Dimakis, Andy Konwinski, and Ludwig Schmidt. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026. arXiv:2601.11868 [cs.SE]. [35] Youssef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure, Mattia Rigotti, Kristjan Greenewald, Ji","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"runtime agent evaluation remains a far-from-resolved task for current frontier models. We release the task specifications, containerized workspaces, grading code, and harness configurations to support reproducible evaluation. 2. Related Work Agent Benchmarks across Environments. Agent benchmarks have largely been organized by interaction sur- face: software engineering (SWE-bench [ 17], Terminal-Bench [24], LiveCodeBench [ 16]), web and GUI con- trol (WebArena [59], WebShop [48], VisualWebArena ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"•Agentic Capabilities: BrowseComp [68], WideSearch [69],DeepSearchQA [60], FinSearchComp (T2&T3) [26], Seal-0 [45], GDPVal [43]. •Image Understanding:(math & reasoning)MMMU-Pro [75], MMMU (val) [76], CharXiv (RQ) [67], Math- Vision [61] and MathVista (mini) [36];(vision knowledge)SimpleVQA [13] and WorldVQA 2;(perception) ZeroBench (w/ and w/o tools) [48], BabyVision [12], BLINK [18] and MMVP [57];(OCR & document)OCR- Bench [35], OmniDocBench 1.5 [42] and InfoVQA [38]. •Video Understanding: Vide","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ramework/harbor. [81] Milton Rokeach.The nature of human values.Free press, 1973. [82] ShalomHSchwartz,JanCieciuch,MicheleVecchione,EldadDavidov,RonaldFischer,Constanze Beierlein,AliceRamos,MarkkuVerkasalo,Jan-ErikLönnqvist,KursadDemirutku,etal. Refining the theory of basic individual values.Journal of personality and social psychology, 103(4):663, 2012. [83] Shalom H Schwartz. Basic human values: Theory, methods, and application.Risorsa Uomo, (2007/2), 2007. [84] Philip E. Tetlock.Coping with T","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (16 contexts).","role_counts":[{"n":16,"context_role":"background"},{"n":3,"context_role":"dataset"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-07-02T22:23:01.528442+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"004b8966-de45-4333-b4bf-54becea9918b","orcid":null,"display_name":"Mike A. Merrill"},{"id":"503a251e-dbb7-4d75-b656-dd43b42e328b","orcid":null,"display_name":"Alexander G. Shaw"},{"id":"e598061b-b452-4a7e-8dc3-5e1d3353d206","orcid":null,"display_name":"Nicholas Carlini"},{"id":"1ae322de-b126-40d4-94de-b8b3d210d411","orcid":null,"display_name":"Boxuan Li"},{"id":"374d3e59-b3eb-4a4e-8cdc-1e0358d67938","orcid":null,"display_name":"Harsh Raj"},{"id":"8c78c9b3-b0fc-48fa-be65-3ab0032edaf8","orcid":null,"display_name":"Ivan Bercovich"}]},"error":null,"updated_at":"2026-07-02T22:23:02.082503+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T15:32:03.169098+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?","work_id":"d0effe15-a689-441a-8e3f-ea35f1c4e4b1","shared_citers":13},{"title":"$\\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains","work_id":"6a8d8dc4-0cc0-4052-8109-abbcdcd4a962","shared_citers":8},{"title":"DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models","work_id":"07c85cc5-4086-4abc-823b-6d0f4ff784d0","shared_citers":7},{"title":"Kimi K2.5: Visual Agentic Intelligence","work_id":"d690be8f-5d53-49b0-b1e7-79668eb8fcdb","shared_citers":7},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":7},{"title":"BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents","work_id":"25adb508-d97c-49d6-ae43-7a70c2478a34","shared_citers":6},{"title":"GLM-5: from Vibe Coding to Agentic Engineering","work_id":"ad29b1a2-bf77-46b3-9ead-fb62b1d2c6fe","shared_citers":6},{"title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","work_id":"ea9e51ce-1e75-4182-92d8-4d25f70d2ee4","shared_citers":6},{"title":"OpenHands: An Open Platform for AI Software Developers as Generalist Agents","work_id":"f1762ea0-e382-4f38-a28c-adc643789859","shared_citers":6},{"title":"SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?","work_id":"a561c78a-4b02-4053-a92a-bc5c7c5f6b9b","shared_citers":6},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":5},{"title":"WebArena: A Realistic Web Environment for Building Autonomous Agents","work_id":"7058ffd2-a339-4102-89eb-248eeb074652","shared_citers":5},{"title":"$\\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment","work_id":"3a498b1a-455f-4667-b572-c5216c99a89c","shared_citers":4},{"title":"Agent-SafetyBench: Evaluating the Safety of LLM Agents","work_id":"96afb8b9-0e7e-442c-93b1-6638599fc041","shared_citers":4},{"title":"AlphaEvolve: A coding agent for scientific and algorithmic discovery","work_id":"76a0f850-d490-4e4f-ab98-8d25df82cd23","shared_citers":4},{"title":"Humanity's Last Exam","work_id":"59ea00d4-16a8-45e1-aafc-290a6f91d9f4","shared_citers":4},{"title":"Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek","work_id":"6eca346e-0e48-4240-961c-1ecee1e71aab","shared_citers":4},{"title":"Meta-Harness: End-to-End Optimization of Model Harnesses","work_id":"5be3c079-4ffa-458f-adcf-204b66f7af51","shared_citers":4},{"title":"Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning","work_id":"0e0b7549-2bc4-4574-aa7f-588ffa16eaae","shared_citers":4},{"title":"AgentBench: Evaluating LLMs as Agents","work_id":"a37549b4-4c94-412d-acc4-4efeb08509be","shared_citers":3},{"title":"Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models","work_id":"c1d09cf9-9176-403d-8028-630bbc0cbc5d","shared_citers":3},{"title":"arXiv preprint arXiv:2601.03192 , year=","work_id":"0ab11b49-e934-49e1-9eaf-a25fb7343010","shared_citers":3},{"title":"arXiv preprint arXiv:2603.04448 , year=","work_id":"47ef6124-7b5c-4acd-94d6-cbd6ab854c9c","shared_citers":3},{"title":"AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation","work_id":"92b7eb9c-c3d8-4518-a376-06fa15dd895b","shared_citers":3}],"time_series":[{"n":43,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T15:32:03.188118+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T15:32:07.923213+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","claims":[{"claim_text":"AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\\% on the bench","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"evaluation pipelines have not internalized an adversarial mindset, and that proactive auditing could help close the security gap for the fast-paced benchmarking space. 1 Introduction The progress of AI is mostly tracked by a wide range of benchmarks. Hundreds of new benchmarks have been released in the past two years, spanning software engineering [25, 15, 13], web naviga- tion [61], desktop computing [55], general AI assistance [35], terminal operations [34], enterprise workflows [50], and tool","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[20] Yuan Liang, Ruobin Zhong, Haoming Xu, Chen Jiang, Yi Zhong, Runnan Fang, Jia-Chen Gu, Shumin Deng, Yunzhi Yao, Mengru Wang, et al. Skillnet: Create, evaluate, and connect ai skills. arXiv preprint arXiv:2603.04448, 2026. [21] Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. Agentbench: Evaluating llms as agents.arXiv preprint arXiv:2308.03688, 2023. [22] Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Har","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Jesse Hu, Christopher Michael Rytting, Ryan Marten, Yixin Wang, Alex Dimakis, Andy Konwinski, and Ludwig Schmidt. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026. arXiv:2601.11868 [cs.SE]. [35] Youssef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure, Mattia Rigotti, Kristjan Greenewald, Ji","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"runtime agent evaluation remains a far-from-resolved task for current frontier models. We release the task specifications, containerized workspaces, grading code, and harness configurations to support reproducible evaluation. 2. Related Work Agent Benchmarks across Environments. Agent benchmarks have largely been organized by interaction sur- face: software engineering (SWE-bench [ 17], Terminal-Bench [24], LiveCodeBench [ 16]), web and GUI con- trol (WebArena [59], WebShop [48], VisualWebArena ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"•Agentic Capabilities: BrowseComp [68], WideSearch [69],DeepSearchQA [60], FinSearchComp (T2&T3) [26], Seal-0 [45], GDPVal [43]. •Image Understanding:(math & reasoning)MMMU-Pro [75], MMMU (val) [76], CharXiv (RQ) [67], Math- Vision [61] and MathVista (mini) [36];(vision knowledge)SimpleVQA [13] and WorldVQA 2;(perception) ZeroBench (w/ and w/o tools) [48], BabyVision [12], BLINK [18] and MMVP [57];(OCR & document)OCR- Bench [35], OmniDocBench 1.5 [42] and InfoVQA [38]. •Video Understanding: Vide","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ramework/harbor. [81] Milton Rokeach.The nature of human values.Free press, 1973. [82] ShalomHSchwartz,JanCieciuch,MicheleVecchione,EldadDavidov,RonaldFischer,Constanze Beierlein,AliceRamos,MarkkuVerkasalo,Jan-ErikLönnqvist,KursadDemirutku,etal. Refining the theory of basic individual values.Journal of personality and social psychology, 103(4):663, 2012. [83] Shalom H Schwartz. Basic human values: Theory, methods, and application.Risorsa Uomo, (2007/2), 2007. [84] Philip E. Tetlock.Coping with T","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (16 contexts).","role_counts":[{"n":16,"context_role":"background"},{"n":3,"context_role":"dataset"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-07-02T22:23:02.084861+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","claims":[{"claim_text":"AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\\% on the bench","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T15:31:57.145191+00:00"}},"summary":{"title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","claims":[{"claim_text":"AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\\% on the bench","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?","work_id":"d0effe15-a689-441a-8e3f-ea35f1c4e4b1","shared_citers":13},{"title":"$\\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains","work_id":"6a8d8dc4-0cc0-4052-8109-abbcdcd4a962","shared_citers":8},{"title":"DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models","work_id":"07c85cc5-4086-4abc-823b-6d0f4ff784d0","shared_citers":7},{"title":"Kimi K2.5: Visual Agentic Intelligence","work_id":"d690be8f-5d53-49b0-b1e7-79668eb8fcdb","shared_citers":7},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":7},{"title":"BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents","work_id":"25adb508-d97c-49d6-ae43-7a70c2478a34","shared_citers":6},{"title":"GLM-5: from Vibe Coding to Agentic Engineering","work_id":"ad29b1a2-bf77-46b3-9ead-fb62b1d2c6fe","shared_citers":6},{"title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","work_id":"ea9e51ce-1e75-4182-92d8-4d25f70d2ee4","shared_citers":6},{"title":"OpenHands: An Open Platform for AI Software Developers as Generalist Agents","work_id":"f1762ea0-e382-4f38-a28c-adc643789859","shared_citers":6},{"title":"SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?","work_id":"a561c78a-4b02-4053-a92a-bc5c7c5f6b9b","shared_citers":6},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":5},{"title":"WebArena: A Realistic Web Environment for Building Autonomous Agents","work_id":"7058ffd2-a339-4102-89eb-248eeb074652","shared_citers":5},{"title":"$\\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment","work_id":"3a498b1a-455f-4667-b572-c5216c99a89c","shared_citers":4},{"title":"Agent-SafetyBench: Evaluating the Safety of LLM Agents","work_id":"96afb8b9-0e7e-442c-93b1-6638599fc041","shared_citers":4},{"title":"AlphaEvolve: A coding agent for scientific and algorithmic discovery","work_id":"76a0f850-d490-4e4f-ab98-8d25df82cd23","shared_citers":4},{"title":"Humanity's Last Exam","work_id":"59ea00d4-16a8-45e1-aafc-290a6f91d9f4","shared_citers":4},{"title":"Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek","work_id":"6eca346e-0e48-4240-961c-1ecee1e71aab","shared_citers":4},{"title":"Meta-Harness: End-to-End Optimization of Model Harnesses","work_id":"5be3c079-4ffa-458f-adcf-204b66f7af51","shared_citers":4},{"title":"Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning","work_id":"0e0b7549-2bc4-4574-aa7f-588ffa16eaae","shared_citers":4},{"title":"AgentBench: Evaluating LLMs as Agents","work_id":"a37549b4-4c94-412d-acc4-4efeb08509be","shared_citers":3},{"title":"Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models","work_id":"c1d09cf9-9176-403d-8028-630bbc0cbc5d","shared_citers":3},{"title":"arXiv preprint arXiv:2601.03192 , year=","work_id":"0ab11b49-e934-49e1-9eaf-a25fb7343010","shared_citers":3},{"title":"arXiv preprint arXiv:2603.04448 , year=","work_id":"47ef6124-7b5c-4acd-94d6-cbd6ab854c9c","shared_citers":3},{"title":"AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation","work_id":"92b7eb9c-c3d8-4518-a376-06fa15dd895b","shared_citers":3}],"time_series":[{"n":43,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"503a251e-dbb7-4d75-b656-dd43b42e328b","orcid":null,"display_name":"Alexander G. Shaw","source":"manual","import_confidence":0.72},{"id":"1ae322de-b126-40d4-94de-b8b3d210d411","orcid":null,"display_name":"Boxuan Li","source":"manual","import_confidence":0.72},{"id":"374d3e59-b3eb-4a4e-8cdc-1e0358d67938","orcid":null,"display_name":"Harsh Raj","source":"manual","import_confidence":0.72},{"id":"8c78c9b3-b0fc-48fa-be65-3ab0032edaf8","orcid":null,"display_name":"Ivan Bercovich","source":"manual","import_confidence":0.72},{"id":"004b8966-de45-4333-b4bf-54becea9918b","orcid":null,"display_name":"Mike A. Merrill","source":"manual","import_confidence":0.72},{"id":"e598061b-b452-4a7e-8dc3-5e1d3353d206","orcid":null,"display_name":"Nicholas Carlini","source":"manual","import_confidence":0.72}]}}