Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:19:32.406859Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 100 inbound Pith citation observations for arXiv:2404.07972.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:19:32.406859Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:06:57.471803Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
68 of 68 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
Observation e87cf1fd-a8e3-4ed9-88a3-6f9f149cb93f · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ACT-1: Transformer for Actions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24b03bd5-ccdd-4702-a2b7-7efb120eb987 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Introducing the next generation of claude
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c6ab962-88a4-44e4-be8a-5adeeefcee2a · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments The claude 3 model family: Opus, sonnet, haiku
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efc94861-acd1-47a8-ac82-95857e7e65c7 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3d84067-01f8-40ea-a830-988413575bc7 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Qwen Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0b874fe-07af-41dd-96cf-1ece28c77511 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-1: Robotics Transformer for Real-World Control at Scale
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37a39ef4-f617-4a6d-abd2-6c000a0eaf42 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 845927e9-9ff1-4b9b-b3fa-45c2748c6dd9 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47fc6f52-90de-4566-8eef-4db63cee2c0e · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mind2Web: Towards a Generalist Agent for the Web
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a5208d7-d8f2-404f-adeb-0f7d94e205c8 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b30dd0c-3f10-46f4-a9bc-d16263bd24b1 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26bbebfe-935e-4e7b-a189-b2f04b35bbe8 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Multimodal Web Navigation with Instruction-Finetuned Foundation Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3821a82-c34f-4022-a157-b50f5a7fa60e · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4eac973-4ba2-4188-a9fb-89b51cebdd30 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a4cc4f2-a11f-439c-ac7a-26463c39fa5f · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 640615f0-63c6-4dab-8214-f12642c0eb00 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 639bd103-164a-48a4-bd0d-82cbf0daa853 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments CogAgent: A Visual Language Model for GUI Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83df3bdd-e0a0-4eec-afb2-4a140ab0c40c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments A data-driven approach for learning to control computers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b1d778c-6d22-4137-8a89-1f50763bdee9 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mixtral of Experts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23455256-bf66-4abe-8213-a40fefe12952 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e815a620-7508-4d43-9732-deac89b7ccc0 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92394b64-b147-4b23-aa4e-831a1940f237 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed91c05c-c1a2-4a65-8ef0-70d490ed9293 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Pix2struct: Screenshot parsing as pretraining for visual language understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12a5dd74-81c7-41d0-b5c7-26816fe11cc2 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f99749b-d19f-48cd-8994-b69744a70545 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments SheetCopilot: Bringing Software Productivity to the Next Level through Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 233cdd86-cc0c-4292-a444-a54f32b7ff39 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Silkie: Preference Distillation for Large Visual Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e22006b-03c2-44ec-83c9-8832f1ef955c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mapping Natural Language Instructions to Mobile UI Action Sequences
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4d586ca-ac0c-479e-9abf-d93c0ead8359 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Natasha Chrissane Lobo, Himani H
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b082a34-97f9-43d2-918d-30cdbf24c15c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24a9e56a-825c-4ac6-87da-d6ccf2b2e8e0 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a485cf8-67e6-4daf-879e-b0f154d78f8c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Visual Instruction Tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67eee199-5d51-4931-bb48-93df08cc5274 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AgentBench: Evaluating LLMs as Agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d395331-a1c3-4a85-95d2-fe597026794d · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Weblinx: Real-world website navigation with multi-turn dialogue
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f76ff87a-3d72-4686-8967-5ffdc37f8e4c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed603c44-1608-46e4-8b74-6f233d56c14e · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Introducing meta Llama 3: The most capable openly available LLM to date, April
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42d7b803-cec1-40c0-92f4-f1c27e04127b · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Accessed: 2024-04-18
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9417b89-2f7a-4391-aa5e-be5c25c69c4e · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GAIA: a benchmark for General AI Assistants
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7e8feae-bdba-4547-ab11-e79a45a4179d · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WebGPT: Browser-assisted question-answering with human feedback
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6562f1dd-626a-405a-9b2c-07441364ccea · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ScreenAgent: A Vision Language Model-driven Computer Control Agent
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cb6cda8-7862-40f7-bddc-353dacf7354b · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GPT-4 Technical Report
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62f91ddf-75f1-4bd0-9b39-51f2e90c0322 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Android in the Wild: A Large-Scale Dataset for Android Device Control
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6a85c12-7bab-4b87-b114-da4a0822f6f6 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c1ec10a-7474-423f-9a02-7857a3055474 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments An empirical study & evaluation of modern {CAPTCHAs}
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b85d019-4022-467c-a17c-d13338099204 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1f25d28-b1e9-4974-ae04-ce95df3823d7 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments World of bits: An open-domain platform for web-based agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43ba48cc-ed59-4645-90a3-6f6c49e8bd51 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Design2code: How far are we from automating front-end engineering?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae375c20-1860-4380-bebc-47473217e6e9 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Hierarchical Prompting Assists Large Language Model on Web Navigation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72d4933d-b12f-4ed8-9cf4-9e3dbfffbb68 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dbbd82f-cbc1-4881-838f-499ed3fed62c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Cradle: Empowering Foundation Agents Towards General Computer Control
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4800fce6-aa93-4591-bbd2-83312df9c5f3 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Gemini: A Family of Highly Capable Multimodal Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27243d42-c013-4d38-9b78-8e8a3ca53f0f · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AndroidEnv: A Reinforcement Learning Platform for Android
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c39cb77a-d7c0-4023-9b23-4ec4fb37272b · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments UGIF: UI Grounded Instruction Following
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d4aba06-231e-4667-bfdf-ddc9ca440267 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation feede790-49c5-4d60-8560-df638d65c377 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AutoDroid: LLM-powered Task Automation in Android
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e60b7d9-e453-49f4-a766-70025ed83707 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1cbfd6e8-54ac-4c04-bb7a-bed3622f7ecf · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e14ca917-fbb4-4c6b-b54e-b0a838ad929c · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e88b8a0-fefd-425c-8f95-4aee49b6af53 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 764786b0-f6f2-4864-bfdf-3f1703ee23d2 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4213af39-1902-4d97-9c4b-8f2ab1274964 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments UFO: A UI-Focused Agent for Windows OS Interaction
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64c90358-97b4-4c97-8592-884c8405262b · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Appagent: Multimodal agents as smartphone users
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f07dabaf-84a0-4a60-a6ef-c1db6abd67e7 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a00b3e49-e9e3-4285-8762-0d6e6b0c7d32 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Large language models are semi-parametric reinforcement learning agents
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19d6f11c-6204-4b84-9921-77b95068b8ab · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments You only look at screens: Multimodal chain-of-action agents
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f9b0c30-bc9e-4c26-9274-16972c236550 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Tie: Topological information enhanced structural reading comprehension on web pages
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e91e920a-ed79-432e-b80b-fa8bd9a595bb · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GPT-4V(ision) is a Generalist Web Agent, if Grounded
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55e8ff2e-054e-4603-b32e-b7f878eff23e · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78c207b4-616f-4c9c-9242-06017dc48947 · outbound
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45ea4ba0-8ec6-4a0d-ab76-d0b674236c40 · inbound
Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c16d6729-7ab8-46b5-a2b8-dd7129eb580a · inbound
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bd9de76-5195-4c14-9699-a4e4fd35a489 · inbound
WebCanvas: Benchmarking Web Agents in Online Environments OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37237622-eb7b-46d3-9022-4158bb0f1338 · inbound
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e9468d6-8aba-438c-8163-6259e7782567 · inbound
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3665c22e-0717-461b-b0bb-7f9dec2a447c · inbound
A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dffd54f-cb0c-4994-b4ac-68bbdd7cb12c · inbound
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a935f98a-3274-46e4-9fc8-32e1a2850359 · inbound
SWE-smith: Scaling Data for Software Engineering Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4804db6-1420-4297-ba88-d5c0bbe94f67 · inbound
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54719fc-d061-4af5-b4c3-0d10d14eea7c · inbound
Self-Challenging Language Model Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba39a1fd-be63-4687-abeb-baa85f06085d · inbound
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4dbe174-a48d-4b6f-b2f3-a242b1104d8a · inbound
A Red Teaming Roadmap Towards System-Level Safety OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb747de6-b1f3-453e-8a2f-ccdbc3835f84 · inbound
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e082b267-d4f3-424e-adf6-f4a1d9132bfd · inbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc31b950-ecb2-4d67-aa18-e3b9c59d6dfd · inbound
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f556dbea-3f9c-4423-b125-323b58224c3e · inbound
Deep Research Agents: A Systematic Examination And Roadmap OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 128
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac92648-2a35-4996-8114-fa90e82c1a78 · inbound
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b19a5cd-06d0-4230-b831-f8dc0a7641ca · inbound
Multilingual Multimodal Software Developer for Code Generation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aaf830b-3ec9-4809-b1b8-55ab9db3030f · inbound
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a687a27-e2c6-419d-8403-c693e0e42247 · inbound
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86cee43f-b5a5-47e1-bf96-ee4e1c71595f · inbound
Magentic-UI: Towards Human-in-the-loop Agentic Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b202f99-2d90-468a-bb16-4eeb23db8f52 · inbound
Instruction Agent: Enhancing Agent with Expert Demonstration OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8cafdc6-f760-4ca0-9b38-f5bee3d790f5 · inbound
Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07be4d4d-cc99-4fee-9539-94cc41f6ed51 · inbound
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8ca01e-9cc6-4fbd-9d53-a162f0550e2c · inbound
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97c84ab-9c89-45ed-be05-c2d32f2cf919 · inbound
StepShield: When, Not Whether to Intervene on Rogue Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904a1599-9831-48bb-94b2-36ee6e0795b0 · inbound
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71f09531-1f2b-40d1-bb63-9db123bdb778 · inbound
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5540bf1-04d1-4ece-bd63-ead3acd460e1 · inbound
Kimi K2.5: Visual Agentic Intelligence OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc6ada6d-7735-42f7-a4ac-6ac66b0584ba · inbound
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a91f53af-6f4a-4505-a425-5dca60de01fa · inbound
SoK: Agentic Skills -- Beyond Tool Use in LLM Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abcc432b-10f2-4292-80d0-40da3a898ca6 · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c213217-51ad-4a18-8c99-a06869988851 · inbound
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 298b6312-b1f0-4571-95a8-d2461734e9fc · inbound
From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8633bda4-be3a-45df-84ff-f2516609ef69 · inbound
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44285b12-253b-4eb4-a757-3f1a25d360d0 · inbound
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40057c14-2a2d-40d2-ba8a-dc9ed847dd4d · inbound
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2398afb7-84a7-4154-aee6-5a595908d505 · inbound
Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 032d5d0a-5811-4825-a3f9-abe5baaeae9f · inbound
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc198b38-781f-46a9-af33-525685b05aeb · inbound
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecc082ca-81bb-41b2-aee2-eddc131da08c · inbound
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2c8c00-37d3-4151-941b-2e26111b9c2e · inbound
ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee85e413-d8c0-4d3b-8823-d78918b9f5f4 · inbound
AlphaEval: Evaluating Agents in Production OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49a3dcac-b105-4d31-93ad-642ccc14491b · inbound
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6e0dfde-0040-4693-b42e-d6b1c9497225 · inbound
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf8215c-e671-43dd-b4bf-76b79fdcf04b · inbound
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2fd0831-55c3-4b44-bcee-94bf73eab6a7 · inbound
Feedback-Driven Execution for LLM-Based Binary Analysis OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcd0509b-45db-4e80-9215-df0592d69242 · inbound
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a843195-225a-4412-b517-bfb9ca7bb020 · inbound
AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3087765e-665b-4789-80c0-4b10fa6ad09d · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c2be504-a7cd-435b-aa65-7e1cada043de · inbound
An AI Agent Execution Environment to Safeguard User Data OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c3fde8a-7c01-4147-8f95-ec5d9399ff71 · inbound
Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39a0e033-0990-49a6-a3da-4e8240976b5d · inbound
Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8618001-0809-4a68-9bfe-fae470de5136 · inbound
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2be37d84-b3eb-4d9f-8621-3abdc73dea44 · inbound
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ee0a034-8c83-48a4-ad06-a7a32eebe56d · inbound
Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 312c40f5-f571-4b74-84e4-f205ae9638df · inbound
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e090455c-764d-4059-bba2-527dd17e2491 · inbound
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7fe53c18-0841-40f0-9444-da361312a449 · inbound
From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cccaf9a3-0569-44c7-954e-1a197f8918cd · inbound
Computer Use at the Edge of the Statistical Precipice OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec5fb971-a6ec-4eb9-9ade-2d989e18b270 · inbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91415c43-8a4a-4889-a950-cf58d03451d3 · inbound
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 129bda9f-d70a-4730-97fc-d03263e9b001 · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 628d2910-eb0e-44f0-a840-a01d93968117 · inbound
ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb2eabd7-0d48-4ba9-8e92-d94058ae0a1b · inbound
ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17efd460-2324-4a42-8f49-25ed16f9ddc7 · inbound
ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9b79432-6d19-41d0-9b3b-99d0ea5a0813 · inbound
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 485adf21-7650-4e82-b235-0a94a4437746 · inbound
TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92dbda9f-7b41-44f9-9101-2dd726c6d696 · inbound
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4b78e58-b813-4555-bd31-05a4205fa2c4 · inbound
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f04903dd-771d-4f29-b4d4-5a70e275a4c1 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4444d1a-6104-4618-8ef6-77bc024d2374 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e875bef5-035f-4dc6-9d21-f8bc8f663f7a · inbound
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba4257a2-04d5-4132-b839-2a476e23b5a1 · inbound
Toward Native Multimodal Modeling: A Roadmap OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b9f29b5-77bb-4699-ae9e-568309c91f0d · inbound
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b538d906-a148-4471-a20c-adb3ef165141 · inbound
JobBench: Aligning Agent Work With Human Will OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7054e13-65f4-4ff1-9991-a6e79549a425 · inbound
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99d73b7a-c0c8-4b2f-9ff9-874eff3809cc · inbound
LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 431c9706-c82f-443a-a499-cd75dff39cc1 · inbound
unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d25a1708-869f-4877-94de-43d5f06fc7b7 · inbound
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 204f3772-d73c-4ebd-819c-6a34d06aa610 · inbound
Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bfbde66-5caf-4e99-b9d3-61db7a082ccc · inbound
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59e18fa3-dd0c-40fd-becf-acb2b4d70162 · inbound
DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 986ea443-d739-45f7-a4bd-2a2424c8ba3d · inbound
Signal-Driven Observation for Long-Horizon Web Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f471f672-402a-4c1f-9f79-c3c8f5afc7ce · inbound
Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 676eedc4-f966-46ac-8a5a-bf48f9198938 · inbound
Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9034f7da-dcdf-4ee1-8b55-ce51053694db · inbound
MedCTA: A Benchmark for Clinical Tool Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecdfb51e-749b-4c1c-8f12-830a4a1967d5 · inbound
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0e318c6-2350-4f65-9b7f-ef12c8db6307 · inbound
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c6f4af5-c226-4b7d-8677-afe15c476945 · inbound
ProCUA-SFT Technical Report OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20c23fbe-6d3a-4871-b31f-22cbee475275 · inbound
Dissecting model behavior through agent trajectories OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fba41a5-172a-477f-9198-c5c2054a91e0 · inbound
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b24081f8-6507-4cb3-9fc4-6f529091bc0f · inbound
OpenRath: Session-Centered Runtime State for Agent Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 634fb884-4b43-4c52-a1ac-aa63bb85ae65 · inbound
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a61e9752-141a-4f5f-9513-355acb2bed54 · inbound
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8cde36b-ff87-4f3a-9bd5-2d4efbadff3e · inbound
PhoneBuddy: Training Open Models for Agentic Phone Use OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed6c4d3c-cf0f-4a99-8330-cd5ab95518ac · inbound
Agent vs. Parametric World Models: Hybrid Planning for Reliable Language Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90aa53f0-1a45-4f9c-988b-3726014708b1 · inbound
Agent-Computer Observation Interfaces Enable Dynamic Computer Use OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc1ea22f-ce92-4b19-bd9c-78e4d25964e6 · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 958cd20e-c09a-4896-b797-a35e52443aa0 · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.