Pith. sign in

Paper Citation Record · LEDGER

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 100 inbound Pith citation observations for arXiv:2404.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.07972 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:19:32.406859Z

measured 168 of 168 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 131 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:06:57.471803Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact42
  • verified fuzzy15
  • unresolved1
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch8

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e87cf1fd-a8e3-4ed9-88a3-6f9f149cb93f · outbound

This paper cites ACT-1: Transformer for Actions.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ACT-1: Transformer for Actions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.622324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:056dcd40f183a283252dd2f8beac6d77f58a1978e98d9f85ed6ad3a635bfe0e1

Observation 24b03bd5-ccdd-4702-a2b7-7efb120eb987 · outbound

This paper cites Introducing the next generation of claude.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Introducing the next generation of claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.615361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:d0b7e94cf17f8bbf5879ce06d5872f82bc44ea40c932f7f6ed9a0cc86731eb7d

Observation 3c6ab962-88a4-44e4-be8a-5adeeefcee2a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.612990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:c0cf29ff1de7682dac1821871fed72dc105a651c28dae0c90a8cbb2dc5baae32

Observation efc94861-acd1-47a8-ac82-95857e7e65c7 · outbound

This paper cites ScreenAI: A Vision-Language Model for UI and Infographics Understanding.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.469196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:a6c018a1750ff7e97ffa14d487c1968deb59ca8e7149ca339f9696dffdcc1f5a

Observation d3d84067-01f8-40ea-a830-988413575bc7 · outbound

This paper cites Qwen Technical Report.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Qwen Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.442790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:add3049c9212da85a2b2caf8f0f4c47cd663a23bcd66c6bb360c9436a0602cb3

Observation f0b874fe-07af-41dd-96cf-1ece28c77511 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.446816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:de053879e2ac76729b471b67a60eb89e884079aff5a005b38256b302c144e4f5

Observation 37a39ef4-f617-4a6d-abd2-6c000a0eaf42 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.450810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:20164422ec12fb845ba94d1af3f7a2ef1a780aea9912f08503e96e84bb687d9b

Observation 845927e9-9ff1-4b9b-b3fa-45c2748c6dd9 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.778093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:db6c7f069c97901929aeaedcf09c1d65f86e1d27d7ed30aa354570805ca0df9a

Observation 47fc6f52-90de-4566-8eef-4db63cee2c0e · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mind2Web: Towards a Generalist Agent for the Web

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:05:16.335224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:9f15276dfefd3b3d68ce0cc9267252c7bbab01a4427a3c0d456578cae103cda1

Observation 5a5208d7-d8f2-404f-adeb-0f7d94e205c8 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:48:05.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:dab0e7d06b2412386679c7b16a20a081617f0ebbb2d20927cc7f66b71261d5a6

Observation 2b30dd0c-3f10-46f4-a9bc-d16263bd24b1 · outbound

This paper cites an unresolved cited work.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-13T01:19:32.607565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:412e1bf85e5ab250cdd22f6e6a7a1ab7fd120fad89c11262e1a184416070cecb

Observation 26bbebfe-935e-4e7b-a189-b2f04b35bbe8 · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.525521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:c6ae19e59592b240b0719c8f98fd7472d941d0e30b0e947603a65a5b70262cb8

Observation c3821a82-c34f-4022-a157-b50f5a7fa60e · outbound

This paper cites ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.528636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:2089ad89f328f5f615cb529ccc8583ad2b51f84ca305d159c8a7e3416a9c3a2e

Observation c4eac973-4ba2-4188-a9fb-89b51cebdd30 · outbound

This paper cites PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.532136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:6c098eaee84fcdde55c9f1da9c41b7ae883c19de49847cca1d9c5ce81d705dd2

Observation 3a4cc4f2-a11f-439c-ac7a-26463c39fa5f · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:37:34.727630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:28e6351ebb3661fcd13282e892f5216fa719b0ccbc943d046d81a83690facabb

Observation 640615f0-63c6-4dab-8214-f12642c0eb00 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:43:35.479641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:1dcaeceabc1c6292a7a3e3bd2de5381b0ce32541fc5af3c742a8f5021e393ea5

Observation 639bd103-164a-48a4-bd0d-82cbf0daa853 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments CogAgent: A Visual Language Model for GUI Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.465423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:0ef47e18644736b4efa36868751c814942d02c7dec0a6521c39e0b1f97520f69

Observation 83df3bdd-e0a0-4eec-afb2-4a140ab0c40c · outbound

This paper cites A data-driven approach for learning to control computers.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments A data-driven approach for learning to control computers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.624649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:f98fbd88f7f92c71a61fc223709473b1540dfa7dfda18485ee203f556f06d5b6

Observation 0b1d778c-6d22-4137-8a89-1f50763bdee9 · outbound

This paper cites Mixtral of Experts.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mixtral of Experts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.541679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:25a0a2ad3a56269590d5124e3022676429826e0926233f4d6d8377c3f6edbc7d

Observation 23455256-bf66-4abe-8213-a40fefe12952 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.545200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:3af5bb791d103ff0062c0167393275217c70e6306f7f0d0c426ecbffd6b9f709

Observation e815a620-7508-4d43-9732-deac89b7ccc0 · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.548729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:218874393625d68a50a913d822b76d35952f5f2709ef11e65ca933ae75c2becf

Observation 92394b64-b147-4b23-aa4e-831a1940f237 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:20:36.653799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:ee5fda9dad6c53acce070abbcaa38f19e30729a531b311d43ebb6c1fca885b75

Observation ed91c05c-c1a2-4a65-8ef0-70d490ed9293 · outbound

This paper cites Pix2struct: Screenshot parsing as pretraining for visual language understanding.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Pix2struct: Screenshot parsing as pretraining for visual language understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.635957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:f12dbdd94e4cc2b25bc8de24d388e8106fdfc967839c12811da3f18715c08c09

Observation 12a5dd74-81c7-41d0-b5c7-26816fe11cc2 · outbound

This paper cites Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.554838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:63951c861b312fc4b78f73ae174b08d8162d47d8f65d5f6d94623e3314d2214a

Observation 0f99749b-d19f-48cd-8994-b69744a70545 · outbound

This paper cites SheetCopilot: Bringing Software Productivity to the Next Level through Large Language Models.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments SheetCopilot: Bringing Software Productivity to the Next Level through Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.558195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:693c3bd2b851b3232b0f83bb1805dc9f921a8602f3ec71f06ffd73f5008e52d0

Observation 233cdd86-cc0c-4292-a444-a54f32b7ff39 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Silkie: Preference Distillation for Large Visual Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.561482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:e0954f38447f8f5272f66f40a4f8bd0de408d4c17481404c08651564ca518ac5

Observation 7e22006b-03c2-44ec-83c9-8832f1ef955c · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.564185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:5d6193e48a53531c45d2f3aef0e50f4455664c7a16fdc7e884769609bfd75d13

Observation b4d586ca-ac0c-479e-9abf-d93c0ead8359 · outbound

This paper cites Natasha Chrissane Lobo, Himani H.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Natasha Chrissane Lobo, Himani H

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.434145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:794c961a39ae842e694ed524498da21873ad3fbe7d6a7e99eb8a5ccbd9ea803b

Observation 5b082a34-97f9-43d2-918d-30cdbf24c15c · outbound

This paper cites NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.567035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:97c6182373117e528cb91bd3871c33d6dc8f046af268d425bd6841acf15947b8

Observation 24a9e56a-825c-4ac6-87da-d6ccf2b2e8e0 · outbound

This paper cites Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.570023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:fcf9197a63d74a6cd174d13146fd7a51a5bc966484fc9b2738afe350fe35208b

Observation 9a485cf8-67e6-4daf-879e-b0f154d78f8c · outbound

This paper cites Visual Instruction Tuning.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Visual Instruction Tuning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.572521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:beced4ad79b63de0f8214f6e11c37d9ca86f57d3f40b9535bd203df75a0b1077

Observation 67eee199-5d51-4931-bb48-93df08cc5274 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AgentBench: Evaluating LLMs as Agents

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.575725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:797af6e9867a827772d3cc17ccbdb12361105574ee5f8cb42e1dcdbfb05ccde9

Observation 0d395331-a1c3-4a85-95d2-fe597026794d · outbound

This paper cites Weblinx: Real-world website navigation with multi-turn dialogue.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Weblinx: Real-world website navigation with multi-turn dialogue

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.578740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:576a71fe7723ef1de43f8096156fdaeb4e1a96ed2262ac8ae01baf2be5c65851

Observation f76ff87a-3d72-4686-8967-5ffdc37f8e4c · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.581967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:712fc986db2549467ea105125f0277b3aeb8a74202f71ee999662b54f355051c

Observation ed603c44-1608-46e4-8b74-6f233d56c14e · outbound

This paper cites Introducing meta Llama 3: The most capable openly available LLM to date, April.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Introducing meta Llama 3: The most capable openly available LLM to date, April

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.633727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:2f123864dcfbd7c7561d1e4adfa9355211a58bdfd0c9d9c316bdbd4404451778

Observation 42d7b803-cec1-40c0-92f4-f1c27e04127b · outbound

This paper cites Accessed: 2024-04-18.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Accessed: 2024-04-18

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.637770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:27442025573b59d8d94b436389e942c618135608029269b05f02b634b3f25986

Observation b9417b89-2f7a-4391-aa5e-be5c25c69c4e · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GAIA: a benchmark for General AI Assistants

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.584711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:c623b884799d21de8fc85021e3696e861d0f803bd2e459ca76f165954fd3a919

Observation a7e8feae-bdba-4547-ab11-e79a45a4179d · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WebGPT: Browser-assisted question-answering with human feedback

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.587578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:999d09fd073d05595a5c09fc6338bffad7db3e297f2aa6e90c0ecf6fd5c06cfc

Observation 6562f1dd-626a-405a-9b2c-07441364ccea · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.591004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:e75a58473761da44d6811a0ab09ab40f0ec116c742144b0083b7ed0922b24df0

Observation 9cb6cda8-7862-40f7-bddc-353dacf7354b · outbound

This paper cites GPT-4 Technical Report.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GPT-4 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.593875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:d0a053b61d6288e005e1ae3f3eabd7dcda61b03644cc3906032ab8d23efc529e

Observation 62f91ddf-75f1-4bd0-9b39-51f2e90c0322 · outbound

This paper cites Android in the Wild: A Large-Scale Dataset for Android Device Control.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Android in the Wild: A Large-Scale Dataset for Android Device Control

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.597044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:8502fe4707eb9b384f5f96f58fd7b270fead7678766be07d2ecf44f3b89a9bf7

Observation a6a85c12-7bab-4b87-b114-da4a0822f6f6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.600721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:4eec623a55824c2c0c98ee25ac1a984056e112f708125c0c96fcdeb280caf4a8

Observation 8c1ec10a-7474-423f-9a02-7857a3055474 · outbound

This paper cites An empirical study & evaluation of modern {CAPTCHAs}.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments An empirical study & evaluation of modern {CAPTCHAs}

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.626988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:079589f826f67a05e334278f0523f3b397571d9dafca30f6dfe2b7433d0496bd

Observation 6b85d019-4022-467c-a17c-d13338099204 · outbound

This paper cites From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.604987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:a0404bc2dd34784021cab01550d6a44678974957c364234670d2da5cc00f5bec

Observation f1f25d28-b1e9-4974-ae04-ce95df3823d7 · outbound

This paper cites World of bits: An open-domain platform for web-based agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments World of bits: An open-domain platform for web-based agents

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.631174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:4cff15adf21bfc41ad302e761ba3bf0975b66e6ad9fe8b7bd91cd8e7827a0947

Observation 43ba48cc-ed59-4645-90a3-6f6c49e8bd51 · outbound

This paper cites Design2code: How far are we from automating front-end engineering?.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Design2code: How far are we from automating front-end engineering?

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.639703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:a6beea474b64ef156d8e38b2fcfad60920536883b9cb8da386f8d6ac27c86637

Observation ae375c20-1860-4380-bebc-47473217e6e9 · outbound

This paper cites Hierarchical Prompting Assists Large Language Model on Web Navigation.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Hierarchical Prompting Assists Large Language Model on Web Navigation

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.472256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:2716ec38803dc092f3fadbd9718b1355a58675775887d1798b3af3858fc60061

Observation 72d4933d-b12f-4ed8-9cf4-9e3dbfffbb68 · outbound

This paper cites META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.475665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:04d5708ac42416ee8431aa868237ea111cd485ca9d499857b525cb2160286a7c

Observation 3dbbd82f-cbc1-4881-838f-499ed3fed62c · outbound

This paper cites Cradle: Empowering Foundation Agents Towards General Computer Control.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Cradle: Empowering Foundation Agents Towards General Computer Control

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.478918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:1ca091635bd19c0766c56d3518e2bcb3e4bf81de254e145e5c23dff74b0075b9

Observation 4800fce6-aa93-4591-bbd2-83312df9c5f3 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Gemini: A Family of Highly Capable Multimodal Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.482620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:3cd3a2527b57d334f77b73654124a366a1e8dd2e24b5abc83029d5ccf8eca9d2

Observation 27243d42-c013-4d38-9b78-8e8a3ca53f0f · outbound

This paper cites AndroidEnv: A Reinforcement Learning Platform for Android.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AndroidEnv: A Reinforcement Learning Platform for Android

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.485919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:21260bb8e8e071a20c64dbd8c62635acce71b223252be479d26fc8ee9d9d6b06

Observation c39cb77a-d7c0-4023-9b23-4ec4fb37272b · outbound

This paper cites UGIF: UI Grounded Instruction Following.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments UGIF: UI Grounded Instruction Following

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.490386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:c3914087089b789cf6150ed83369189ee69bf2e93a1a2068a5e6891f21ef8c73

Observation 7d4aba06-231e-4667-bfdf-ddc9ca440267 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:19:28.048419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:7cebcc1d5ff610257b1d34f19a30379a4847a0925a08ff517c2a981dfe2af342

Observation feede790-49c5-4d60-8560-df638d65c377 · outbound

This paper cites AutoDroid: LLM-powered Task Automation in Android.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AutoDroid: LLM-powered Task Automation in Android

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.497484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:50b52a31b9707922157e65630887fd416aae8dc0a9b9672d7e4edeb38055f1ca

Observation 2e60b7d9-e453-49f4-a766-70025ed83707 · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 55

Resolution
malformed identifier
doi_truncated, observed 2026-05-13T01:19:32.437346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:cce6f310c716f4dcf44c016f3b0b059f25aeaca0632999b0aedcbd8256f48193

Observation 1cbfd6e8-54ac-4c04-bb7a-bed3622f7ecf · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.500326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:e45fdd9fc77ce916f7ca47c378555a6af43eb29d83b0112a780bdfdd6f341a2b

Observation e14ca917-fbb4-4c6b-b54e-b0a838ad929c · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.503391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:18f8f64862c0a842e9645cbf64a81fe8d6eedf0e2e3c73f9cd0d79ea9868d17b

Observation 4e88b8a0-fefd-425c-8f95-4aee49b6af53 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.507302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:7824b725f3c028a0e3d5ca0575ee59ad9740d8a01fdb529ed1199304a5730cf5

Observation 764786b0-f6f2-4864-bfdf-3f1703ee23d2 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.641863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:3ee665aaaf05b9aa481148e354ac8d09f977095a0626674a9013dadd0e20ccba

Observation 4213af39-1902-4d97-9c4b-8f2ab1274964 · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments UFO: A UI-Focused Agent for Windows OS Interaction

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.510117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:476ba277c6dd5290a10b8870cc73c33b63faa318479fa2adec1f9c3bc0aacb2b

Observation 64c90358-97b4-4c97-8592-884c8405262b · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Appagent: Multimodal agents as smartphone users

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.617660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:c2df449e360059d26b8709591f2d6673597c0e5fc219f4cedf6710c58e649b53

Observation f07dabaf-84a0-4a60-a6ef-c1db6abd67e7 · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.513206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:ea51d15da6ab246ad3a3883494b47463a1d56dfc503f8f45bacfba125249c1ca

Observation a00b3e49-e9e3-4285-8762-0d6e6b0c7d32 · outbound

This paper cites Large language models are semi-parametric reinforcement learning agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Large language models are semi-parametric reinforcement learning agents

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.629086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:c3a2832ca717b159e2bd66fb3fe93de45a1db887a073b75489f97d2c4aded9f6

Observation 19d6f11c-6204-4b84-9921-77b95068b8ab · outbound

This paper cites You only look at screens: Multimodal chain-of-action agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments You only look at screens: Multimodal chain-of-action agents

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.610216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:ffc4cf5f15e7092296cefef41b6c9b96356ae3b6e9a239fcda679ba79b62e17e

Observation 1f9b0c30-bc9e-4c26-9274-16972c236550 · outbound

This paper cites Tie: Topological information enhanced structural reading comprehension on web pages.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Tie: Topological information enhanced structural reading comprehension on web pages

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:19:32.620251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:beabe75d1fcdb5810a377ba12a324cd1e80e99d3f9c58167e9787a67e3128908

Observation e91e920a-ed79-432e-b80b-fa8bd9a595bb · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:40:14.169994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:b306d850bfe09f5eaaa6da15fd7fb79abee21b5ec498c2f03bfa1425cf6239a3

Observation 55e8ff2e-054e-4603-b32e-b7f878eff23e · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.519332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:8b502b3bacaa721435a5abcde18e0950b4e58ec03caa632a710596c6348590ff

Observation 78c207b4-616f-4c9c-9242-06017dc48947 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 68

Resolution
malformed identifier
local_arxiv, observed 2026-05-13T01:19:32.522499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:ed6d3efdc874398b0ecf01c1ce00102f7f8c159176ffbfc30c8fd0595825ebfb

Pith citing papers

Observation 45ea4ba0-8ec6-4a0d-ab76-d0b674236c40 · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:57:26.848568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:2a93e6731e77c73bffee5b8f4078a6183319faefc958dcb6ba668e6442be5392

Observation c16d6729-7ab8-46b5-a2b8-dd7129eb580a · inbound

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents cites this paper.

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:06:13.844846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T12:06:13.697487Z digest=sha256:20d632d411994170311c384727885584f81b1dc1f9f289b3f981a8cb6039b30c

Observation 5bd9de76-5195-4c14-9699-a4e4fd35a489 · inbound

WebCanvas: Benchmarking Web Agents in Online Environments cites this paper.

WebCanvas: Benchmarking Web Agents in Online Environments OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:29:14.970132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:29:14.918147Z digest=sha256:2d93fb8590f76ba9405d62863b57f7323badd017aeef846a8865069bef14a054

Observation 37237622-eb7b-46d3-9022-4158bb0f1338 · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:29:27.553165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:4e166961081067267938d5cf5ec40354f951130358d5523bc84970bfa71b9507

Observation 7e9468d6-8aba-438c-8163-6259e7782567 · inbound

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction cites this paper.

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 113

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T04:09:41.649412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T04:09:41.494136Z digest=sha256:8378b9eb37d23d421d016d6bd1a304dd6daa135565a114ec57480dabd3a524ab

Observation 3665c22e-0717-461b-b0bb-7f9dec2a447c · inbound

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions cites this paper.

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 173

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:02:35.561245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T04:59:36.994758Z digest=sha256:8a7222e564e57cdda81db084c105f9829e15d8f464be045fc0d8d58a468e7830

Observation 1dffd54f-cb0c-4994-b4ac-68bbdd7cb12c · inbound

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks cites this paper.

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T21:32:18.687785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:32:18.491541Z digest=sha256:32511d8e629a807ad0d85b770cf61281dba9b294c0129a1c177fb423817a73cf

Observation a935f98a-3274-46e4-9fc8-32e1a2850359 · inbound

SWE-smith: Scaling Data for Software Engineering Agents cites this paper.

SWE-smith: Scaling Data for Software Engineering Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T10:22:06.984822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T10:22:06.951531Z digest=sha256:10ea7fd346e67541848caa9589c2f803e08431f7dcd79c91996b166edbb3cb94

Observation e4804db6-1420-4297-ba88-d5c0bbe94f67 · inbound

Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking cites this paper.

Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:57.471803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:06:57.471803Z digest=sha256:9dbb489f72325a231063badfd069e95be3de87a353bb46dd5a22a0289f1991f3

Observation e54719fc-d061-4af5-b4c3-0d10d14eea7c · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:58.052425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:58.052425Z digest=sha256:0462d0e87d1bd846c04de31c1a5cb36e773a34632834d5eae1170c0371d0a7b6

Observation ba39a1fd-be63-4687-abeb-baa85f06085d · inbound

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback cites this paper.

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:56.287971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:56.287971Z digest=sha256:f0ad2db9c4e2b8f1b73b500497bcda9c77dde96ca015f33e3658bf7a9daea54c

Observation a4dbe174-a48d-4b6f-b2f3-a242b1104d8a · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:25.259960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:25.259960Z digest=sha256:c1b869d502426cb58b0b81bf2f708ed794345d0ff7629505a75e263f692fd1e5

Observation bb747de6-b1f3-453e-8a2f-ccdbc3835f84 · inbound

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities cites this paper.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.285064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.285064Z digest=sha256:a7675c6fed974d038cbdb4f7e13febf5000006772173ba390db3fcdafb09b3c9

Observation e082b267-d4f3-424e-adf6-f4a1d9132bfd · inbound

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards cites this paper.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.492855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.492855Z digest=sha256:61c0f1c97988e4c131c959ea73eacb370395e6d9c4b2cbd4ead58d83ea916ce3

Observation bc31b950-ecb2-4d67-aa18-e3b9c59d6dfd · inbound

Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents cites this paper.

Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:18.569925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:41:18.569925Z digest=sha256:d2a618f770438ee24f834d40698b9aac6519d6282898142d342b9f6a1d46b283

Observation f556dbea-3f9c-4423-b125-323b58224c3e · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.979221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.979221Z digest=sha256:505fd6f69786252a995a46663fb865663d849f337dfef56b73ada4d188173c1f

Observation eac92648-2a35-4996-8114-fa90e82c1a78 · inbound

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents cites this paper.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.585127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.585127Z digest=sha256:ce07524962c97e1aa4533b6626dc6d7b48643106d1703018b09be00a44a49371

Observation 6b19a5cd-06d0-4230-b831-f8dc0a7641ca · inbound

Multilingual Multimodal Software Developer for Code Generation cites this paper.

Multilingual Multimodal Software Developer for Code Generation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.348232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.348232Z digest=sha256:c7b1d3f03d183832faafe089d196ec007035af7799b3bff7c95c1d85227e359e

Observation 6aaf830b-3ec9-4809-b1b8-55ab9db3030f · inbound

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models cites this paper.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.738182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.738182Z digest=sha256:c11eb811ba1642ca47b92746d1cacd8d4cc411239c0b79a474f2c85391675ac8

Observation 7a687a27-e2c6-419d-8403-c693e0e42247 · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.352421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.352421Z digest=sha256:60d18649789b7430e607700d9a4617d94ff58b591121a0358e6dfa20d430cb15

Observation 86cee43f-b5a5-47e1-bf96-ee4e1c71595f · inbound

Magentic-UI: Towards Human-in-the-loop Agentic Systems cites this paper.

Magentic-UI: Towards Human-in-the-loop Agentic Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T11:52:12.346411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:52:12.346411Z digest=sha256:802cbb82d9ed964ddde086f8b82a6a925afa030676b0184c4720fad4c851275c

Observation 0b202f99-2d90-468a-bb16-4eeb23db8f52 · inbound

Instruction Agent: Enhancing Agent with Expert Demonstration cites this paper.

Instruction Agent: Enhancing Agent with Expert Demonstration OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:20.925734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:55:20.925734Z digest=sha256:6282ae98fd87f4495fc0d18dbe958dab23d4211b24c5c16142869075196b2596

Observation f8cafdc6-f760-4ca0-9b38-f5bee3d790f5 · inbound

Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes cites this paper.

Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T11:31:10.146970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:31:10.146970Z digest=sha256:bf2879eb25f45f9c8e89944ca82e006b17c1fcea36557040dab5294bcbc3f9c4

Observation 07be4d4d-cc99-4fee-9539-94cc41f6ed51 · inbound

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering cites this paper.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.101569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.101569Z digest=sha256:cabc5a3885d3a22b39b0153d9a1af432d6b4709b7d48414fab8b7fde653fa7db

Observation aa8ca01e-9cc6-4fbd-9d53-a162f0550e2c · inbound

InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training cites this paper.

InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:10:24.865140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:10:24.865140Z digest=sha256:7ece1335cd334ff1c9a7e185fe833001e7c6c3f60b04ac29b02c6b8dbed8f113

Observation d97c84ab-9c89-45ed-be05-c2d32f2cf919 · inbound

StepShield: When, Not Whether to Intervene on Rogue Agents cites this paper.

StepShield: When, Not Whether to Intervene on Rogue Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.428948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.428948Z digest=sha256:1f4ac78f6b44f008fafc2b724d55dcda99a8f113715d23875278bca133c730b2

Observation 904a1599-9831-48bb-94b2-36ee6e0795b0 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:30:45.607360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:f2e975cda3f0f5963226b6a98004482a54dae3fd027d310c7a75cf9130251c3c

Observation 71f09531-1f2b-40d1-bb63-9db123bdb778 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:10:16.464451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:263fc7424bdaa746b0cdadc5e979ecea6cf45056f6368585b524164c5aff08d6

Observation e5540bf1-04d1-4ece-bd63-ead3acd460e1 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:7990d3dde88b5f625248a6d72a75fff53a0e4cc5eb073bc3c13d51a1ee438ec8

Observation fc6ada6d-7735-42f7-a4ac-6ac66b0584ba · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:14.009549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:14.009549Z digest=sha256:22c7cc00b17a0ff24be8e80456239ae8e0b3a0ed065f3db7c23e113a28c6eb3b

Observation a91f53af-6f4a-4505-a425-5dca60de01fa · inbound

SoK: Agentic Skills -- Beyond Tool Use in LLM Agents cites this paper.

SoK: Agentic Skills -- Beyond Tool Use in LLM Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:19:31.260996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T23:19:31.024268Z digest=sha256:6650174a534d7ccf2e034e8b876b02a5e6d713cd5b9d7f0270cd1746f40c9d13

Observation abcc432b-10f2-4292-80d0-40da3a898ca6 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:21:23.252058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:fada51008b7282b071a318fa06d7baec86d37350ef28ace6005f59e02c7ae599

Observation 8c213217-51ad-4a18-8c99-a06869988851 · inbound

AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning cites this paper.

AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:49:49.335166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:45:59.036473Z digest=sha256:6d35ecc106fe848ed31ea51e40cd6df32224f3b9d4b667de87b562144b5195f0

Observation 298b6312-b1f0-4571-95a8-d2461734e9fc · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:23:27.095484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:b57b4eabeccfddfe7522082532728d77015f261c56bb77f2969691535e2b403a

Observation 8633bda4-be3a-45df-84ff-f2516609ef69 · inbound

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents cites this paper.

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:08:21.527478Z digest=sha256:b9df7fecdf4ad25ef3f69dcba04e5e2c89a61608057f1b3b82b6432c9323aec4

Observation 44285b12-253b-4eb4-a757-3f1a25d360d0 · inbound

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents cites this paper.

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:00:27.121649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:56:56.794206Z digest=sha256:aecfb3a4e0f695bf9509b8acd3536cb522583be02346b154e9c803b820be1c71

Observation 40057c14-2a2d-40d2-ba8a-dc9ed847dd4d · inbound

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks cites this paper.

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:49:32.983778Z digest=sha256:144ace3ef0a7d426e7533592c05fba32d83e60f7892671a3ec0cfd43e6e57811

Observation 2398afb7-84a7-4154-aee6-5a595908d505 · inbound

Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems cites this paper.

Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:22:52.879124Z digest=sha256:7eccfb25b63846994c6b498f8b5248fe25ff7ddc7fd7cad68036808bfbb56f32

Observation 032d5d0a-5811-4825-a3f9-abe5baaeae9f · inbound

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web cites this paper.

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:00:34.401698Z digest=sha256:9a6a4e6f00fd3a83b5db339d3fac3d0c74401e8ac45f9ef195ab772377fc34a8

Observation cc198b38-781f-46a9-af33-525685b05aeb · inbound

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment cites this paper.

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:52:08.991150Z digest=sha256:5b7153f28e3dddf5aced0db7577410644dbe680d38bc475376e901343acb4c83

Observation ecc082ca-81bb-41b2-aee2-eddc131da08c · inbound

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment cites this paper.

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T23:36:55.286437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:36:55.286437Z digest=sha256:f03ce7e25301d29ce92ac3446270f72d38cf9a5ab1f036ad39c19d20f227ca0e

Observation 6d2c8c00-37d3-4151-941b-2e26111b9c2e · inbound

ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents cites this paper.

ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:23:07.187666Z digest=sha256:2ddb0f30e804fcc1f8e51fecca6854ab0ccde499af6907128b6fdd41744e86b1

Observation ee85e413-d8c0-4d3b-8823-d78918b9f5f4 · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:e802b48d087d76cb569171f763a1310fdd4bfc8b0b42165b564fc2380e0f941d

Observation 49a3dcac-b105-4d31-93ad-642ccc14491b · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:39:15.820777Z digest=sha256:d4bedbe62416a325c52bdece5e575f77ee04192227e83062b3f018dd6869eb0d

Observation c6e0dfde-0040-4693-b42e-d6b1c9497225 · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T20:25:41.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:25:41.773345Z digest=sha256:e0136d0658d9148309b07078630d67b46cd649a868d8bf9240f3726495e21184

Observation acf8215c-e671-43dd-b4bf-76b79fdcf04b · inbound

Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation cites this paper.

Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:19:54.809577Z digest=sha256:e3f0995f17d62b5721aec9e2fbce0f749523dbc6a4bafdd58b4035bfb9410f9c

Observation b2fd0831-55c3-4b44-bcee-94bf73eab6a7 · inbound

Feedback-Driven Execution for LLM-Based Binary Analysis cites this paper.

Feedback-Driven Execution for LLM-Based Binary Analysis OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:40:32.133423Z digest=sha256:f009f5179c0e79979aced18e8fd7f928df6af5c6e0c0967cce99d56a311bfaa3

Observation bcd0509b-45db-4e80-9215-df0592d69242 · inbound

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs cites this paper.

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:45:17.043107Z digest=sha256:1b9cdf996e28b405b804af61b57772b28d68c02006878ab54d61e1fb36df1094

Observation 1a843195-225a-4412-b517-bfb9ca7bb020 · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:30e91b564992df1fbceb1412023b09068b08a540db0fde3b81eec824bf0c66ed

Observation 3087765e-665b-4789-80c0-4b10fa6ad09d · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:545190f8842a63011188b8cf214a0c44866c3a1a60af2efece40c090a4375502

Observation 1c2be504-a7cd-435b-aa65-7e1cada043de · inbound

An AI Agent Execution Environment to Safeguard User Data cites this paper.

An AI Agent Execution Environment to Safeguard User Data OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:14:40.639143Z digest=sha256:dda2ca68e8df87f4afd896dc6d1532536c92d8bae5f4d6e4c2742c7e48f4b0e2

Observation 9c3fde8a-7c01-4147-8f95-ec5d9399ff71 · inbound

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption cites this paper.

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T09:33:27.714427Z digest=sha256:26db666b1dc78d2fdd442d70644a40fc7f8c3cc823c4a2656d7c9fcacc725875

Observation 39a0e033-0990-49a6-a3da-4e8240976b5d · inbound

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption cites this paper.

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T23:53:51.353667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T23:52:17.061899Z digest=sha256:365ebcb6598b44b58f4c9e65e85f7a4480b8932fb447eb88f3b2362b05d2acea

Observation c8618001-0809-4a68-9bfe-fae470de5136 · inbound

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles cites this paper.

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:54:08.156511Z digest=sha256:c3ef6972011796a5e4c0897b13bbbad748440de25cd959d55a482fd7ca4fb0cc

Observation 2be37d84-b3eb-4d9f-8621-3abdc73dea44 · inbound

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles cites this paper.

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:59:49.257627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:56:44.316029Z digest=sha256:7ef5c996d32823621309ce028b4873fde0eda67757e9b53de4485dbc08c4fb46

Observation 7ee0a034-8c83-48a4-ad06-a7a32eebe56d · inbound

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents cites this paper.

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:52:16.162336Z digest=sha256:9f9bcb8494fe94a2b8de810472d77dcd57991f65094581f6e98a3143034276fe

Observation 312c40f5-f571-4b74-84e4-f205ae9638df · inbound

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction cites this paper.

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T14:53:52.019677Z digest=sha256:d1382bd1db2d65ea1ed3c278465e48b73ebad7e99d42f0722584c84b56c02f13

Observation e090455c-764d-4059-bba2-527dd17e2491 · inbound

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction cites this paper.

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:51:21.719603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:50:00.415616Z digest=sha256:17a82ce5296bf3f014d6ae4a0f513e7dc06e1de552466d25eff6cb606029c218

Observation 7fe53c18-0841-40f0-9444-da361312a449 · inbound

From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work cites this paper.

From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T09:50:30.639962Z digest=sha256:57be06fa730b0d1b8784e54c21e0f7897e41197ec4ff0e1c70d609044b5a6f23

Observation cccaf9a3-0569-44c7-954e-1a197f8918cd · inbound

Computer Use at the Edge of the Statistical Precipice cites this paper.

Computer Use at the Edge of the Statistical Precipice OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T01:08:15.736001Z digest=sha256:2c9228ca56ee9ada2c187ef5feef73283c14683c4d390e9737c2a0f9a62264ce

Observation ec5fb971-a6ec-4eb9-9ade-2d989e18b270 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:b8ca496ed70bbad65bd874788458c07cd5f2968fb3969b8a87186ec37acb8eb6

Observation 91415c43-8a4a-4889-a950-cf58d03451d3 · inbound

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks cites this paper.

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:03:22.390574Z digest=sha256:b7c0beba6d745092d1138ce2f3874da717d19cda43aa5ecaa6c014c4e7a98608

Observation 129bda9f-d70a-4730-97fc-d03263e9b001 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.835280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:277a8d169bf576e4c88853ee6c21dff1cfdcc36089abd73436dd1cd9d154f243

Observation 628d2910-eb0e-44f0-a840-a01d93968117 · inbound

ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections cites this paper.

ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.178212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:12:50.686474Z digest=sha256:4579026d28df1579ef07c4745c8f23b53612a0ae0511a74fdf0a3d814c538fd2

Observation bb2eabd7-0d48-4ba9-8e92-d94058ae0a1b · inbound

ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation cites this paper.

ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:19:46.267389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T05:17:37.850154Z digest=sha256:49ab322f371eed5ee195fd008a7475262d40bbdf352be852d9372274c1e20f2a

Observation 17efd460-2324-4a42-8f49-25ed16f9ddc7 · inbound

ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation cites this paper.

ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:33:43.129581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T20:32:10.141343Z digest=sha256:5888c76b840b80d6278cbac65b20a030c3b2081cf06cf3eb05008d65c0058c0c

Observation a9b79432-6d19-41d0-9b3b-99d0ea5a0813 · inbound

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? cites this paper.

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:48:48.826044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T17:45:02.896703Z digest=sha256:c5674161b0c3c0cb5ba6beac6013dc9fb7321c2a016b2fff9c62c37c6d52eb44

Observation 485adf21-7650-4e82-b235-0a94a4437746 · inbound

TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents cites this paper.

TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-19T22:57:50.193076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T22:55:28.637039Z digest=sha256:e49af75fe57c85392280884143385cb9c4dad74f51b2730d598b0d56dcddf1f6

Observation 92dbda9f-7b41-44f9-9101-2dd726c6d696 · inbound

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding cites this paper.

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T22:12:50.568653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T22:08:27.511727Z digest=sha256:036b66674cc6c8ccce6f78b8522e19222e22d581dae969081fff5e325ba1acd2

Observation c4b78e58-b813-4555-bd31-05a4205fa2c4 · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.662722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:46f5d4e1848b80bdee00fa960793eafabc9397d81eea6c6a306e861ae6d55f68

Observation f04903dd-771d-4f29-b4d4-5a70e275a4c1 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:08.002848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:f1843913afbcd3c33b63f672e4d30da0c170a49a99e8ec09abcb848b4b303ead

Observation b4444d1a-6104-4618-8ef6-77bc024d2374 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:06:42.834544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:14ffb473dcfaa93f74b5f7cda6e31d81d15b28f7ba4265ff3c322a7a17b3e89f

Observation e875bef5-035f-4dc6-9d21-f8bc8f663f7a · inbound

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis cites this paper.

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:04:37.692134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:57:59.013811Z digest=sha256:e39ca00100b29f3d544e3b017fa1ed31c08a789d5796edab39d0a0285f63ed9c

Observation ba4257a2-04d5-4132-b839-2a476e23b5a1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 133

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.790934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:e9485d4eaa30932dd5f4a01856e7e4d33ba32b2500944dba89d0ef7a5d05bb6c

Observation 0b9f29b5-77bb-4699-ae9e-568309c91f0d · inbound

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents cites this paper.

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:54:45.652180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T14:44:53.335178Z digest=sha256:bd7b21b951f7eff13f0d320b372ba8ac7aad4f0d6ed7a206658c6941ceff5f6b

Observation b538d906-a148-4471-a20c-adb3ef165141 · inbound

JobBench: Aligning Agent Work With Human Will cites this paper.

JobBench: Aligning Agent Work With Human Will OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:33:59.457856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T21:24:29.232709Z digest=sha256:50efc73487c56f0cc157ccd59233504a4a2e7d7c6670d57657a037ed016efc7b

Observation b7054e13-65f4-4ff1-9991-a6e79549a425 · inbound

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents cites this paper.

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:53:26.367017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:52:20.911788Z digest=sha256:58c7f6eccc837022738e3f9eccb42c9299d930e9e39546dc6ef633b21503604b

Observation 99d73b7a-c0c8-4b2f-9ff9-874eff3809cc · inbound

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis cites this paper.

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T16:13:36.055090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T16:09:27.852716Z digest=sha256:395371b9e7c69547a00339c16d047063e0896b92872ffdfa55e038975f5a709f

Observation 431c9706-c82f-443a-a499-cd75dff39cc1 · inbound

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning cites this paper.

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:23:20.863611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T11:19:38.959705Z digest=sha256:e6adac80cc7b95d07770d4771687f82cef69bda8356aa74e94600db96deb58d6

Observation d25a1708-869f-4877-94de-43d5f06fc7b7 · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.120068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:bb4707855235dbba20b6970fb883c09d1f88b4f84a05c73ce39c183a67c4e005

Observation 204f3772-d73c-4ebd-819c-6a34d06aa610 · inbound

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming cites this paper.

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:06:48.247592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:22:03.848058Z digest=sha256:d21d1fb1e91994b2f4fb91288ee48552542065228291a6a3c15ac796a7a74376

Observation 5bfbde66-5caf-4e99-b9d3-61db7a082ccc · inbound

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments cites this paper.

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:56:56.823942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:45:28.693098Z digest=sha256:48c18f4c45fa596c6371d80c2b9b2d8e3b93551eec57bf489337356deea12695

Observation 59e18fa3-dd0c-40fd-becf-acb2b4d70162 · inbound

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions cites this paper.

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:41:29.382021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:37:45.812577Z digest=sha256:5ee1655b0a93b2ae03df7126304612717d040abdc3848ec238623f81968dd24a

Observation 986ea443-d739-45f7-a4bd-2a2424c8ba3d · inbound

Signal-Driven Observation for Long-Horizon Web Agents cites this paper.

Signal-Driven Observation for Long-Horizon Web Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-28T01:31:29.182702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:25:11.149977Z digest=sha256:64226687a1dfb4c86e3447b085703ca76f07b11424514eb611cce43f58bb468d

Observation f471f672-402a-4c1f-9f79-c3c8f5afc7ce · inbound

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents cites this paper.

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 4

Resolution
malformed identifier
local_arxiv, observed 2026-06-27T19:51:12.581395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:45:10.664592Z digest=sha256:c41a1b7f68bc57089c2451e8d7162d1ec448ff103ec0674afeda2159a93b6a79

Observation 676eedc4-f966-46ac-8a5a-bf48f9198938 · inbound

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces cites this paper.

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:37:29.834859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:07:39.645843Z digest=sha256:c096ca1e111cdb77f801bbcd9983919c81dbdb0ac021ac4c309bd9970a133363

Observation 9034f7da-dcdf-4ee1-8b55-ce51053694db · inbound

MedCTA: A Benchmark for Clinical Tool Agents cites this paper.

MedCTA: A Benchmark for Clinical Tool Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:07:47.973992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:32:01.387453Z digest=sha256:298a32c8039cb7294582bb395b1b446a34122e27aaf7209a95d9e8bc97fa71c4

Observation ecdfb51e-749b-4c1c-8f12-830a4a1967d5 · inbound

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails cites this paper.

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.570651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:42:05.886984Z digest=sha256:6220b0a72f961ee335c0f23d099f87e3cc545ffe7df44346a600cf5c0efec490

Observation b0e318c6-2350-4f65-9b7f-ef12c8db6307 · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:38:44.116785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:471d0b0c76c51e84f6c33cd1efb3e1e78a97b04f746b46fd25690445d1ecd03d

Observation 4c6f4af5-c226-4b7d-8677-afe15c476945 · inbound

ProCUA-SFT Technical Report cites this paper.

ProCUA-SFT Technical Report OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.692464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:18:04.149281Z digest=sha256:e670471b1a6292b3c1903c432f7c4af4bf19a9606ca0aeba6fec4cd8773b0bbf

Observation 20c23fbe-6d3a-4871-b31f-22cbee475275 · inbound

Dissecting model behavior through agent trajectories cites this paper.

Dissecting model behavior through agent trajectories OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.935465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:27:39.812496Z digest=sha256:e738ecc7d88dbd64a2642727f2d36e3a2e1ec12fa777c58deb2a2496c55525e6

Observation 9fba41a5-172a-477f-9198-c5c2054a91e0 · inbound

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning cites this paper.

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.201947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:26:41.376368Z digest=sha256:e8d2e147b064771dd0e7a3b23efbb874eae1827372953a792acdd03df2c95288

Observation b24081f8-6507-4cb3-9fc4-6f529091bc0f · inbound

OpenRath: Session-Centered Runtime State for Agent Systems cites this paper.

OpenRath: Session-Centered Runtime State for Agent Systems OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-04T01:29:23.057631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:18:48.243426Z digest=sha256:f48516f7c95f822a03995159848c2c3f4638c608fc241e7975604251045c39e3

Observation 634fb884-4b43-4c52-a1ac-aa63bb85ae65 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:19:23.924611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:b01aa347562158674d2adda9cb4503aa3c5d53d21d55178253d14a682efc34b1

Observation a61e9752-141a-4f5f-9513-355acb2bed54 · inbound

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents cites this paper.

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T21:08:58.347763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T00:54:00.944706Z digest=sha256:ae5e5f65880b2f8b48bd0ea1c844821b23fa7b386a7c22a9adeaab322ef3e780

Observation b8cde36b-ff87-4f3a-9bd5-2d4efbadff3e · inbound

PhoneBuddy: Training Open Models for Agentic Phone Use cites this paper.

PhoneBuddy: Training Open Models for Agentic Phone Use OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:59:46.504700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:15:49.428124Z digest=sha256:8dc768afa78c1eff767757a837a6b1c8caf89639db8e925f6604eb0873126bcf

Observation ed6c4d3c-cf0f-4a99-8330-cd5ab95518ac · inbound

Agent vs. Parametric World Models: Hybrid Planning for Reliable Language Agents cites this paper.

Agent vs. Parametric World Models: Hybrid Planning for Reliable Language Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T11:37:35.542449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:37:35.542449Z digest=sha256:8fd0f37f4fc6174ff4b85db809e961f53fa2f975e1ac86b524649b4674c5eaf3

Observation 90aa53f0-1a45-4f9c-988b-3726014708b1 · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:04:21.271037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:30f22b5e34640a2ecb4aaacac296a655725340fb5bebc34403bbfefc8e167370

Observation bc1ea22f-ce92-4b19-bd9c-78e4d25964e6 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:21.100795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:894b337b1d7175f24d440b40c1d1bf98ca055cf9462cf90cfccd238d30fe0004

Observation 958cd20e-c09a-4896-b797-a35e52443aa0 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:bddfd6676f4f27f014b756b93d309414a605e9ad650d9a56b7638288648ced9e