Pith. sign in

Paper Citation Record · LEDGER

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 80 inbound Pith citation observations for arXiv:2509.02544.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02544 v2

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T10:13:58.774968Z

measured 170 of 170 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 80 of 80 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:03:39.391482Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact54
  • verified fuzzy31
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch4

External citation measurements

31
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation de3d4e3a-9a2a-4cfa-aa79-ef2c6740547d · outbound

This paper cites Introducing the model context protocol.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Introducing the model context protocol

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.284826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:92099ec9f840ce5c5797f6ecdf82fcb86d32c496cc75f42fe4d76967684e809e

Observation a5f31d35-4d4e-4790-89e2-d5bd1c744ccc · outbound

This paper cites Developing a computer use model.https://www.anthropic.com/news/developing-computer-use.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Developing a computer use model.https://www.anthropic.com/news/developing-computer-use

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.193462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:dc01c2f7df651fd7a8543155d745055eb5dc4955eaa69952eadaeb4082138dea

Observation dc912f36-e514-47f7-a022-426085393707 · outbound

This paper cites an unresolved cited work.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Unresolved cited work

Reference 3

Resolution
parse uncertain
raw_fallback, observed 2026-05-13T10:13:59.197868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:2694790672cc5d0ffdad7da1c8906d4153858e2f8298c00f7d3c500c0c8947fd

Observation fb943cd8-d57c-4b1e-9693-70659930bdf8 · outbound

This paper cites Claude 3.7 sonnet system card.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Claude 3.7 sonnet system card

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.202053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:1d8516695146ead331c428b02b41d217855bbc536a8a69d15739ad717b4dddf8

Observation 95e7cf4f-282a-4cf2-87c9-7b9a89a0c899 · outbound

This paper cites Claude’s extended thinking.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Claude’s extended thinking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.206172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:26436720a2c62db30aabe49e0bff016eb7d7bb231c930f0a4404bfbe76243096

Observation 5726bbdd-ce81-49b8-93a0-532a3664fd6e · outbound

This paper cites Introducing claude 4.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Introducing claude 4

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.210724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:f45d3c218bc1d3781b312d4e86ba2704a712942e7ff35a17064a8a8263d7b2b5

Observation 59ff5609-7a39-4fa1-b975-cd340c4e02c8 · outbound

This paper cites Scaling data collection for training software engineering agents.Nebius blog.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Scaling data collection for training software engineering agents.Nebius blog

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.215282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:58ec4650028c03e408419ac25783c4a3c48b21186b1724253851dd97bb22aaf9

Observation a733e41c-1e01-4974-a726-211e4251776a · outbound

This paper cites SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.066796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:ddc119a7cfa889b2996c93e4b6db197178fa1ced2096366e95e3f320e89a038a

Observation a3e7f836-8647-43ef-96d6-d0883fdd8915 · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.Advances in Neural Information Processing Systems, 35:24639–24654.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Video pretraining (vpt): Learning to act by watching unlabeled online videos.Advances in Neural Information Processing Systems, 35:24639–24654

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.220414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:5c2b30af6ca0ca1db6a6a427c95a9786bf9d3c02c1a8191771cfa8f557f6a5ae

Observation 70f7b388-d9fe-4fa4-bcf8-c349a60ff57b · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of artificial intelligence research, 47:253–279.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning The arcade learning environment: An evaluation platform for general agents.Journal of artificial intelligence research, 47:253–279

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.225234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:dbf71ab1b97e1cf976882cf98dee69f7900e090e0c9b1f0b34cc2eaa62384545

Observation 042026ee-35a6-4706-a152-bdbe14b7f2d0 · outbound

This paper cites Windows agent arena: Evaluating multi-modal os agents at scale.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Windows agent arena: Evaluating multi-modal os agents at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.230301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:08ce5b29f29b1207410e70c48ab7724528c81a4d38988942b17799029025bbf6

Observation 1497993b-6d28-444b-acb7-7fa68e3173c1 · outbound

This paper cites Seed-thinking-1.6.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Seed-thinking-1.6

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.233720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:d91e5745b147e62911d49c341260acf5cc99685a94069d8a9d055d1954579b9a

Observation 5084a61c-7bf2-4f8e-8e74-8c64999d6241 · outbound

This paper cites Mindsearch: Mimicking human minds elicits deep ai searcher.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Mindsearch: Mimicking human minds elicits deep ai searcher

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.085850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:cf69a98ee3c0a5334158c859be3be6089d4e77963b2bd16bde46dbf186389ae9

Observation 379666f6-5021-49b5-9f28-3990481137ae · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.778093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:471939f00d62c8b74d67625b40bf709c52457d4b568a8e08a355bd759186d1d6

Observation 4684eb12-2aa2-405d-a34c-f7c6b8dd8e72 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.060146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:e939374b6354a863e06bffaef0390eb12235031593664fd67c7cd68d74c1eecd

Observation 758f78d0-9afb-4382-a7d2-caa6bd7ed31d · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.236547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:11ee2ac0b4bb52200d9206057bac4b5f07a66178f01d3b73856211d9ea3ba3d4

Observation adf86baa-3cf9-435a-89fa-4d2d1ea0c52e · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Mind2Web: Towards a Generalist Agent for the Web

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:05:16.335224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:800762728f685fe000e43c92f3cd363e16f5810f4cbc4fc66c737cdd53fb5c69

Observation 704c45d5-1748-4bac-b1a7-73e15fe6e9e0 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.240350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:3aecccb15ec61b563ff7b73c3d2ad786a6c6b9dd443754a4d39fdb02150fb7e2

Observation 15f6c244-687b-4b84-9943-2525f47874aa · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:42:39.160125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:560118f7d65af159713bb4ae6dc2eef0051339620134b7a494f4c81d4dbcb158

Observation 615ff5bb-32cc-4172-971e-b01b9679b31b · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:19:36.299641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:275530747239980dd60cc9be01ac7547403c161de08319d37e5821d012c40ea4

Observation 0a96da0c-de26-49d8-baa2-e6fd0fa5b25a · outbound

This paper cites A Survey on LLM-as-a-Judge.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning A Survey on LLM-as-a-Judge

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.103962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:dbf55e5558a96ee0150ab8d8f303020322ec686ce524564300afe220fa2d34c5

Observation bfd63874-1468-450b-9ab4-868a68cc8e85 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.109629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:49ecff1e3b7c48d828756fab2acad36df5d649df8ef8dbf9055aa4cb13a29b6b

Observation a3d3700e-8dfe-4eaa-84ad-4f7d0f00f94b · outbound

This paper cites OWL: A Large Language Model for IT Operations.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning OWL: A Large Language Model for IT Operations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.115701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:bf04efaa7f3586daad5545682942eb5468d8a7dde30703b713de7b07301c2c4f

Observation 19e7957f-4545-4596-b6ab-f35abe9f0984 · outbound

This paper cites Cogagent: A visual language model for gui agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Cogagent: A visual language model for gui agents

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.244186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:878aee35eedd0eec691cb60ee99ad98278f17139a34eaa52155072efbe4f38d2

Observation c87a4664-3ed6-421e-bf53-762d56b3599c · outbound

This paper cites lmgame-Bench: How Good are LLMs at Playing Games?.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning lmgame-Bench: How Good are LLMs at Playing Games?

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.184183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:7c751b52962086721a9243c1911405da946d9cc6a4dc8281b70cca9a800274d0

Observation 944110c9-c093-42b8-b993-20daf9b17846 · outbound

This paper cites OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:58.845816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:879b9640f74a77f957e4c45fdcf962d8070690d0d6c1c7eb7e8875636600d955

Observation 58b2e9f9-264c-4002-962f-95023d9d2a37 · outbound

This paper cites ManuSearch: Democratizing Deep Search in Large Language Models with a Transparent and Open Multi-Agent Framework.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ManuSearch: Democratizing Deep Search in Large Language Models with a Transparent and Open Multi-Agent Framework

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.880912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:eb4fcabe42efb3b836ca0e9fff6212f5e65783310e06daadb82e9a7b3e4d66d9

Observation 2fccd53a-de22-41b0-ac7e-6d0151c13033 · outbound

This paper cites OpenAI o1 System Card.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning OpenAI o1 System Card

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.886621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:87bf6f9047850f608a95c22ffd83914730500a4ac054fdbe89b7bcba47d9f38b

Observation f1a0c3ee-5aba-4d9c-8e26-4eb22e3fa5bd · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.918697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:47091b01e18d5d070e46bca952f396dc1f41ce9486a32b2bb416380181209c6d

Observation e8a8c77a-9766-4de9-b7af-5796656fdabe · outbound

This paper cites MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:31:08.404728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4ef588c26e2c926b1b9fdd5086fd7ba87551d6c36a527a0e9c65557736115e45

Observation 30ef546f-4f12-4a81-8aab-221dfeebc95d · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:144e81fa08a45699825aa5ff3b8f71220dc95ec489aa3c34d843eb1af3313fe8

Observation 2eff03f3-e254-4d06-88d9-e3454eef8b1c · outbound

This paper cites JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.004983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:8e8b55272fcf31b9f06a1a573f1f3f8d65913bd08352e5e455319bccdb6939bf

Observation fb9bb40b-15bf-4aed-b1e6-e1d0f0caba50 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:36:27.820666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:1ccdcebfe24530e952ab3f9acd54cebfc2d4fb75bf42d7cd6cbad507df7f9124

Observation b9f49a87-c1ee-4848-adad-06a39be94915 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.036699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6d3d77a452b4973d66d0d008e2a69f42f455abaf720f2f043d09ee7655e586a4

Observation 1ded9e82-d560-4c8f-a2e2-06e552dc3762 · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.041986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:5747360508697234e0bc43c023433fbbd8d2a31169c182fe18214bee1c0f7ac4

Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · outbound

This paper cites ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.047730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6048b98b3b499865ac49da8123ee8cc8e7532d49520a36a854f75e1666c78917

Observation 2178612c-4e14-449c-ab0c-e5596a49b28a · outbound

This paper cites RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.053557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:b6f550a81e1efe331c663113997695a25b511ee166934186b967fa8ab7e9835c

Observation fca6930a-71da-436b-b1eb-12b645b31734 · outbound

This paper cites Large language models play starcraft ii: Benchmarks and a chain of summarization approach.Advances in Neural Information Processing Systems, 37:133386–133442.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Large language models play starcraft ii: Benchmarks and a chain of summarization approach.Advances in Neural Information Processing Systems, 37:133386–133442

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.248418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4ec8f62772dc7a626c11c607cd60544ad70cf413662f90e421676559ee88d680

Observation 51f24dd3-f343-4273-b841-ff8ec47cecf6 · outbound

This paper cites Human-level control through deep reinforcement learning.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Human-level control through deep reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.252471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:ace25787877917d6cfe5f6e6898883bebeb973a10539853c590c0cafd287955f

Observation f0a3ea3f-9fdc-4838-87e0-48ab8c7dece4 · outbound

This paper cites Kimi-researcher: End-to-end rl training for emerging agentic capabilities.https://moonshotai.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Kimi-researcher: End-to-end rl training for emerging agentic capabilities.https://moonshotai

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.256633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:e10e8f982186c61a94f1db0d802ee97b9a192a84018a99aac4f990bc2f283527

Observation 3fca4f54-49c9-44c3-b37e-2375d06705c7 · outbound

This paper cites Gui agents: A survey.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Gui agents: A survey

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.079422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:9797626a3b71cb6c6ee94a28cc47c1778407157ed97483627e6bdad4b5e19651

Observation 2af4f713-f4f7-48d2-b794-1b627170012c · outbound

This paper cites OpenAI: Introducing ChatGPT.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning OpenAI: Introducing ChatGPT

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.260936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:c1ec2748abb68baadfb07d2f9070f826917ff3638c957f2b6e6c3cc2342218d6

Observation 0e53297b-4131-41a7-97e2-ac3dea75efcd · outbound

This paper cites Introducing gpt 5.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Introducing gpt 5

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.264676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:601948b2079791ce241ef8d757c23a3248443afeffd3367e1faaffe1b5624cb0

Observation c7a3b863-07b4-4ed4-8573-e5e7702f2c80 · outbound

This paper cites Introducing deep research - openai.https://openai.com/index/introducing-deep-research/.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Introducing deep research - openai.https://openai.com/index/introducing-deep-research/

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.268878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:e06f854718cccb66ad4a708b90ed828caa40fef1b07028ec94793dbcc05f80df

Observation 44dbbce9-e221-4f13-93b3-a5e01a1ddc91 · outbound

This paper cites Openai o3 and o4-mini system card.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Openai o3 and o4-mini system card

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.272797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:7901e1e63a13f40fd813a1dfdd99e6d487d1943dc41173d02c817933aac8c971

Observation 467dbc45-7e12-4acc-b672-b556c1b6e17e · outbound

This paper cites Computer-using agent (cua).

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Computer-using agent (cua)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.276563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:ce3a94cca2b833e14140ab46e0596bc235c9ba1dfa85284c0191107f8b6a2dfd

Observation 86231617-ca47-499d-8c31-7f2100926fbc · outbound

This paper cites Operator.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Operator

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.280752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:d790e3839777b32e95a2b0308b2ac2aff83c4ebe7f21c2908bca087b3ad7bcff

Observation 1c0a0eb8-eeb9-489d-beb4-0464fa0eac73 · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:40.573423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4cd5a66a32341e1c4bb15bb157c693d8d2b12f4b90a5241146b570c095b48944

Observation c95661c9-6d35-49f3-9023-96457b42d33a · outbound

This paper cites Exploring mode connectivity for pre-trained language models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Exploring mode connectivity for pre-trained language models

Reference 49

Resolution
verified exact
doi, observed 2026-05-13T10:13:58.832471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:d7ff0c28673d96262f10b536c0d697eb7b6fec00e63722affa8639590640cdec

Observation fd4f7368-5270-49b9-a630-759aade2bdd3 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.127832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:5d8bab062e8d45c8ce3c56b819875c3beef95688eddf075eb3791cfb844da817

Observation 93d14fd2-4d9f-44ae-9485-adc41df0567e · outbound

This paper cites Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.132799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4cb761f1bb172ee58cf0eba302a61604ad54cae452940cf456d25e7a2d566b8f

Observation 4081edfb-eb3b-47dd-be6c-6aea0658c3a2 · outbound

This paper cites Scaling Instructable Agents Across Many Simulated Worlds.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Scaling Instructable Agents Across Many Simulated Worlds

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.137373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6efa59113e00e493021a4df8ee18251a29045a38ea195fff132ed92e092bcc61

Observation b8218296-d4d1-4fb3-9549-86778b499ced · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:06:13.928391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:9f6a25307eae45fff487e7a9710c8fcbbb9972a84cbb17d5f684a265581c018b

Observation 7773e536-5c0f-46cd-af32-387bace27bcf · outbound

This paper cites A Generalist Agent.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning A Generalist Agent

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.148204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:90d3a0758f46910692e5a280bf52c8deae33f120090002965d94a512350cc46a

Observation ee4951a7-96b4-41e8-a93d-52ead39690a1 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.153816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:fd5abbb50bf98dc2453ae0e90ede7ab0494b18529e44f71facbd8cf852f8b059

Observation 7b3ec73a-373c-4837-b353-3632485fd75e · outbound

This paper cites Proximal Policy Optimization Algorithms.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.159784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:b38017cd9553a65d9b566806769d1ac256b30937c49a644a101ffbd949caed83

Observation 7f5de416-8b05-4eaa-941f-81e6c38fd012 · outbound

This paper cites Ui-tars-1.5.https://seed-tars.com/1.5.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Ui-tars-1.5.https://seed-tars.com/1.5

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.188438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:d5a41a6a6359df0d9c155e3489416b5bceab71d54cfff4bd38eac2ad4c1d3782

Observation fff9fb30-d0d8-450b-b418-2852e4dcca0e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.178104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:48d13fa5627e355fffb8b03c9d6c3e7cb48a877beca4fa42d882c9d15cd0d77e

Observation 63f560cc-b7a7-414c-9742-3d60d8d1ffd6 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.172725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:81aac775b7bab722f4c00748e99839a1d1f8edd78769fcf4f335089db7c2de96

Observation c068340d-82e5-4bfc-8a97-358fb9225445 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.288928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:51762be6ae5a1fb50e943bdc5dcfeffcf7d682128109b351b9d145247a6441e8

Observation b597834e-65ac-4184-aa7b-55a3f85b7d78 · outbound

This paper cites Mastering the game of go without human knowledge.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Mastering the game of go without human knowledge

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.293160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:11801f753b355e1352e017ca7ec3117c77a7335c934d736ae23ef5fad5922915

Observation 5190ed19-ff39-4025-9f0b-1b0a40ef037d · outbound

This paper cites R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.852583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:af5be7cb56268947a9e03fabd1a17cc636c4704162430a558b7c1cad86351e21

Observation c5d5f36a-580b-4632-9c2c-f75b2d64bdd4 · outbound

This paper cites Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.838223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:1ed7ca9940cfeaa0ed1e3374367c0f223323c57dc60fef3e9f932686efed74b7

Observation 3b79331d-4abf-4c99-ae87-80df533c1337 · outbound

This paper cites A Survey on (M)LLM-Based GUI Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning A Survey on (M)LLM-Based GUI Agents

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.859933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:647f4b3f66a3e1fc18aedabc2ba6f224d1d16b9d7ccc6c8a9d6c6ec68b2dc65e

Observation d1ca66a6-f7b9-400e-9271-02650cf71938 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.866070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:bb964a227f04a4589ed9d73d9d768b57b361aab096774ad639ff77dbc0fa170b

Observation 27400637-b343-4f51-a6c7-949a21e57c1b · outbound

This paper cites Qwen3 Technical Report.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Qwen3 Technical Report

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.873345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:dd674eb135570b4d85186e0e126daec7794a195f49bb838bfed134b01f6b2e49

Observation a72b6e33-61c0-4fe5-925f-537cf3c6a119 · outbound

This paper cites Terminal-bench: A benchmark for ai agents in terminal environments, Apr 2025.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Terminal-bench: A benchmark for ai agents in terminal environments, Apr 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.297516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:28a8276779d556b726928756ada9acdbddcdf5778bb9353a48b800faf504790f

Observation b4f033b3-4bd2-4ac2-8986-e97faffd5525 · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.301733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:f491a7da06f544e64abcc30dd6879d397124fa0f1f312636cc11d1d7a6c40288

Observation 61cca9ef-94aa-4447-a9c6-b488d97009d6 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.892299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:cebd3de22bce95774f9379bee393ad7747c18517dee38e8da206bb72bd6501c3

Observation d4973e55-2653-4d3b-b064-4fcac0a9ad7a · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.899300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:726b87e81d9df48797cb06ba7c25ca93121d18197bf97211e155c9d8392023c9

Observation c7440ca6-9008-4b1a-b009-db1466a79df5 · outbound

This paper cites GUI Agents with Foundation Models: A Comprehensive Survey.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning GUI Agents with Foundation Models: A Comprehensive Survey

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.907184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:844e4b430f256bbe2555ab0e3d506895ea80a770be16ce8e6660be043e618cb4

Observation 1c7bff89-4e5a-4402-a062-ff05086fc6aa · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T10:13:58.913195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:decf8eaaf26632a9e41670db1939e4eb86a0893df765715f09e1e9f0885570c1

Observation f30ef1b8-4f0d-479e-ab63-386ec104bc76 · outbound

This paper cites Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.306336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:3b6bec933cc937feca6214c7451891a3713acea75660ef98fbf7f58bb4e3f2cb

Observation e03d7d87-8c84-4017-a82a-6caa432876bc · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.926566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:556fde99cc6154f89f0b9b4079827f0a1a59ef3380a57cff6b45e8dabba289fb

Observation 08fcf8ea-563e-48e3-911a-1627a6198afb · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.931773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:909363212f71031c965f082e64a3e01b3892503355bee43e7a63013cf4ac3135

Observation f5cef7a7-b027-4238-8b53-f758e5e068c2 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advancesin Neural Information Processing Systems, 37:52040–52094.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advancesin Neural Information Processing Systems, 37:52040–52094

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.310786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:125a6a02c04afc2c0a6a9530cbbffd34cbffb3b9b7a1363611b4e90416b9ba5d

Observation 1828ce06-2366-41ec-a450-18bcf600ff8c · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:09:42.002663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:c446def6bd355dc20723ef10c42eb668a2316f4c4a91ac46c10f8267dc4e1acb

Observation bfdb9337-e75a-430a-840f-bed423bbe53b · outbound

This paper cites An illusion of progress? assessing the current state of web agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning An illusion of progress? assessing the current state of web agents

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.948017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:44b9b46e5e84ad0c80f7dd61008bcfdd18d1e7338dedd28979d0747107e025ae

Observation 45c4b309-5d99-45f3-8a83-4f91d7c6f9cd · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T10:13:58.953503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4ed82e34e65c8989366726edfe727ae018c6e0751ad285f0a838a00f1fd8f162

Observation 35636f95-d60f-40f9-b194-1bce3654d687 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning React: Synergizing reasoning and acting in language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.314798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:58b6902efd56ae0d8e7036a9ac9eaee67f79e4b2a4e42de026973727d4b647aa

Observation 494210af-a3a7-40e4-ad71-358a6e0f58b5 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.966213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:e9c693cc7a07fb728aa8a721285c2c39c96c8b8a4f9d0280d868e332870c36d6

Observation fe9b2a5e-ecb2-4902-9605-207c538bea61 · outbound

This paper cites Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.973256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:ecc8d78a1ef8740d4c8d16f2c6b3b63a37521bb3a7f24a5077b66c33613f4f9d

Observation 65780489-24fc-4086-8b40-c06b37a4b90a · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.979808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6b03623e05d3e40cf0d3846fd62ad5296107eb1c05ec5de455008198be11195b

Observation a4c1acac-b9ed-4ac5-a94a-3d44a3cb00b3 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.985799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4c7c4742f67f553aeed91ff4aeb25d6965a269050c0245611d317d5c353bb52c

Observation fce82d0d-5011-463c-908c-89594276b620 · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:48:50.578231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:34b5d2a07fae771fd81b8c0a01f6a1e37cb147eec6a097f8155e74e33145a593

Observation 325cd7da-af23-4399-a85f-939f70e260d9 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.998108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6ae810f48fe9f38dbaa5b2229a17a99a91c9b5570e5b87db8e5f8c0dbfcf7ce0

Observation 02579d63-69a6-483f-9ecc-3f05b6c82663 · outbound

This paper cites Proagent: building proactive cooperative agents with large language models.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Proagent: building proactive cooperative agents with large language models

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:13:59.319576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:89609889034244b49223a2d98ba39ad35fb1f5774a0614915f220d5f51ec7ad2

Observation f342a48a-feed-429c-b143-49a246ced4b1 · outbound

This paper cites Large Language Model-Brained GUI Agents: A Survey.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Large Language Model-Brained GUI Agents: A Survey

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:28.109115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:3141a7151d915d0d53ac03617a8e94dda235b8e20455e8a54d1594a3421e07c2

Observation d64d02d7-b344-4ba1-a763-12477effec23 · outbound

This paper cites BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:04:50.009871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:5c3a10adad4d6eab4595dc6ad72ee6837ad1cd55939c5ffc6e235fecb9fb9c78

Observation c11eb651-6cb9-48b8-86bc-db002b5dbb9c · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:59.024073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:6018d329e63962ac0774b4d248ef585d40f96dd261f9d5e94a1ffb6367eec1ea

Pith citing papers

Observation 22e84a3f-2680-4883-bcbe-c72e9efddaef · inbound

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action cites this paper.

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:03:39.391482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:03:39.391482Z digest=sha256:b4cb8f92ce094c57d3fb579fc2acb6f8a8f8e1bfaf603a7b98d46cca459ae2e3

Observation 2ec08c40-cb89-4650-908c-9bf4befba1d3 · inbound

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning cites this paper.

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:30:52.482488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:30:52.482488Z digest=sha256:9f3b8fe08ab40bbe803e14fd6dcc43326b61d3b7c2193cf81ca07a8bfd218c83

Observation 0c1bdfb4-9d4b-486d-8d94-bd7e8ec3c024 · inbound

AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management cites this paper.

AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:43:42.332690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:42:25.086602Z digest=sha256:968499220db41b862b1239d23c4b37061ec92eb60f38b31002a833affc6c353d

Observation 3daea074-dd6b-45c1-b3b6-49926c3eb257 · inbound

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents cites this paper.

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:32:48.323772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:31:16.087579Z digest=sha256:7aa2ce2891ff8a9c5f865a2afce59142581efb5704cbaab8d3a36155e7873700

Observation ec197a47-1211-4fa5-8657-ee18f9a693f0 · inbound

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward cites this paper.

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:39:45.272421Z digest=sha256:29f3c2ab6cd37553ccb5faf704d75e40fa84051dcd7b91fe33b5dee449a3c92a

Observation 87a6a619-899d-40c3-8372-6c59c8a8604f · inbound

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward cites this paper.

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T23:50:45.688267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:50:45.688267Z digest=sha256:e1cb6af39bbe6b02a39a4e6a371fd3dc44a256bf13a7442ffe04560eb0d22ded

Observation 3b732d0d-0c29-4eba-98fa-75812375cd22 · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:44.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:44.108333Z digest=sha256:7da230680db6d94e6c50b6f81ffb8a9f0749b1fb4cc6a97f0c6a7259bc63ba85

Observation 716cf421-b31f-43c4-8004-eda73b532bb6 · inbound

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces cites this paper.

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:36:17.512296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T16:33:58.746944Z digest=sha256:515ff5d13181de982662f1ed179e58b723c57ab7ef9cd20bf8d09f1ca645f55d

Observation 5e90e0a7-cdd2-4b50-82c6-f726033d0aa3 · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:00:31.594427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:99445d4e954a039bab8ea7cf3e8912d90840090292d7c23a8e57f15dc5f5b761

Observation 4cf683ed-4773-490d-8877-f19e112cd5a0 · inbound

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection cites this paper.

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:29:27.373741Z digest=sha256:87a3cbfd0c250dc7dcc06a79bbfc51a8960508f8884e0e9ef6ad2112c12eb047

Observation c7c82b1d-a33a-4814-88b3-60366538fc15 · inbound

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web cites this paper.

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:00:34.401698Z digest=sha256:d8a5a7cc6c1300709fbc85de097ba449a7b54610aea1c77150aa9813de622560

Observation 7a8887ca-c3ef-4cfa-a792-be6318c1445c · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:55:48.506049Z digest=sha256:673a90cf34189063f3a777ffb8006fdaa35fe4ebf3fd9a58358be481d0a90d68

Observation 6dae6368-e8a8-461c-a76d-aabaaf608fec · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T23:07:40.699400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:07:40.699400Z digest=sha256:86b61beeb0026dd4f3efe4ce9263dff97a40b58b3fa80d88376c5831b57ca584

Observation 6d295fb4-bc6f-4f2d-9444-1cf534063f8f · inbound

RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management cites this paper.

RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:40:55.540869Z digest=sha256:151812b6bbd05bb6ba415a0038e8709f8c9f61dcdea6fb2a499a112bdda78196

Observation 0644b139-8f8f-47ba-8ab0-173250a0fbf0 · inbound

HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents cites this paper.

HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:19.717895Z digest=sha256:3d2aa9c91b682a43d4c861217bdb8b1aefe416262bbf914e2ca60f44f471f143

Observation 3c4b806e-8a9e-47e4-bb89-f3e74a1cec06 · inbound

Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging cites this paper.

Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:09:59.957333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:06:50.635891Z digest=sha256:df1357c9c799287f28f5029d4ebd153d3754700b6ec88efcce2e8f71bbae1a9b

Observation ed4adf55-a94e-4a39-9156-6a7395238083 · inbound

AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents cites this paper.

AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:58:12.704566Z digest=sha256:bbeb811b09a25ccd38da8b9e7062b0809280b60939e47995c63069fa625891e4

Observation e9d3e4ba-ebca-4457-8ea2-1a355c1042ba · inbound

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation cites this paper.

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:24:45.045405Z digest=sha256:06ffbbf24cbc977630d89a617d2b573a237b74a011b556d03b48279780961caf

Observation 587f851c-727f-4d67-b961-4fcd22577b0e · inbound

Supermassive Black Hole Winds in X-rays: SUBWAYS IV. Tracing Radio Emission and Unveiling the Role of Winds cites this paper.

Supermassive Black Hole Winds in X-rays: SUBWAYS IV. Tracing Radio Emission and Unveiling the Role of Winds UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T18:39:55.906516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:39:55.906516Z digest=sha256:34b973f6aecd645ecb0222c162d8783082059cb5278b75ce0fcbd934fbc5184f

Observation a5b68a6b-ce38-455e-858e-e8846c863f11 · inbound

S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images cites this paper.

S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:26:03.080281Z digest=sha256:5c9b8e79bb4dff7c58d5cbe56db48152ed1e50ba872aab2aec2bb3fc4c2ab89a

Observation 83e4936c-8eb8-448e-b159-49288ca75470 · inbound

SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents cites this paper.

SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T15:48:28.869796Z digest=sha256:6ea2d006dcc9a3331f889bf59a677993dd5c4b9d73c5a11cd3f6eee8df5c4669

Observation 19542d3e-4b13-4b8f-a200-85ec055f4c5e · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:05e2878a6f9fccc1538ae1dc3fd3e1c7b76ffbad8390bc9e1ee9cafee7c98f6f

Observation 28557078-8ceb-4463-bd2b-e378c8b9bf24 · inbound

Faithful Mobile GUI Agents with Guided Advantage Estimator cites this paper.

Faithful Mobile GUI Agents with Guided Advantage Estimator UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T15:17:05.477466Z digest=sha256:87d14b0c34af28e328f99ab1698e5fe45a5045e964d0c5171f9bc95e63ab1a9c

Observation 274bd553-e222-4cc3-b0f2-3d7c3119970b · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:9544c0a64bcfc36f853eca172f546ca2f280e21488f70a7524d432cde3c596ca

Observation 622f0497-0040-41a1-8772-6d8f62d32762 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:2a845e1b48e8dffde1f0117180f4831b287841cda5d5003856abe21267e90d3a

Observation c4adde78-d67f-45b1-a84f-34f959825199 · inbound

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability cites this paper.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:15:19.239355Z digest=sha256:0f25c20a00bba469e9aac627a6f1cb1bc28aa235ecd9da50a822867865f48d26

Observation 689b7449-db70-4b29-abb6-4629f00329df · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:00:01.549385Z digest=sha256:935d2c6788d63d102e6b78de6303f05075c4c45cb410af5d9f4de619229e2065

Observation 26176fcb-ca42-42e7-a80a-b631adf5a1b6 · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T00:58:25.205386Z digest=sha256:d9295a17648e73c4a85b93186f00f3bc5609ff0f82261cb43b3100e853ea116c

Observation 776e3726-0ad3-41da-8ae9-acf41f2fbc81 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:25:24.844859Z digest=sha256:64d7ef7c4fa50a34d899bbd18044a7a6b7abf48042a200c5cb12d3452ea5025e

Observation f3a1372f-0a50-43d6-9f27-512be323ea65 · inbound

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse cites this paper.

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T06:44:28.552513Z digest=sha256:ee7508ebbb9b59fd7849e60a3ca1202345edd861dc29615a94767ef36dec04d5

Observation 64dcf418-87f0-4cbe-b885-64871c049ceb · inbound

How Mobile World Model Guides GUI Agents? cites this paper.

How Mobile World Model Guides GUI Agents? UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:28:34.562344Z digest=sha256:b01cbb9154a08faa37c3cb8742c07f022484198a12aa760df0b943518e5db660

Observation bb1fa078-1b11-4f65-86f6-2df62fbdbb15 · inbound

How Mobile World Model Guides GUI Agents? cites this paper.

How Mobile World Model Guides GUI Agents? UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:05:27.283138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:00:34.560405Z digest=sha256:ebad499141e84e63a40078a3c15b81dfff9dbeac661b8d8e647e90893dcd6523

Observation 1b07ec9d-ecdd-4142-a904-4399555431db · inbound

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction cites this paper.

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T02:12:42.449883Z digest=sha256:647324d0c7af6035e41a7548e7ad2fb960ecdf31bfc8fb8d9afcd55ef39e3d5a

Observation e4074840-7a5e-41ed-b3e3-b55e3f173dec · inbound

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction cites this paper.

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:57:59.129990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:56:56.216287Z digest=sha256:ec5597728dec500b2dca00c81492ad5482ce0a2d0aac13d0cd03c2082f1c0b68

Observation cbb15daf-f61a-4bc9-9529-53a8572cd99b · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:1d69fe969d966fa476d46c4e433f2c063e55a88cfb1c35a2767928b352d64286

Observation fa28f032-804b-43a1-8fe3-277d565b71ee · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:215dee1052b799e5f982aed8e529e92d2a5aafaff466b0367eb34a2dab5e06ea

Observation c64de1d5-80ae-44a1-a6d5-82e1bc2feb51 · inbound

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark cites this paper.

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.321363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:18:23.310925Z digest=sha256:ce5ec0bc9b5986490c614b55b992e35d88280f1c6bd2eac7c326d6c8efafebe3

Observation 430a40b0-0bcd-4b6b-830b-a34a521b90d5 · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:58:28.795280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T01:58:19.295247Z digest=sha256:689870f54ccbfaf526acdcc430a0c13c5e357c6bdf3151cdbf1b3d95375345f0

Observation fb8f7c68-e46d-4557-ba54-0414c7fbbde8 · inbound

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment cites this paper.

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:47:40.395327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T16:45:16.963802Z digest=sha256:ab9dc21a0c2583a827fc16dee42b9b6c750d5fbd04225d956c67bcb07594ebe8

Observation e3bc0aec-27b4-4b2f-9c0f-42bb85553c82 · inbound

WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections cites this paper.

WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:45:49.512095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T20:16:13.413064Z digest=sha256:b3b0ce3674a8299dfdca8432e36dd7664c2e55185e92bab891ba419f38bc73ad

Observation a7e987ca-5414-4fc1-8154-2366ddbe870c · inbound

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments cites this paper.

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:19:07.600603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T22:17:22.560632Z digest=sha256:52e6805a296445a2af4f5e206e03810fe86d0dc526dca66fc52a4ee578bf2f8c

Observation 7a7ce9c3-5de4-4de5-a7e0-c99cb9ead3fa · inbound

SE-GA: Memory-Augmented Self-Evolution for GUI Agents cites this paper.

SE-GA: Memory-Augmented Self-Evolution for GUI Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:42:46.390420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T20:38:37.264654Z digest=sha256:335b679273b715e1eee6caa3e6bafcf2ef9cbe8886cdb4d925e058394ae5d4b1

Observation 2538ddc5-c0bd-4865-8ea8-b9e11477b1c0 · inbound

Code as Agent Harness cites this paper.

Code as Agent Harness UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 226

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:14.520774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:54:54.558241Z digest=sha256:388fb8ea3f3a6926fb17e8ce1b605f4ee2663b10802483c534b3f5f78a5b19a4

Observation b7d9a004-dcaa-48ac-93e6-a0b34ee58850 · inbound

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay cites this paper.

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:42:59.307518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T02:42:24.746442Z digest=sha256:c0c94e39c9f02512f33d41128ac8952d2b5340b98e4f8d1ab20a9f026454fe26

Observation 528f3253-1bbd-4475-b670-7e364a09efa8 · inbound

CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision cites this paper.

CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:18:05.543191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T06:14:43.288597Z digest=sha256:9588a10d59627a4195f5d8d90906821c3f1cf9c4f46a268bd3ecdf2bbcd8411d

Observation 68497067-71c7-424c-9026-714cd0cb1ecf · inbound

PANDO: Efficient Multimodal AI Agents via Online Skill Distillation cites this paper.

PANDO: Efficient Multimodal AI Agents via Online Skill Distillation UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:04:38.748410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T12:02:31.227405Z digest=sha256:dc6080e171cfcf13b3b2336d45458ed7106e0e5581acd52a004347a341f9c271

Observation ce3ba494-d484-4fee-88ff-f4734119a684 · inbound

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents cites this paper.

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:43:59.220893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T21:40:36.033696Z digest=sha256:c7945b5629125f1a50c4a696b5382088453d53dc6409a01d50a84d93b7e4e9fa

Observation 1f334b2d-df65-4715-a031-ddf6748762e2 · inbound

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration cites this paper.

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:33:50.068791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:33:42.199405Z digest=sha256:190b640a59b5fb27645b71b65c2dff7fad7702991c543b29f949685b3d16ec35

Observation ae8ed666-5841-47ba-a8c6-d33d333b113f · inbound

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems cites this paper.

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:33:24.884989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:25:26.177071Z digest=sha256:07c628f23a42e93c4fab913aa96d3871af812d46694b483dd67c7f807b86fb74

Observation 4dd3407a-b8e7-4496-be85-87767697246d · inbound

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems cites this paper.

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:48.617757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:48.617757Z digest=sha256:2d4fddcc0c33c804c6940f46421ee759c9b3ea8f171bcb87565f9aa37da283a4

Observation 80d68cab-5d84-4325-be18-4c6757b14be5 · inbound

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents cites this paper.

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:23:30.866947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T14:15:55.180284Z digest=sha256:a91d5bf3de4f79b8ecda709ed21862c407ade6b4463a22910ec6a424e5449ff1

Observation 187d3462-8be0-4600-ac16-127bbdd1bff8 · inbound

PhoneWorld: Scaling Phone-Use Agent Environments cites this paper.

PhoneWorld: Scaling Phone-Use Agent Environments UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:43:13.439917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:41:20.001845Z digest=sha256:a59784b51d13e712ad75efb63b50002349461235f8cd3acf602a3fd1c8045ac5

Observation a834332c-8069-405b-9012-3c03c1f089ba · inbound

What to Format and How: A Benchmark and Workflow Approach for Document Formatting cites this paper.

What to Format and How: A Benchmark and Workflow Approach for Document Formatting UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:26:22.665054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:21:53.716738Z digest=sha256:0bf40654d025f58ef4f0cc5e79ab25f4c19d500411d69994fa58fc0c2e1d9279

Observation a813d14d-4c60-4e1e-8b10-d9d48d09dc68 · inbound

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents cites this paper.

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.448597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:46:50.587684Z digest=sha256:f624f1859536d96366fa4425c57ece9287065f608fa86c910b43e40611384a96

Observation 207b66e6-4f6d-4c6f-a11a-d052bbde2d00 · inbound

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials cites this paper.

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:28.739173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:33:26.954959Z digest=sha256:8d913f609de7f3b6d2dbd6f0c8f2b852813bdc2d552e408f4270aaeaace5fb47

Observation 8be8acd6-2d8b-4747-bf27-bdfbdde802f6 · inbound

AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents cites this paper.

AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.509350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:b262b2f895b7dea613ea8866dcc45036365e89eddfccbd3ed01768ba1a76d3d2

Observation 2b7e993e-1c70-442c-a279-94e2834f144f · inbound

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints cites this paper.

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:51:29.260143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:a08f2924f3638b0531f03d8c1ff5174f50be020f8b5ee227b09e0ae3a3849361

Observation b164bc1c-f59b-4dec-a95d-359f19b60087 · inbound

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints cites this paper.

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T15:06:31.394162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:06:31.394162Z digest=sha256:3ad8d30a13073a76bd7fdf388905374fe2e094ca111c2fd3f67fbccfcfd3d292

Observation 8dbee079-90b8-42c5-9699-dfde25b82da4 · inbound

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions cites this paper.

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:41:29.394778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T01:37:45.812577Z digest=sha256:ded413a39a3fc66edfee68adc3f83c9660b2cba93104cfbfe3e5c07b201380ee

Observation 61b87412-0c23-4e34-ba98-37a358088fd5 · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:47:09.711345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:8c1639c4d9f5f39d81180f6bac600bfead183da3d90c214bef17c5838c336d19

Observation b7733ab8-6c72-41d3-afda-9fcf7eb762aa · inbound

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning cites this paper.

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:47:31.132295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:22:05.424293Z digest=sha256:b4b8079131349d2bdc0b7af4a96d9cdffaf08a28326d2ee9b23cf9a1f93147bb

Observation 3b536508-5ced-4349-901c-1d1cfa4170e4 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.662721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:1ac7d02530506ccbe90c8cf1dadfd079e63b2b1a02bea6711436c0dedded7f60

Observation 4b4a3584-ce2e-4e57-af1c-679cc85e6fd8 · inbound

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields cites this paper.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.654126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:ec1225e97f2e51f74da41f8ec2f88aa827621fbb6ec53fc61128d6186072ddce

Observation edde6dd7-3a4d-4390-bcbc-b06f366838ab · inbound

Beyond the GUI Paradigm: Do Mobile Agents Need the Phone Screen? cites this paper.

Beyond the GUI Paradigm: Do Mobile Agents Need the Phone Screen? UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T21:38:58.305205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T00:23:37.068764Z digest=sha256:4f20cf797c76e851a0a7280f7ab801b3a8582881a26d9a2992395f232dfbe499

Observation f42e3eb9-715c-4ca7-a37f-780200ad97ba · inbound

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks cites this paper.

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:49:38.859596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:05:28.423719Z digest=sha256:940b3e82ce8c31d4d3f932dbd16474ee8c78790b2791b0a5d27b35f9b6d9f6f8

Observation 8b64f3ab-5de9-4220-883e-507c9005d6c8 · inbound

PhoneBuddy: Training Open Models for Agentic Phone Use cites this paper.

PhoneBuddy: Training Open Models for Agentic Phone Use UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:59:46.445801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T08:15:49.428124Z digest=sha256:b078af8cec445ffdc4dd25ee1453dceb6cc20e4a88e0015c6fcb5e12b6442c9d

Observation 12934652-4122-4f44-b0b1-fe64b0bcd911 · inbound

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes cites this paper.

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T01:44:09.321015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:539f13c99b3ba55baad2b7c94ecc6248d1d210e7422223d0e19572d0a05a8110

Observation 4a0bb389-216f-4111-b4a6-2e1f9a323639 · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:04:21.282341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:33a7fca3519b6e7d2eb271cf76e852911149d9b8d075f1e2fb4d05c1989c9dde

Observation 42d9e712-ea09-4402-bf03-080fd076cd74 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:21.222921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:a93e63acbe64af01c78f08fd886adf0c1762a0ddb33993ffcf7258830891f0d5

Observation dacf9851-5ddc-4f1e-938f-8d69895c8b3b · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:ee0002ee795311a801a417e70656073b34097cefbd34acb6813e09738e7023d7

Observation 06b788bd-9d91-427d-8e49-dff4d90a12af · inbound

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots cites this paper.

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:44:19.210310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:39:12.591090Z digest=sha256:28117330d37ebf864cd818d565d26bc843f5a164682378354cadb42c68b6930d

Observation bfa6bb6d-0d89-4410-8ef1-ccbada18e4f0 · inbound

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding cites this paper.

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:04:21.757698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T06:58:01.734975Z digest=sha256:46b6688ea9b2420ce09762ef2411462da25b1277716acac846ef53f980d170b3

Observation 6f5dbbed-4510-4e5f-9b7c-f65b47514306 · inbound

Xiaomi-GUI-0 Technical Report cites this paper.

Xiaomi-GUI-0 Technical Report UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:55:40.316053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T06:07:48.137351Z digest=sha256:b0b7b34f0e6ad2bd543e9e824b62db290636fcd2378d203eb7f5dbf38b20cdb9

Observation 662f5640-c0f2-4a86-bb35-3ed66badfccd · inbound

Xiaomi-GUI-0 Technical Report cites this paper.

Xiaomi-GUI-0 Technical Report UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:18.996918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T19:37:49.661596Z digest=sha256:83cc7176461da78f6911d3839f7f67af51ec35cde671d93ed3d8c723b5e98ca4

Observation 698230d0-796e-4c86-8307-9caf40dc4127 · inbound

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States cites this paper.

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:25:42.425459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T05:26:53.441833Z digest=sha256:6f23bef07ad66ab972501f3a87bf6c6bc9deb8c4c51fd0094d3770df4b9403b4

Observation 2215a800-d821-457c-b693-f2b28e98b87d · inbound

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States cites this paper.

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T22:08:58.620596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T22:03:26.625209Z digest=sha256:945474ca56375aa2069747e3d1006f288d72b4708afa9725c78be49743e5c58d

Observation 3d42dae9-7f89-4a89-8db2-12c2e3b14198 · inbound

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning cites this paper.

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:37.302399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:37.302399Z digest=sha256:36ff5be5048f61d297467631484e186ffeb7da75f3e25bed35a7016d451ad5e8

Observation b74c6ba6-0d27-49b2-b84d-a9b00bd6fbf5 · inbound

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL cites this paper.

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T06:25:23.264527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:25:23.264527Z digest=sha256:fb27e3694fb03b00e9568a8e02b449f9046322c11de6474a6916bb0e01e0f18c

Observation ea7e4970-021d-4755-9be3-c5cce40bb408 · inbound

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents cites this paper.

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T23:58:28.708772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:58:28.708772Z digest=sha256:ab43986210b23c60bd14927fa97d513844a5aeb2cf815ab6d26d7fd4b36708fa

Observation d20070ad-eb74-4861-96d7-f33af5dcc337 · inbound

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents cites this paper.

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-31T14:04:44.771905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:04:44.771905Z digest=sha256:be4d87dfff9248bf3c5ced59f197ef043346be9cde626819ef31e78ca3033073