Pith. sign in

Paper Citation Record · LEDGER

DPO Learning with LLMs-Judge Signal for Computer Use Agents

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2506.03095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03095 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:48.561600Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84a62ef0-d64b-4099-8946-7d03bf5faaa2 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:46.803506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:46.803506Z digest=sha256:2a4c6f2cae39b7b141e7eeba9d5fd2f21a75c4ae1fcf011e7587dfbcc1fd318d

Observation ea1053d5-30ed-4248-9cb2-640ec29e577f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:46.892805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:46.892805Z digest=sha256:50e1652fca4a4244f1616d741788569461e973a4d27d0473ce7fc8ffb38cedf7

Observation 05cea6c6-4d04-40f8-9aa7-ee208e20f0b7 · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:46.979114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:46.979114Z digest=sha256:d345852dd5114f0bc4cf54c02f191d76c9112335a7c4ace180d9803342308fa0

Observation 7d4bee5d-5850-4a70-884f-3bf4e00f7f47 · outbound

This paper cites Deep reinforcement learn- ing from human preferences.Advances in neural information processing systems, 30, 2017.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Deep reinforcement learn- ing from human preferences.Advances in neural information processing systems, 30, 2017

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:51.007382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.046229Z digest=sha256:b4aafedeafa802d76c5838b7c8465a5e5906cbc982653f09fb97f70eafe5361b

Observation 658122ca-daf3-4bc7-a0ad-804fd2eb11c1 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Mind2web: Towards a generalist agent for the web

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:50.797671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.127378Z digest=sha256:c74120eb37554d72a3f2f2db59f6968f3018e83a4802a0f9d188551a04bce396

Observation e9c7f30d-34c8-4991-8b0f-56f0c395b6bc · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Detecting and preventing hallucinations in large vision language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:50.533904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.230549Z digest=sha256:23f900b3742c5d027912f236dbb726bb1651bc4fb283b3b937b8d546a93a95ac

Observation 0a2d819d-b698-43dc-bfaa-8625d3a61fd5 · outbound

This paper cites From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge.

DPO Learning with LLMs-Judge Signal for Computer Use Agents From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.315267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.315267Z digest=sha256:e3985a5b3585f4a37ed1f4437ff0f4171442c2e47c9ada66b76cc03a57cc144e

Observation 2851502b-aa0c-47f8-bb4b-5bbecad3798b · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Silkie: Preference Distillation for Large Visual Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.407418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.407418Z digest=sha256:6175ca2b7236199d63fff31b8332e7f885a9d16f2ffbe72115e318acee6c0f9e

Observation 6492f508-04b4-4ca8-9b5f-0aecedd78d40 · outbound

This paper cites Visual instruction tuning.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Visual instruction tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.487075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.487075Z digest=sha256:7fe860b04e47407caeb4cbdac1b5773b960d5ace7475b27dd775d87088a49858

Observation 3d6fae3a-0208-4387-b14c-424fea65d435 · outbound

This paper cites InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection.

DPO Learning with LLMs-Judge Signal for Computer Use Agents InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.579318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.579318Z digest=sha256:d8f13243c8a518103b4429bff1b4daf3d44e50b1c0d6b1a95a8b7bacb4525b3c

Observation b84ccf14-510c-4c8b-894c-088bc0000284 · outbound

This paper cites Hello gpt-4o.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Hello gpt-4o

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:50.238584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.641146Z digest=sha256:ebf2e2e989f190e00c19b5885cb9dd1bceb14f47f8a40375ab9b18701892c0df

Observation d39b2950-ae32-43f9-99d1-9a49b6d4d6cd · outbound

This paper cites Training language models to follow instructions with human feedback.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Training language models to follow instructions with human feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.985964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.702656Z digest=sha256:536cba52c10a4adf352dbdb3f7d468824e1cdef925b31bc1cf101ba85da50f80

Observation 4d16142b-d8d2-4b88-b24d-c4fd0f912613 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

DPO Learning with LLMs-Judge Signal for Computer Use Agents UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.763545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.763545Z digest=sha256:89682b082f5b85d966a7b8327362c4e20fbbe7b374b6af995af651eb97f62f5d

Observation bef87bf5-1916-4d8e-bc38-e68ae7ee8514 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.695754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.819248Z digest=sha256:f1e5bc0d02388b66cb0e4d7e73a382878dd96047c9bbb2aa97927391029f3bb6

Observation c0f62f9f-6d31-4201-bc4e-345989574692 · outbound

This paper cites Androidinthewild: A large- scale dataset for android device control.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Androidinthewild: A large- scale dataset for android device control

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.492838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.869310Z digest=sha256:2202ccece541e0b664d4de22897b7ad743cfe68bdecf133ebbfbb033513574b3

Observation 85780b80-96e0-4ffa-8e75-495e650e76ea · outbound

This paper cites Proximal Policy Optimization Algorithms.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.932567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.932567Z digest=sha256:4df683a96b85950ee2366afd3e5e846a89d0e81da398adb31ec1aff107147fd5

Observation 52fc45d8-d8ae-47c0-8146-e8dab32c6959 · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

DPO Learning with LLMs-Judge Signal for Computer Use Agents mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.992971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.992971Z digest=sha256:17aa2baa175fd73dc3a8e63e7aebd9b488ab6daaf0dd1644cd7afbd03a5619ca

Observation 13a7da60-6ed7-43cf-b502-223ebe2c8838 · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

DPO Learning with LLMs-Judge Signal for Computer Use Agents OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.072942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.072942Z digest=sha256:95d4afe6b14b6b363553bf75cb1b09d2b5a04dadd7c63d0b5d3634783c140d03

Observation 2c4b1f6a-fc90-4e32-b79d-28036744c2b8 · outbound

This paper cites Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.294060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:48.154432Z digest=sha256:069923a1b9c708dc9a238653c8afe642e30590a0e9b51ce204414fcf7f627c25

Observation b940e2a0-ab9d-487c-ae4c-5a991da57c58 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.195477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.195477Z digest=sha256:15cca242dfc8578c2d920c325cd54f9e8e0302e73d02a5ecb19227285c0901ab

Observation f3e53efa-5ea7-40c2-a989-a6f99ef929aa · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

DPO Learning with LLMs-Judge Signal for Computer Use Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.250484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.250484Z digest=sha256:5f90ef8d315c98c39aacad1557431bda8cd7d618bfc5ca2eda0a5291f91ac782

Observation a6d1d13c-c917-42b2-b910-fd47e5cd3e2d · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.321722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.321722Z digest=sha256:80155f0b591b6524519dae8ff0c34f8406839f13a4ee18f3b9a8a14437ad53be

Observation 2e782ce6-1f08-47c5-ab9b-cc56e749039c · outbound

This paper cites Gpt-4v (ision) is a generalist web agent, if grounded.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Gpt-4v (ision) is a generalist web agent, if grounded

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.083792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:48.389540Z digest=sha256:674250131305c127e5de513af095e20d569bdc001e3f10addc9f10af852cce4d

Observation 2a5c22b9-21d8-49cf-93b3-dda022dd0684 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:48.925875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:48.452167Z digest=sha256:626eb781bd2c07e402bc9381ccdfdadce5b7e2a0f1850cd1a53a4cb43cd7c579

Observation b83d7520-3411-4e9f-a70c-3042e5b5d92c · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

DPO Learning with LLMs-Judge Signal for Computer Use Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.502654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.502654Z digest=sha256:f5eedb97d8553bb9f3525075a72a2e0e561831cb9dae1230368d01146ee3e40e

Observation 552f0184-948a-4b4d-b4b4-5ce9554e9345 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Fine-Tuning Language Models from Human Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.561600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.561600Z digest=sha256:7b89ec2e9d5a9ba1acffc5fcf37f5296a317e527f2b296c217f7f941532a62ac

Pith citing papers

No inbound Pith citation observations are available.