Pith. sign in

Paper Citation Record · LEDGER

Benchmarking LLM Judges for Mobile Agent Evaluation

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.11434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11434 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:32.198253Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 459565ae-63e7-4dd6-a812-ad3413836f2f · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.911176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.911176Z digest=sha256:7a9c4f8580050eea9bc54b7d050733197f9f8056ed32b40cb245556b85d18b23

Observation 4c37776f-8270-447b-8722-155012e39f7b · outbound

This paper cites A Survey on LLM-as-a-Judge.

Benchmarking LLM Judges for Mobile Agent Evaluation A Survey on LLM-as-a-Judge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.917445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.917445Z digest=sha256:804dd56bd155eaa78ef51a8501176b46e7235506292496b26e7e48ae791e29e7

Observation 859ba63b-7930-4f89-a197-ccde841f800a · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Benchmarking LLM Judges for Mobile Agent Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.923269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.923269Z digest=sha256:de9719dbc67ad67c4d0dbc2ebd6ab8c9daefcf7871b8821cf1391fcf2361d88b

Observation 22ce04ac-f32a-4eea-8542-f17a247c9d6e · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation Agent-as-a-Judge: Evaluate Agents with Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.928785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.928785Z digest=sha256:ed0c5176c640e4fc9dc24045564ca4e80093a0bea48fe4f4e77fdb584211f74c

Observation 4159f734-2f57-418f-9a2e-69c61e1657ef · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Benchmarking LLM Judges for Mobile Agent Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.934258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.934258Z digest=sha256:e0e6e185f9c6c83ed81d7e83db2af0c8e68b1ca5e3f6e6f09a28b9546106dc15

Observation d1fa04e0-638a-4413-8c14-96337f107169 · outbound

This paper cites arXiv preprint arXiv:2502.01534 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2502.01534 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.939821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.939821Z digest=sha256:501488db90077abaa96ab2cb5e832a2e6d73fbb0ac46bb5045de8a6f0f8c275a

Observation 59a8ad62-7e33-48fa-b8b8-bdeb29d4b960 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Benchmarking LLM Judges for Mobile Agent Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.945269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.945269Z digest=sha256:d1be618f660bb6ffaa7b69952a9bcf09d0ccd567f2c81b9070314d7039007524

Observation 04cf6c93-592c-4d2a-afaa-6fbdbd114091 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Benchmarking LLM Judges for Mobile Agent Evaluation LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.951140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.951140Z digest=sha256:79521632944e8792c0989404eebd62b112041ce28493a1506ced747a1f2b1ec4

Observation 53e2ec39-3c31-494c-baaa-6429e480fa76 · outbound

This paper cites arXiv preprint arXiv:2504.08942 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.08942 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.956961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.956961Z digest=sha256:ac7a213247e9d057018d240d67f21732a20f53b4050ea3f3d3719c34d6d275f9

Observation 032d1b18-e46a-41d3-90a3-a4e053a48ce1 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation Autonomous Evaluation and Refinement of Digital Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.962001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.962001Z digest=sha256:c8a23214bb0fdc610dc37f94404577931d3796b82683f10e7ae2ddef80074c19

Observation 098aef06-4c47-4b4e-82b4-11f191ecb904 · outbound

This paper cites arXiv preprint arXiv:2503.02403 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2503.02403 , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.967653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.967653Z digest=sha256:fc83151b3fb998ea61bdef237204c9196b2e9535bf26161f5de0f193af2aee8a

Observation 416e280a-2822-4573-b6ef-49d966f55767 · outbound

This paper cites arXiv preprint arXiv:2504.01382 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.01382 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.972605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.972605Z digest=sha256:3e013c7c27f32162e25e90e361acc19401ebdd39a41740e1be90b9650aeaebf6

Observation 785bddc2-7648-4494-be09-4673ab7c0821 · outbound

This paper cites Agentic Reward Modeling: Verifying.

Benchmarking LLM Judges for Mobile Agent Evaluation Agentic Reward Modeling: Verifying

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.355751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.978148Z digest=sha256:2bad3480fc966ba7c010df4bea61ef02765cc03a1c4507b3419307f126d2f867

Observation 8963812d-7c3d-4d88-8712-ca0ae5ca3e97 · outbound

This paper cites ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration.

Benchmarking LLM Judges for Mobile Agent Evaluation ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.983082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.983082Z digest=sha256:e1f5439698c49e472b0623cf6c0e834e9ca5618790b8e90b3d12f97169fce80e

Observation 84019af1-727f-40e6-a9ce-30c178af32cc · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.988574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.988574Z digest=sha256:4a3a403de71319bf5932d778d1a380e352796a8210f92bf397121734c1354298

Observation 51f932d1-9049-4148-b46e-9b8eb2ec1d99 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.326938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.993527Z digest=sha256:8820b89c758786f8e7323798b5ae1f36180b17c1bcce8f4b1ad5763701c8550a

Observation f2883150-765b-44d4-bbb6-7f7d9e876b27 · outbound

This paper cites 2025 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.310381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.998716Z digest=sha256:137ae1cdef94f4dfad920602bc6e6ddbee505b32c5d7287fedd7fad64ded885f

Observation 8ce83736-aed8-46d5-b7aa-a39987ff5fee · outbound

This paper cites Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.295066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.003906Z digest=sha256:69ad0694e67b178e57239cf0ac8b851112dc1bc822ba773091c68d0c13051fb1

Observation c76913cf-4f77-4859-8e4d-d017b199c749 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.008836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.008836Z digest=sha256:3c26bfe4b3ab9b12b1ce7f5123126e1fac7dd580aff1ddba3d23577c960855cb

Observation 20a63972-cfe0-4996-ad35-5a3959ff5151 · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Benchmarking LLM Judges for Mobile Agent Evaluation Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.013673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.013673Z digest=sha256:d770bc966d863654ae885e38bd69f29acbd19f862f7130d213db5d22f2ab25f9

Observation 0734db5b-9877-426b-a87b-8099740f23e8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.018875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.018875Z digest=sha256:5c5c71b3e17aee04a7b5d09ec26ce9f0e63ec85799dcc728ba3c46820d0182df

Observation 856bbb72-5629-41d5-9e79-8e999dc5a31e · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Twelfth International Conference on Learning Representations , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.023625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.023625Z digest=sha256:50d12f39a3a87f48405c7d6be133e21cccf2e97ea644d1affa3bab61d800375e

Observation 40982bc3-e1a1-48a3-a05a-b807dd17939d · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.028406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.028406Z digest=sha256:e2927d160f1ef1c723ec654c53a37c9d56cd5b78f77529a32fecc979b97309d7

Observation c06a0372-89c0-4a33-8383-574f9f29d717 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.033137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.033137Z digest=sha256:3eff67909df49dc7f650c9864828d562b22786eda00c87b39e7b734afdff657d

Observation 3edaa242-13ea-49b3-ba36-5ac3246d2f18 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.038168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.038168Z digest=sha256:2f8cbc94ac74892745adad00f058daf88e0942a79f75d9b5e5bb728f838a647e

Observation e6b39dcc-4d15-4b58-b11c-f377d615c79f · outbound

This paper cites AgentBench: Evaluating.

Benchmarking LLM Judges for Mobile Agent Evaluation AgentBench: Evaluating

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.042790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.042790Z digest=sha256:e6f40e7a7b504b7e998390ea38ac57d527905f2dd616d41e8ed0cd2b47c77af3

Observation 44b90881-c4fa-42fd-9c9a-737df97b0901 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.047694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.047694Z digest=sha256:2e97a38950bcb9fcfe721a8021521b0d14e6533bf85053b257ed9bb9407cef25

Observation 174e94ea-8eff-427a-8fd1-b262c726d22e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.052364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.052364Z digest=sha256:f313f7072b50b0b896ecba7bbae9f7957bd12f2abce022507f6e0eec3aef6e9c

Observation 0d672b7b-1725-44eb-bf82-ced1257f74e2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.057692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.057692Z digest=sha256:b65918baa9c6d8fd054961ca7f244980865a41d772cd52a54a49410f801d70d4

Observation 172e35b6-9782-4df1-82a2-1f2450f4ab0a · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.062580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.062580Z digest=sha256:16edb6ea418c73461db72726931336b93afdde532d6adec9edca0687b76abf04

Observation 6cbe5f76-0556-41be-a8aa-af24c2fe83a4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Benchmarking LLM Judges for Mobile Agent Evaluation Constitutional AI: Harmlessness from AI Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.067446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.067446Z digest=sha256:9b05b8525102a78f0fd620805281f4466fadab72ebac6349af70ba8d86933758

Observation e9bd76ce-9bd4-40ed-b964-3c170892c2a6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.073219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.073219Z digest=sha256:890558d8f155cc2eb066a9b60327810765db9dc76b56f533b5e7e52de3fddbdf

Observation 03355267-5ae0-4dfa-b3dc-df89f733b8ed · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.079033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.079033Z digest=sha256:58c5d12f80c68f7e417e61ac8c0aa2929d881ffc3814fa65138153916d198832

Observation 296ab333-57e0-4942-948f-8679fa6474f2 · outbound

This paper cites arXiv preprint arXiv:2509.18119 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2509.18119 , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.083855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.083855Z digest=sha256:f0dddf88d40589ec13002ee09a0c08b464caa480827a3f03a928470324382a64

Observation 0f56b061-3d55-465d-a1bf-fc66b6dc4e69 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Benchmarking LLM Judges for Mobile Agent Evaluation MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.088344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.088344Z digest=sha256:5d4b861567d0b6340c7f068b8061aa5ab5a4020aa927835f884aebf551bb1755

Observation 65c1156f-ba41-4315-a70f-56cb60bbc993 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.093292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.093292Z digest=sha256:492a37c5856dfe23376b79a9cb63211679ec2e9e032d5067e1fa93370eb62e71

Observation 77aeb3c1-2a88-4476-9888-5e1845537f1c · outbound

This paper cites International Conference on Machine Learning , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation International Conference on Machine Learning , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.098421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.098421Z digest=sha256:10f3a262a33ccf8c46499db49676572edf0d3f73732e856563b21b70c8fb2333

Observation fdc80167-d9ce-44c6-bdd7-9fc0bea76df0 · outbound

This paper cites arXiv preprint arXiv:2409.15922 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2409.15922 , year=

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:16:32.473314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.104134Z digest=sha256:c5611826cf5ac8bdf908c1764cc80649ca2fb85c5049a10d1e2cc34ce69a31b1

Observation 596d336c-524c-4359-956c-5136dfa0cb19 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Eleventh International Conference on Learning Representations , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.109139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.109139Z digest=sha256:caca1ed8577cc81d05c7e7a323518cacfd79efb87d57c75fcf1bc193b47bdd4b

Observation 8ab4694e-adef-4b5b-a7e2-ad6578e1705d · outbound

This paper cites 2023 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.113664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.113664Z digest=sha256:935a0dccc5d6fe423e40b7e7311eb25439fe2bde83e7910e36c83eb2fa01bbfd

Observation 42469195-bcff-40ab-b82e-fcd6fd95bd22 · outbound

This paper cites 2023 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.120720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.120720Z digest=sha256:99bc23e8e0fd428294c21aa3204c0d08ac2e70cde32f5673749317957be6281d

Observation 467705c5-13bc-4cf1-bce6-df9e156c3a3c · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.125755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.125755Z digest=sha256:ef5af93c8d10931a3ece43827f520255caf9b61b3ee29380860e86e4002b263b

Observation da50c88c-d8e9-4378-ae52-fba1444e3721 · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.130803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.130803Z digest=sha256:8409f6afc7f64062adeb7b3a3aeba5297f761533603c937efdefb8a785dc86c4

Observation 813b31c9-9785-484a-896d-4cc80db9f165 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Benchmarking LLM Judges for Mobile Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.135869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.135869Z digest=sha256:2154d2afb2764460e4574d5a9b6b8e7965a0079eac6105e18520569d632260e0

Observation db4c014c-b9a1-4712-8858-0f5ad09dbf9d · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

Benchmarking LLM Judges for Mobile Agent Evaluation NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.141202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.141202Z digest=sha256:93aa0cee2c67950633844ba54d76ec3c22e7d3690422afbc6909d94cde98166e

Observation c169458d-480d-4331-b420-f2e2239d0b6c · outbound

This paper cites 2025 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.146221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.146221Z digest=sha256:4d5c70632aa0b787421738336df192ca8cf34865330c743c813fb0927a00006d

Observation 65c21e3b-1e6f-475c-a3cf-39d8ca993a1e · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Benchmarking LLM Judges for Mobile Agent Evaluation ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.151042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.151042Z digest=sha256:99d997f065d8b1e0ecb936f66edd5289a5736263e7a1599d8c6056f5c94c3e9a

Observation c2ca3d5b-87f6-4e61-9d75-a6c634fbc8c9 · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking LLM Judges for Mobile Agent Evaluation The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.156301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.156301Z digest=sha256:0e42b84a6270de3cb7aef567645aba2b1eadc168c4f6ce5ea396a1abd13bccd3

Observation 9d0df293-bc95-447d-9e23-29eba29c88e9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Benchmarking LLM Judges for Mobile Agent Evaluation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.161977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.161977Z digest=sha256:89088b46fe01b128c06dbeee6ec22086539645d1c6efbb019071f72256da468c

Observation 2fa48a6d-d961-44df-b778-5fc06d7491e0 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.063532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.167383Z digest=sha256:416bfc72bfaf8d2439785c47aa35356aafadc3988a3577d36e4f0e3c486e6865

Observation ecc1b69f-87eb-4487-9980-26640e937fbd · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.046366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.173421Z digest=sha256:e8173bdfe5a5703a28d57e33ed07e8c6571f1ce97f968f1b22a5a1fb5e0c18de

Observation f791c250-faf3-4e16-b380-307aca291450 · outbound

This paper cites 2025 , note=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.027992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.178476Z digest=sha256:0b672988a22abe941710e43f57b008c621fc8059bb4995ec23e5b56c01b8239b

Observation eb6527f1-0fa1-40b0-8443-0b67c865ece4 · outbound

This paper cites 2025 , note=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.183383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.183383Z digest=sha256:da43a310e4ad1aa8028ec66d8f8ef891cb338f8972db63bbbf6fe45b95013e99

Observation 248cf1fe-a7ce-41c4-a806-0b0cc83d4ad2 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Benchmarking LLM Judges for Mobile Agent Evaluation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.188278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.188278Z digest=sha256:436ccc2fbb0c74dfd6b214ead983dc5a73166d0a9978336a5c94328b5d0e9397

Observation 98d9f3cb-4d8f-45a5-83ab-d70e1bdd8a69 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.001354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.193184Z digest=sha256:8024dd15d31f96c7a7fa6a38f58433abd6c22a2206471a5592c495c117614c19

Observation 347f05eb-0442-49f7-9010-a9d9604287db · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:32.983736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.198253Z digest=sha256:d12ae9c5566fe7bd7bd29ec4e63128eb1846091d4b70141eb8745548a2bfaf86

Pith citing papers

No inbound Pith citation observations are available.