Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:32.198253Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.11434.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:32.198253Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 459565ae-63e7-4dd6-a812-ad3413836f2f · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c37776f-8270-447b-8722-155012e39f7b · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation A Survey on LLM-as-a-Judge
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859ba63b-7930-4f89-a197-ccde841f800a · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ce04ac-f32a-4eea-8542-f17a247c9d6e · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Agent-as-a-Judge: Evaluate Agents with Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4159f734-2f57-418f-9a2e-69c61e1657ef · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1fa04e0-638a-4413-8c14-96337f107169 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2502.01534 , year=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a8ad62-7e33-48fa-b8b8-bdeb29d4b960 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04cf6c93-592c-4d2a-afaa-6fbdbd114091 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation LLM Critics Help Catch LLM Bugs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e2ec39-3c31-494c-baaa-6429e480fa76 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.08942 , year=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032d1b18-e46a-41d3-90a3-a4e053a48ce1 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Autonomous Evaluation and Refinement of Digital Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098aef06-4c47-4b4e-82b4-11f191ecb904 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2503.02403 , year=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 416e280a-2822-4573-b6ef-49d966f55767 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.01382 , year=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785bddc2-7648-4494-be09-4673ab7c0821 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Agentic Reward Modeling: Verifying
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8963812d-7c3d-4d88-8712-ca0ae5ca3e97 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84019af1-727f-40e6-a9ce-30c178af32cc · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f932d1-9049-4148-b46e-9b8eb2ec1d99 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f2883150-765b-44d4-bbb6-7f7d9e876b27 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8ce83736-aed8-46d5-b7aa-a39987ff5fee · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c76913cf-4f77-4859-8e4d-d017b199c749 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a63972-cfe0-4996-ad35-5a3959ff5151 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Benchmarking Mobile Device Control Agents across Diverse Configurations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0734db5b-9877-426b-a87b-8099740f23e8 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856bbb72-5629-41d5-9e79-8e999dc5a31e · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation The Twelfth International Conference on Learning Representations , year=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40982bc3-e1a1-48a3-a05a-b807dd17939d · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06a0372-89c0-4a33-8383-574f9f29d717 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3edaa242-13ea-49b3-ba36-5ac3246d2f18 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b39dcc-4d15-4b58-b11c-f377d615c79f · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation AgentBench: Evaluating
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44b90881-c4fa-42fd-9c9a-737df97b0901 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 174e94ea-8eff-427a-8fd1-b262c726d22e · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d672b7b-1725-44eb-bf82-ced1257f74e2 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172e35b6-9782-4df1-82a2-1f2450f4ab0a · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbe5f76-0556-41be-a8aa-af24c2fe83a4 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Constitutional AI: Harmlessness from AI Feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9bd76ce-9bd4-40ed-b964-3c170892c2a6 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03355267-5ae0-4dfa-b3dc-df89f733b8ed · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 296ab333-57e0-4942-948f-8679fa6474f2 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2509.18119 , year=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f56b061-3d55-465d-a1bf-fc66b6dc4e69 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65c1156f-ba41-4315-a70f-56cb60bbc993 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77aeb3c1-2a88-4476-9888-5e1845537f1c · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation International Conference on Machine Learning , pages=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc80167-d9ce-44c6-bdd7-9fc0bea76df0 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2409.15922 , year=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 596d336c-524c-4359-956c-5136dfa0cb19 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation The Eleventh International Conference on Learning Representations , year=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab4694e-adef-4b5b-a7e2-ad6578e1705d · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42469195-bcff-40ab-b82e-fcd6fd95bd22 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 467705c5-13bc-4cf1-bce6-df9e156c3a3c · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da50c88c-d8e9-4378-ae52-fba1444e3721 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 813b31c9-9785-484a-896d-4cc80db9f165 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4c014c-b9a1-4712-8858-0f5ad09dbf9d · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c169458d-480d-4331-b420-f2e2239d0b6c · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65c21e3b-1e6f-475c-a3cf-39d8ca993a1e · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ca3d5b-87f6-4e61-9d75-a6c634fbc8c9 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation The Llama 3 Herd of Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0df293-bc95-447d-9e23-29eba29c88e9 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa48a6d-d961-44df-b778-5fc06d7491e0 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ecc1b69f-87eb-4487-9980-26640e937fbd · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f791c250-faf3-4e16-b380-307aca291450 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eb6527f1-0fa1-40b0-8443-0b67c865ece4 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248cf1fe-a7ce-41c4-a806-0b0cc83d4ad2 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d9f3cb-4d8f-45a5-83ab-d70e1bdd8a69 · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 347f05eb-0442-49f7-9010-a9d9604287db · outbound
Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.