Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.896432Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2607.14989.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.896432Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:38:16.681296Z
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee70646b-4454-4e07-8009-a1e1c2d75d7b · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98ecf2b2-2b0c-4c1a-8cae-d008ec31a237 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d89396-9c19-42f5-9f52-506cad14fab0 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems, 37:5996–6051, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e1de46-1928-4ce4-ba49-12da0a77ca13 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Acebench: Who wins the match point in tool usage?arXiv preprint arXiv:2501.12851, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2222a21-b736-4cfe-8275-ce8d485c1f75 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ac0475-4325-4b21-8e75-df4f7b354e95 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f80d6b94-f83a-4286-9116-df602d194120 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfadd12e-87fa-42be-8641-d1fbb1574dde · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Gaia2: Benchmarking llm agents on dynamic and asynchronous environments.arXiv preprint arXiv:2602.11964, 2026
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c345dfa9-7688-4b25-b7ca-c55a6fe11a38 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios GLM-5: from Vibe Coding to Agentic Engineering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ad8539-97f9-47ff-874d-8edd8998011c · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.arXiv preprint arXiv:2509.26490, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d532e613-26c8-4670-99cf-ee37c8d0485f · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de01f234-0ee2-40bb-a02f-ae99aa420a8e · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios The tool decathlon: Bench- marking language agents for diverse, realistic, and long-horizon task execution.arXiv preprint arXiv:2510.25726, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae5b7737-562e-4b11-abc5-8bf9457f9a66 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcec0a53-67d5-4c3d-bcb9-1ae873084715 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d0ddb8-f4d3-4ac4-920f-1a718e7be0cb · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57e6884d-83e8-4a3e-aa7c-f3d8bac50212 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13722e1-b478-40ba-a0d6-96a425633d9a · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146c6ecc-0f79-4997-83da-47bd6ce570af · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d90dc87d-f0b7-4e55-b017-cfc8530cc58c · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios QwenClawBench: Real-user-distribution benchmark for openclaw agents, April
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d291883-7561-47ed-b00e-4e53a70b11b8 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a43c760-0979-454d-9dcb-148f7ee0153a · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios URLhttps://arxiv.org/abs/2603.04370
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55cf75c-035a-4b25-a140-5a352ebe7b0d · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Kimi K2.5: Visual Agentic Intelligence
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb12d6ba-b891-44d2-be04-6d568462a14a · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af412506-04f8-471c-97d8-06db166c7e8e · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380c0c35-124a-4edf-9d35-aae77bbfd9a5 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7303418-7fa0-46ef-9090-43c43d571a79 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe38867-3a50-4928-8cc1-323641177e9f · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Deepplanning: Bench- marking long-horizon agentic planning with verifiable constraints.arXiv preprint arXiv:2601.18137, 2026
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 532400e7-adb4-410f-acfd-6b22622c9a7c · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7099790e-86de-4bb8-89bc-b4636ce2f1e0 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72246456-5f0d-4556-9518-d2adf7b427e4 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67998d9a-8708-4d4e-b59d-12469d9c2195 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios This protocol evaluates not only task execution, but also clarification, constraint tracking, adaptation to user feedback, and state maintenance across multiple turns
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 714e253d-0a79-4272-8875-05d696ed5080 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77acbd66-6a2b-476f-bf2d-f0b21dcb3264 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef31ed00-4acd-458d-83e4-0b45f1df8868 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios 4.VerifyCodechecks the trajectory and final observation and returns a binary pass/fail result
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7af32ea-81ca-4c1b-bc37-92a81124aa3d · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios [...additional description omitted ...]
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd424a9-9c6d-47a2-8ee4-9f4dde5f68d8 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios [...additional tools omitted ...]
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c3a754-e51a-4626-9ef3-498b001c9a3b · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Your goal is to complete the user’s request in an interactive environment by gradually calling the available tools step by step, and to proactively communicate with the user
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85bd7733-9245-40e0-89a5-575d6b867983 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e457aacb-d23d-48fc-92db-2e1a7301f1e2 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios <Judge reason> The trajectory correctly resolvesLF-2024-PI-024, confirms Yunhe Foods / Lin Qiaoxue / He Shan, verifiesVER-004 as current, and flags the self-referentialLNK-003link
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b900804-95e7-4063-82de-3f3bce04cd52 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios store surveillance screenshots and incident timeline explanation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008c583d-9452-4fa5-936d-392a1db75110 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db42b9d-0575-4d7e-9996-7269cca3ffab · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd688285-357f-4f79-a60f-60a168ba8c39 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf144a3-b8d9-41c8-9c6e-b03b254a77ae · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2861883a-b8c8-47d8-b643-7ac88668d0c4 · outbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc7ea6a-b277-4acd-9cb3-21332b85d0d0 · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.