Pith. sign in

Paper Citation Record · LEDGER

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

As of 6 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2607.14989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14989 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.896432Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:38:16.681296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · outbound

This paper cites Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.686510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.686510Z digest=sha256:7ea321622a51556f417c3aa23c60b8fcbc9e7af89a4a20f3afd72bb5ffedb612

Observation ee70646b-4454-4e07-8009-a1e1c2d75d7b · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.692389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.692389Z digest=sha256:813496d0f38d91d0550f996a4a3789b475c97171d580ab05b5aef4e28f90e0bc

Observation 98ecf2b2-2b0c-4c1a-8cae-d008ec31a237 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.697891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.697891Z digest=sha256:a896743c124346c4d02574b8c1dd63658b227f54871256fbdd050f40d1914999

Observation 69d89396-9c19-42f5-9f52-506cad14fab0 · outbound

This paper cites Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems, 37:5996–6051, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems, 37:5996–6051, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.703133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.703133Z digest=sha256:6b445e6cf4529ffcb6f4017c2f4f106092baf266d004faa95aa829ebf6609511

Observation 92e1de46-1928-4ce4-ba49-12da0a77ca13 · outbound

This paper cites Acebench: Who wins the match point in tool usage?arXiv preprint arXiv:2501.12851, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Acebench: Who wins the match point in tool usage?arXiv preprint arXiv:2501.12851, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.708055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.708055Z digest=sha256:39c8dafeb71d40ce54fe300221fcce861c99fc4f443f8a6ed915b2886fe01051

Observation d2222a21-b736-4cfe-8275-ce8d485c1f75 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.712554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.712554Z digest=sha256:3e07c7044136e861d4f525da864d31aff7a83a00fc48803906d771e652720c88

Observation d9ac0475-4325-4b21-8e75-df4f7b354e95 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.717827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.717827Z digest=sha256:b68c79677bfffb6cbf34769dce641a2774ae7a5ab5cc04505892159869553637

Observation f80d6b94-f83a-4286-9116-df602d194120 · outbound

This paper cites Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.722166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.722166Z digest=sha256:686c7ff0739744c0854fd0fefa00e61576d6f2057fdbba2175f66edbe55346a3

Observation cfadd12e-87fa-42be-8641-d1fbb1574dde · outbound

This paper cites Gaia2: Benchmarking llm agents on dynamic and asynchronous environments.arXiv preprint arXiv:2602.11964, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Gaia2: Benchmarking llm agents on dynamic and asynchronous environments.arXiv preprint arXiv:2602.11964, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.726706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.726706Z digest=sha256:7307bc5fb519d672e56b8bfdb533aa86becb4af267199d4289b7aa7941e0417a

Observation c345dfa9-7688-4b25-b7ca-c55a6fe11a38 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios GLM-5: from Vibe Coding to Agentic Engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.731165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.731165Z digest=sha256:e11c92b6b6372bc404125a8950b54f1ded115fef13378a7eac9c6189342eb8fe

Observation a1ad8539-97f9-47ff-874d-8edd8998011c · outbound

This paper cites Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.arXiv preprint arXiv:2509.26490, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.arXiv preprint arXiv:2509.26490, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.736318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.736318Z digest=sha256:8d72e5713dc10cec9523cd2dd64ca38e4843aa7f47284b60ce442da7f14bef43

Observation d532e613-26c8-4670-99cf-ee37c8d0485f · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.740857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.740857Z digest=sha256:2983ccaa7d1487283dc35ef60d7dbde0594a5994968271f9dd995bcb40be7fa3

Observation de01f234-0ee2-40bb-a02f-ae99aa420a8e · outbound

This paper cites The tool decathlon: Bench- marking language agents for diverse, realistic, and long-horizon task execution.arXiv preprint arXiv:2510.25726, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios The tool decathlon: Bench- marking language agents for diverse, realistic, and long-horizon task execution.arXiv preprint arXiv:2510.25726, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.745842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.745842Z digest=sha256:dc6a20aa43f25dce3ed869795bc1e16952b376033d8c43f4cb77a3e8d6395cdd

Observation ae5b7737-562e-4b11-abc5-8bf9457f9a66 · outbound

This paper cites Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai, 2025.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.750217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.750217Z digest=sha256:4817a09aa4846eb6048df980d4b73f713ae3b9f7e4ae4e615dddc66f3bd75b36

Observation dcec0a53-67d5-4c3d-bcb9-1ae873084715 · outbound

This paper cites Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.754402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.754402Z digest=sha256:28db859711bdc421e30f448f54a6f7a6d353d31cc5638bcbab35269a758c1427

Observation 35d0ddb8-f4d3-4ac4-920f-1a718e7be0cb · outbound

This paper cites ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.758892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.758892Z digest=sha256:baa679de4dcbec254fb549e29a5bf1d417f7195d7e78acc819869c46907a5bc0

Observation 57e6884d-83e8-4a3e-aa7c-f3d8bac50212 · outbound

This paper cites Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.763968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.763968Z digest=sha256:f92170d289fe67c4af1aa286dae05dbff55d9e31d5a5ccb13a34700e5968da3d

Observation c13722e1-b478-40ba-a0d6-96a425633d9a · outbound

This paper cites The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.768484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.768484Z digest=sha256:2fbf51c4c47c985d18703e8d602ff4b5ecaa2f2e7fbf9a975f828532b530d2a2

Observation 146c6ecc-0f79-4997-83da-47bd6ce570af · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.772715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.772715Z digest=sha256:04052c30cb676ed698cf940373b458e2358b08402908437c702952e67a186dbe

Observation d90dc87d-f0b7-4e55-b017-cfc8530cc58c · outbound

This paper cites QwenClawBench: Real-user-distribution benchmark for openclaw agents, April.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios QwenClawBench: Real-user-distribution benchmark for openclaw agents, April

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.777037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.777037Z digest=sha256:eb971662be94b8d1dc96c2ea54f968a391a08bddb381e814a4bdefbeb288e3c9

Observation 8d291883-7561-47ed-b00e-4e53a70b11b8 · outbound

This paper cites One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.785663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.785663Z digest=sha256:745a212ab9bfc84817c5e66a89e4012ea2379dd878d84db5fd11265a71772555

Observation 8a43c760-0979-454d-9dcb-148f7ee0153a · outbound

This paper cites URLhttps://arxiv.org/abs/2603.04370.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios URLhttps://arxiv.org/abs/2603.04370

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.789506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.789506Z digest=sha256:7e3eaa9348dedeae33db4c21871091c25b37307ca624ced9037d102b37db1bf2

Observation e55cf75c-035a-4b25-a140-5a352ebe7b0d · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Kimi K2.5: Visual Agentic Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.794297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.794297Z digest=sha256:ac2150686827fe2bc978bcc0c44526f4b3949a047583ce70d508f98fee8246a7

Observation bb12d6ba-b891-44d2-be04-6d568462a14a · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.799456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.799456Z digest=sha256:c8ef689ff16b4965c0b4f6c752d225895b24cdd624707c8cd8a79f40db327434

Observation af412506-04f8-471c-97d8-06db166c7e8e · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.804245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.804245Z digest=sha256:7041447915f7d9c126d6661e70ccea8dfa2d3d4e1c3bd14988ab74b40a5d6d8a

Observation 380c0c35-124a-4edf-9d35-aae77bbfd9a5 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.809267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.809267Z digest=sha256:31392740e962fcd39e0f228d35c7d339dc3ac85f61a29375fa253f39e3bf526e

Observation f7303418-7fa0-46ef-9090-43c43d571a79 · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.814013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.814013Z digest=sha256:2afe8d17970263be7744535711310a063164af33f2c586e237d080101984f0ae

Observation 6fe38867-3a50-4928-8cc1-323641177e9f · outbound

This paper cites Deepplanning: Bench- marking long-horizon agentic planning with verifiable constraints.arXiv preprint arXiv:2601.18137, 2026.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Deepplanning: Bench- marking long-horizon agentic planning with verifiable constraints.arXiv preprint arXiv:2601.18137, 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.818900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.818900Z digest=sha256:be8bc213143a843cd2952b6c4d8aa37b82d3f94e3805a7938bee9a688043fef4

Observation 532400e7-adb4-410f-acfd-6b22622c9a7c · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.824559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.824559Z digest=sha256:47a86e12dc77051bae441c666aa2fedc0bd0a1d31157809175ddafc4cae92a3f

Observation 7099790e-86de-4bb8-89bc-b4636ce2f1e0 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.828684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.828684Z digest=sha256:4c16908c72c55f9edcb202eb0b4350e904def12c281b04ca716b3439cc4bde73

Observation 72246456-5f0d-4556-9518-d2adf7b427e4 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.833186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.833186Z digest=sha256:18b2c799090902bd31595b19bf706e91d2346d17d20d485a1ae3ba1c8c5a42ee

Observation 67998d9a-8708-4d4e-b59d-12469d9c2195 · outbound

This paper cites This protocol evaluates not only task execution, but also clarification, constraint tracking, adaptation to user feedback, and state maintenance across multiple turns.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios This protocol evaluates not only task execution, but also clarification, constraint tracking, adaptation to user feedback, and state maintenance across multiple turns

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.837265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.837265Z digest=sha256:8b7906c66c2a2c5e5ffe25a7a8a5e7650a76561dbb930b424e288ccad3c8b40c

Observation 714e253d-0a79-4272-8875-05d696ed5080 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.842384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.842384Z digest=sha256:d6734c6f18f594d446bce4de25042c99ae028eb67a7e4decea304fb83bc637a7

Observation 77acbd66-6a2b-476f-bf2d-f0b21dcb3264 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.846871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.846871Z digest=sha256:91f2a199028a887cde2baefca640189c545fd81da6b93abe783f6a1dd4812774

Observation ef31ed00-4acd-458d-83e4-0b45f1df8868 · outbound

This paper cites 4.VerifyCodechecks the trajectory and final observation and returns a binary pass/fail result.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios 4.VerifyCodechecks the trajectory and final observation and returns a binary pass/fail result

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.852162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.852162Z digest=sha256:efbcbfedac6eb0056a140a71216cd8aae5670aea26e0ccb796b4efcaf7f5ab8c

Observation e7af32ea-81ca-4c1b-bc37-92a81124aa3d · outbound

This paper cites [...additional description omitted ...].

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios [...additional description omitted ...]

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.856453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.856453Z digest=sha256:090979b7495f2e50b2555a70e53cd5389c0866dfd25ccb2a141c56d53e75b658

Observation fcd424a9-9c6d-47a2-8ee4-9f4dde5f68d8 · outbound

This paper cites [...additional tools omitted ...].

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios [...additional tools omitted ...]

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.860557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.860557Z digest=sha256:3b30827c49cccabbb2b41544b6979011169522d66cc4dd221799037985df34c8

Observation 10c3a754-e51a-4626-9ef3-498b001c9a3b · outbound

This paper cites Your goal is to complete the user’s request in an interactive environment by gradually calling the available tools step by step, and to proactively communicate with the user.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Your goal is to complete the user’s request in an interactive environment by gradually calling the available tools step by step, and to proactively communicate with the user

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.865246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.865246Z digest=sha256:3284f599af32611030b061d68f516952cd02d3c928d081f99ad4b46b42def16b

Observation 85bd7733-9245-40e0-89a5-575d6b867983 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.869397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.869397Z digest=sha256:f880d96be0339e3db14d657d748bfa81ae14773ee5f2ed91b4eb9e54c00c1074

Observation e457aacb-d23d-48fc-92db-2e1a7301f1e2 · outbound

This paper cites <Judge reason> The trajectory correctly resolvesLF-2024-PI-024, confirms Yunhe Foods / Lin Qiaoxue / He Shan, verifiesVER-004 as current, and flags the self-referentialLNK-003link.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios <Judge reason> The trajectory correctly resolvesLF-2024-PI-024, confirms Yunhe Foods / Lin Qiaoxue / He Shan, verifiesVER-004 as current, and flags the self-referentialLNK-003link

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.874308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.874308Z digest=sha256:4e24fe7beab330168d33991d85d27a7067f85c2700bc697a3514faa17fdab54d

Observation 3b900804-95e7-4063-82de-3f3bce04cd52 · outbound

This paper cites store surveillance screenshots and incident timeline explanation.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios store surveillance screenshots and incident timeline explanation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.878606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.878606Z digest=sha256:7c9cb4a959b5862e623ddfdf4e705ca594c278109330eaec78363076ecec27f0

Observation 008c583d-9452-4fa5-936d-392a1db75110 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.883109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.883109Z digest=sha256:fb9ba786e6a793fed10839d2de75c3657d504bb64d55eb3aa38529c216ab0e70

Observation 8db42b9d-0575-4d7e-9996-7269cca3ffab · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.887535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.887535Z digest=sha256:c2e1bb1b81b20d453ec6df2266b8f8953fa2c764bf2ec0bc1617c70b82b50b1a

Observation bd688285-357f-4f79-a60f-60a168ba8c39 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.892243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.892243Z digest=sha256:53346bb6d8b78ecaa064a0d15bab6b0df8dd34346ffcff39fb7411fde5cc165a

Observation ecf144a3-b8d9-41c8-9c6e-b03b254a77ae · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.896432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.896432Z digest=sha256:a6af8419bb414a9bd75ba51e8ded034fa349369fb9c087670f6dc82dcd1a8aaf

Observation 2861883a-b8c8-47d8-b643-7ac88668d0c4 · outbound

This paper cites an unresolved cited work.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.781239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.781239Z digest=sha256:feb3e0293bbc45f853ad7fc1006d7f707401ce3183d1daba3d434d976b0d6d01

Pith citing papers

Observation 2bc7ea6a-b277-4acd-9cb3-21332b85d0d0 · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:16.681296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:16.681296Z digest=sha256:24e10a71ad48f8bfcedf29827e6d7c5d545c97ef63deb3279b27e05333aa5210