Pith. sign in

Paper Citation Record · LEDGER

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

As of 24 August 2026, this Paper Citation Record lists 100 of 153 outbound references and 0 inbound Pith citation observations for arXiv:2607.23722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23722 v1

Coverage vector

measured 100 of 153 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-30T15:06:50.929024Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 153 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d23089d4-61f0-4a7b-bbdb-fe94e923f2a2 · outbound

This paper cites The claude 3 model family: A new standard for intelligence, 2024.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The claude 3 model family: A new standard for intelligence, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.432689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.432689Z digest=sha256:688ece99bf56c5754e71cc43bcd93001a69da8a5073a3d566cb554a1cb0a814e

Observation ea6b8ca4-b165-4f04-9d0d-af0f6d52fe4a · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.437950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.437950Z digest=sha256:14a7710829349193fa8e392dc0eef62524d643ba45b473b6ff8e28c3fe5d27e5

Observation 41670fb2-1f50-4766-b240-27ec4199b402 · outbound

This paper cites VitaBench : Benchmarking LLM agents with versatile interactive tasks in real-world applications.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios VitaBench : Benchmarking LLM agents with versatile interactive tasks in real-world applications

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.445628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.445628Z digest=sha256:48bc1bef1f55fe33f1f39eacaf2b16a27f07707325b505b082b5f7e025bcfa46

Observation 1611af1d-2d1e-4c57-ae80-d623a95ffcc4 · outbound

This paper cites Measuring massive multitask language understanding.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Measuring massive multitask language understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.452283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.452283Z digest=sha256:e6db4a541b408bc5c7455609ae895b9fcf2c092d727dd95666457c9aab1dc81d

Observation 7baaf7c9-c91d-4cb2-82de-31141964e9c8 · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.457225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.457225Z digest=sha256:04aea57e36252c9ee28e7228857d0eeb1e8161009d208edfa4983f1ead2b3758

Observation 436ef764-3d88-4ead-8003-b91e06f282c6 · outbound

This paper cites Agentbench: Evaluating llms as agents.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Agentbench: Evaluating llms as agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.461755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.461755Z digest=sha256:50d4da1e3e43f9119b47e2597cff3156833ea6f12d4d81dcefa36200ac9e2880

Observation 9faf4ce9-fe33-44b6-8e92-0c9a08061033 · outbound

This paper cites MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.485800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.485800Z digest=sha256:214cdb0f9faae287e266cbc6a66e22329a8f16b8f12bdf90ff303e341b319fd1

Observation b2da1654-1f38-42bc-b4af-0eceeb491e6d · outbound

This paper cites Gaia: a benchmark for general ai assistants.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gaia: a benchmark for general ai assistants

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.490515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.490515Z digest=sha256:b26a4718561d4107e24bcca8622bc3f360c50a32e97bdecdbd91a4bd42349eef

Observation 749906e3-22e1-4593-86d5-f25ef8f9a211 · outbound

This paper cites LiveMCPBench : Can agents navigate an ocean of MCP tools? arXiv preprint arXiv:2508.01780, 2025.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios LiveMCPBench : Can agents navigate an ocean of MCP tools? arXiv preprint arXiv:2508.01780, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.495088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.495088Z digest=sha256:bd10b64c4a08bdeea3f6c76a116cc06e7d4129c68bb96b9c99b49ad39d5dc981

Observation d9ecc0af-2ffc-4c5c-8908-5d515d2a7732 · outbound

This paper cites Gorilla: Large language model connected with massive apis.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gorilla: Large language model connected with massive apis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.499658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.499658Z digest=sha256:b3175dcbcf374a6b427e2aa187c0e1e4147a0e6027000f5c1b22b9b1d61a6acf

Observation bb7287b1-9d9f-4d33-85bf-9880568fbfa4 · outbound

This paper cites Gonzalez.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gonzalez

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.503794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.503794Z digest=sha256:6c9ad0ec899fde74db099aa67d74303fe8df776553880ff762555442493b6213

Observation 08aa23dd-41a1-44cc-990a-026ceebb4a5b · outbound

This paper cites Tool LLM : Facilitating large language models to master 16000+ real-world API s.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Tool LLM : Facilitating large language models to master 16000+ real-world API s

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.507876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.507876Z digest=sha256:3f6aebc9dc646635c82aaf71e8d95fffd8686f59916aa223276298fb143e77bb

Observation 83dddffe-a3f9-4319-a24b-1bd4c218947e · outbound

This paper cites an unresolved cited work.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.512146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.512146Z digest=sha256:e1459059f0db3d66c208b32f0af202d51140e3779b941c6a51c5c168fd4c7e44

Observation 8b4ae660-82e9-4357-833d-4fc9b015d405 · outbound

This paper cites Appworld: A controllable world of apps and people for benchmarking interactive coding agents.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Appworld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.520853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.520853Z digest=sha256:c743319b4c20cf5dacacf2dc0eb9c903474ade47d89a48cf827d38e3a1a467f9

Observation 9a51692f-8694-48a0-8fb4-0a9e6671e89a · outbound

This paper cites MCP -bench: Benchmarking tool-using LLM agents with complex real-world tasks via MCP servers.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP -bench: Benchmarking tool-using LLM agents with complex real-world tasks via MCP servers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.525234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.525234Z digest=sha256:9d363710ec0dc1d6c36ffcf36725757a549b20361f00550eea47486c2d89155d

Observation 4556963c-f17e-475e-bf66-26d72c930d26 · outbound

This paper cites MCPMark : A benchmark for stress-testing realistic and comprehensive MCP use.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCPMark : A benchmark for stress-testing realistic and comprehensive MCP use

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.529517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.529517Z digest=sha256:1217e2c345db7997bce03fdc729fc1aa6cd2922c6d741db4621157a050c36182

Observation 11c59c88-6112-4111-9690-e715c6280ce2 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.538636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.538636Z digest=sha256:dfae734a3e8206db2cf261fa4de14360b19862f4c22d9d7b2fc5e43f11b6d8ec

Observation 13800e47-32a7-4fbd-bf4d-4e79a626ef32 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios React: Synergizing reasoning and acting in language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.543751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.543751Z digest=sha256:1bf812e8f759c48c23442e57126df2c7b19406b6f2b2742b2e850803b910f7de

Observation d9b2c2f2-d770-4007-ac4f-d7ee608520ca · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.555403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.555403Z digest=sha256:14fe9a494791fde183689c00a66d0ac1f94fb11909b53b4cfbcc856b3bbf5a6e

Observation 1bfeb167-351e-4ac1-960f-08be979f9906 · outbound

This paper cites 2025 , month = dec, note =.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , month = dec, note =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.560119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.560119Z digest=sha256:0fa63eeb00edd769a68150b203c048034dd616b5a240a2cb4efa1fa5634c1d60

Observation 5f00fb57-a8d4-482e-ae93-fe4f7f129307 · outbound

This paper cites an unresolved cited work.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.564540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.564540Z digest=sha256:be0e70ed3c2dbecc21c27ae3f954aee70a2772506bc6fd455b04af4c504b8d45

Observation 3fd35665-6d68-4fa9-98de-d00e66f5caea · outbound

This paper cites DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.569263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.569263Z digest=sha256:8d94211fda19fb3837eeeac931d77a272de7b20c3bb2e72a700d9990b99aea62

Observation a604f60e-9eb3-4895-9865-55edc3c049b6 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Kimi K2: Open Agentic Intelligence

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.573963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.573963Z digest=sha256:af430178a518e36f517116d41667e381d38fa2e52e98b7346513aa9d2d608153

Observation 663ee2bb-2bb6-4082-b032-e71ac3cff926 · outbound

This paper cites OpenAI GPT-5 System Card.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios OpenAI GPT-5 System Card

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.578780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.578780Z digest=sha256:cb3e6f7523593e0893272c5d71878e71a4f0ccbd1a8bd98f46e59a6b5f57e2fa

Observation c449061c-aa8b-4dcc-81fc-d62d550b3feb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.582903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.582903Z digest=sha256:bc68d57f904edd8bbb63625a48e5fe1a02e2c257c602286d0bb220e87828112b

Observation dbe9c79f-a62d-4bb6-8ce3-eb7fe2a23cfc · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.588264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.588264Z digest=sha256:907355c77ec0b5d4e2dc306b96d64489e43e43afc730cb7eb56e5c230c9d299c

Observation 1cb0b3b0-f6ae-40e2-b0ee-2630d6a15095 · outbound

This paper cites Qwen3 Technical Report.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Qwen3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.594115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.594115Z digest=sha256:250c222861b9afe36f4b098e041068d6780093513e5510db95af1c8d2b4b4b9e

Observation 16e483a5-4a28-4113-a7a6-3c8b6099ecb0 · outbound

This paper cites Nature , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Nature , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.598862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.598862Z digest=sha256:66da7b7a5d6bc994095a8d34844031d0bbc3c4df79b35fc26be5774e5f391ad5

Observation 8bab084d-e0ec-41f0-bee1-7510612efe14 · outbound

This paper cites International Conference on Learning Representations , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Learning Representations , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.603080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.603080Z digest=sha256:2ff2a76a4549a9417c89e2dd8ad6054408988010d0ed81d8bab6a5835df21a8a

Observation 9234a514-2ac9-414e-af0b-cd5de2157a31 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.607321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.607321Z digest=sha256:9c6fadd5e12b3b4e9a8e80884bc0158430687d5e0b89c093311ec10446a0d0e2

Observation aaa14e9d-6605-4f5e-8ffd-a9b9581794f6 · outbound

This paper cites 2025 , eprint=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.611715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.611715Z digest=sha256:391507fd12a76f8fba3b36fb50f47b86edebd9f0a4668ceffeabaeeb3e9f65f3

Observation 8fa87c48-6d32-4642-ad70-85ea410183e4 · outbound

This paper cites 2024 , url=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , url=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.616348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.616348Z digest=sha256:f224add879217b4f5d34d94750117e41c7e4bdb53332fe737e8b69809d6a024f

Observation 393b61d8-457c-4100-94ef-a4e40bda92a5 · outbound

This paper cites International conference on data intelligence and cognitive informatics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International conference on data intelligence and cognitive informatics , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.620691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.620691Z digest=sha256:6e9c15809de4e033d09bd729ba4d5f0b78bc1b31e0f511a325034ca937dbf918

Observation 0c1ba00c-dfbf-4acb-ac63-f8327b034434 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.625877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.625877Z digest=sha256:607683675a6fe99969de13f1b6562ad29b968246b0a93df795bee7bfa364fbde

Observation 6ca60aef-8939-4105-9801-dd3572331f9c · outbound

This paper cites ACM computing surveys , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios ACM computing surveys , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.630869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.630869Z digest=sha256:e3d020b0c44d1b325497e8100ed752328725e9601cdae23efb3cbdd119f2abfa

Observation 60c3fdd2-6314-4444-a1d4-c82a8e2f7e83 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.635692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.635692Z digest=sha256:4fb3910121fb34cf3fdd12a88127789742732c9751933505d47b0f4006217535

Observation 9c946090-8afb-40e4-9133-4177b74e51be · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.640357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.640357Z digest=sha256:40003c1b0553fed4fc344fae7b30b9ef1698efa7bf567a7165de6a35d0af64dd

Observation 7251f17f-8f50-45cc-80e6-1614488b5b91 · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.644972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.644972Z digest=sha256:84945cac1555812e00b89ff18e92bad188ff7864b34ef57265732e3d666801f3

Observation 9e36905d-20f9-4b70-8fb2-2e35175f3742 · outbound

This paper cites 2022 , journal=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2022 , journal=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.650448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.650448Z digest=sha256:791bace47da149a7c5a70880784d4b9b5044a66ef70c75502e31b6e4d10bb95b

Observation 63a5b5e9-2be1-4739-b10c-f7420d29410b · outbound

This paper cites International Conference on Learning Representations , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Learning Representations , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.654902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.654902Z digest=sha256:afa2bfede8ccb976d7b81cfa7933e7991f3f56c9e2d57fa571963aec0832bb0f

Observation bb504571-9572-4c4f-9bd9-0fe85975bf0e · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.659419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.659419Z digest=sha256:ec44d8b3cef3f80d8ee1438d4c4d688b538a9b67a5ee4658b794073286f0e68c

Observation 2a15518f-fd94-4a7f-a107-2c879c6434d5 · outbound

This paper cites International Conference on Machine Learning , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Machine Learning , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.664166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.664166Z digest=sha256:ea5450a86ef78e75e958b653d87ce0dbfd6a953e7aedbea2670e3234a3a5ff75

Observation d8d671a6-b954-4dc0-bf68-3c63b21c6738 · outbound

This paper cites Memory in the Age of AI Agents.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Memory in the Age of AI Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.668671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.668671Z digest=sha256:a98d9668564505edea3cccde7fdd751e76850ba1ab3562c2c863c4c944f6878b

Observation fe566331-bc70-4d1f-b9dc-51469b636247 · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.673864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.673864Z digest=sha256:0c25517f627708b16830874a3b1e67e3615414c1a9c6bf50bcb575b27e3787c7

Observation c48d34f3-b585-4117-8a88-71cbb1297464 · outbound

This paper cites A Survey of Context Engineering for Large Language Models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios A Survey of Context Engineering for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.678820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.678820Z digest=sha256:bd8fd972c6743f23fbd7e80e9c5543731b6c47b9a99c96c5bedc6fe8cbd49913

Observation 5f5bf318-db56-4cb4-a594-d6be71eeae3e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.683687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.683687Z digest=sha256:437610f07bc1d2ea02919c1445ca165496ee5581562e18258a02f50274e698a3

Observation ffbb6e0c-0559-4ca2-9dab-95044fb4ce8a · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MemGPT: Towards LLMs as Operating Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.688185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.688185Z digest=sha256:d212a2ba975500f554250ebceb0337cd967c274281b21260c625cd080a357318

Observation de8299ba-918b-491b-ae34-85297261db5a · outbound

This paper cites 2025 , eprint=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , eprint=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.692894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.692894Z digest=sha256:639ea2e14f00e89d7e8aea2d1e3746d9e6f51de0f41ff655fcc08f4115d6cc15

Observation c21d6de1-cd3f-4f4b-9ae3-f47fe13ccb80 · outbound

This paper cites 2024 , eprint=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.697385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.697385Z digest=sha256:5bb59798641971163ed71e31747457606bcd4c86e5efa5e5780849cac0b27dc0

Observation 7d6423f3-5828-457a-9d94-7e868d908ee4 · outbound

This paper cites 2024 , url=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , url=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.701724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.701724Z digest=sha256:345f9d1c9495fa8fb8c4a49dbc528cf661efa93c795a877f49a0b1f16ac88bc2

Observation 4f83c223-930f-448d-8718-dd0c8db160ab · outbound

This paper cites GPT-4 Technical Report.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios GPT-4 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.706215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.706215Z digest=sha256:a28ac2c92346ae46a26b012aafde40cf0e55080f057a8aca7be6f887949699a4

Observation 41c79bf0-9144-4042-9a74-ddc86e93b0c1 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.711289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.711289Z digest=sha256:82bdac8cc6957247298215496afc5bd4e9b403c241793155cc5966d25e88f904

Observation f367cb90-dd75-499e-9549-9c6145770500 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.716601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.716601Z digest=sha256:d4261c3e34ef481336d6e9c14fa7f82e21354c728f6ac9c0ce67ac6ef36ae9e9

Observation 53d45abc-89b5-431f-b5f9-ed8d395d7e18 · outbound

This paper cites The Llama 3 Herd of Models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The Llama 3 Herd of Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.721081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.721081Z digest=sha256:190230062fee363dd46dc01760616e8c3ecc99802d746fbc1474a7b70a893995

Observation 7e32090a-12d8-4798-bc4b-421caf897f2d · outbound

This paper cites Findings of EMNLP , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Findings of EMNLP , year=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.725998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.725998Z digest=sha256:5b5f0315ec9feb5905291229ed305fb174b4e1ee4ac90664a5c1e6ef3e1bf964

Observation 1a67e815-4097-455c-9aca-8374aae3a6c8 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.730622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.730622Z digest=sha256:d71a76c898d4b70cc38e3b88513c0891791728c2af57015ca3a7cb3407067c76

Observation ce1d9751-1236-4328-b462-3c4cbd4ad505 · outbound

This paper cites Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.735511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.735511Z digest=sha256:157e3e7cba3009111b0262bb0e206fb1841fae1a07e3f6fa6171553acbe35b13

Observation 9486c4ca-82c5-4719-a6a4-d9ef725c0c17 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.740519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.740519Z digest=sha256:7348012aaa0b9ae79269201c834b09c66802be73e86266ee1ec0a59be1bd1bd3

Observation 406c0b95-5214-4750-984d-4ebdcdbd4249 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.746188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.746188Z digest=sha256:5a5f5723fa5383f1a0a2af0f052ed47a8f15c7e775bea43e23d412f0113b1956

Observation 77af0014-4020-42b7-8390-5fb5d7ca707a · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.750693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.750693Z digest=sha256:d19e979fbb31caf717fc1aa522c5b8e0817ece2fe49b36a2cebbe79e0c00c268

Observation eb91d6cb-16a4-428f-b9b2-63c5b50d4734 · outbound

This paper cites 2024 , url=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , url=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.755689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.755689Z digest=sha256:6d2736f04be1f6f5b474342c04456c89710393d24510c752b94bf8b8e3e671f5

Observation 7ed6ac1d-fda7-4d4a-9110-b6dbfa8945d8 · outbound

This paper cites HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.760396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.760396Z digest=sha256:d783e65b913ff1ab3082ec7c75f6aac363f8ba03c0746e70dfe66052e94d9b22

Observation 282e8ede-bdcc-413a-911b-d59fe69cb688 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Transactions of the Association for Computational Linguistics , volume=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.764950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.764950Z digest=sha256:05b932fa63474e992e3ff7601cb0fd05b53d330e736893f73dab538cc86a58d0

Observation 675b1cd0-4f2c-4c40-8ab1-15a7a90acc31 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.769104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.769104Z digest=sha256:1443e69dc46c43fa67abeabdadc607a0bbea36be4f2a56f2ef925ed00ccdcada

Observation 5b474b89-7286-437b-b831-70e4656ac215 · outbound

This paper cites Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.773339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.773339Z digest=sha256:8c2ef1a769728aa6dcff5244bb465cd193927a82271022af6a09c5ff59876ad6

Observation c1621d39-f659-46ca-bc26-366eb2c5dd1d · outbound

This paper cites Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.777967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.777967Z digest=sha256:a953117220c5d0f297f06ee8f560e16e1ffa4deec889685fcd15d9bdead11b7a

Observation 7528bc7f-a0f9-4df1-849e-436e67b079dc · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.782018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.782018Z digest=sha256:4a100227afc803976320ce697e5e2e4841c943b3ed61a6d287823caeebb94414

Observation 99ab2634-dbab-408f-aacc-94ed92c26b89 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.786270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.786270Z digest=sha256:90c4bf917e7f4dd2f3e56fd95f769159f44d9b6877f2bead154dab610af204f3

Observation cb0d518f-abed-445a-92b8-ad7e3f72a157 · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.790666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.790666Z digest=sha256:c1a2810a6517e274505da8d1a62982119295c5cc7f4c2c9b90ca157c071ad4d5

Observation f4e0c5d8-a6ea-4edc-823a-ea00f21c9de1 · outbound

This paper cites 2023 , howpublished=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2023 , howpublished=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.795216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.795216Z digest=sha256:796100949671fe53bc836af37f763edd7faf34612dc5d54772e96b3b6adb0e9a

Observation 8afbb7c1-3b38-46ca-ba0d-27d0bc9faff1 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.799581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.799581Z digest=sha256:4f02e42aa6660455a8dd16fdbe21eb2487f2c4d5b557a42b94e8019143e88662

Observation 8859dca0-bb2a-4149-a0d5-6097e69d4df5 · outbound

This paper cites Proceedings of the 31st International Conference on Computational Linguistics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 31st International Conference on Computational Linguistics , pages=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.803559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.803559Z digest=sha256:418ffa09de060d797e44ac2a26c32defda4ddcd1c8efb08b1a13c4c5b36c6b38

Observation e156d0fb-af87-4346-8ef1-6d2152d6efbb · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.807788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.807788Z digest=sha256:b9ed84bf02650f19edfc06d7d2dd2c6d59b1e192798875781bbaee48cb8c2098

Observation 44f87baf-dc03-4aed-ad53-d6bed3fdc143 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.812271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.812271Z digest=sha256:8a2c09ae7642ed944e2eab99eb24b06dbb7c7f1dfbd1baf510f3329ad123a030

Observation 381ba74f-8ae1-4531-a890-2d81d47b7e2e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Instruction-Following Evaluation for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.816811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.816811Z digest=sha256:b757d3c73904c6fd532f4be44830f6181fe8253af53a45a957fe0d12b2d1cfc0

Observation 490982e3-58d5-4116-9428-249c276dca2e · outbound

This paper cites 12th International Conference on Learning Representations, ICLR 2024 , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 12th International Conference on Learning Representations, ICLR 2024 , year=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.821686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.821686Z digest=sha256:2a83b8ebff3974eb8ba8c5f64e7d729c9c38bf360aa85f4570a7ff71dc393bed

Observation 66aabc91-d026-49a9-9114-13af427fd794 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.825968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.825968Z digest=sha256:e4141ddb4e8cbbf7378fea5b89bf3cd14a241075aa26db605eabcd072e3e5081

Observation c62427a9-6ed0-4651-98ae-77923081505d · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Findings of the Association for Computational Linguistics: ACL 2024 , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.829949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.829949Z digest=sha256:265e5ddf01cc4ee8af242dd3885f53f78bb3e14dbbe65b5100a5b218cba0f616

Observation 782e5bcf-6efd-4c10-af8f-53dfd08f4ac6 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.834143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.834143Z digest=sha256:230c604e7a68f06e0524c001dce0bfa66e99841fcf7b5273dddd250a8ddea319

Observation 72afd11d-82b9-420c-8bb8-3589a8d170bb · outbound

This paper cites 12th International Conference on Learning Representations, ICLR 2024 , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 12th International Conference on Learning Representations, ICLR 2024 , year=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.839059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.839059Z digest=sha256:acb3e83912a52aa114c50af9522abf2deb643348a9c7cbd20919baa1957713a5

Observation 0f36f000-37f3-4b01-93b6-e936a14d57d9 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.843281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.843281Z digest=sha256:72cfe5d89d6ee2cbe7c3bc7542f56268a40187e6e31d714e3231b7d05e6c106d

Observation 543eb8a8-7eae-49d6-88b4-e0cbab2b4dfc · outbound

This paper cites AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.848292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.848292Z digest=sha256:d4e80e664843fd351c46a542bfc0e8452b96179db5de122326ff2ec1d41cea93

Observation 7aa57ca6-37ac-430d-84a2-c058854ff408 · outbound

This paper cites arXiv preprint arXiv:2511.10507 , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios arXiv preprint arXiv:2511.10507 , year=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.853058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.853058Z digest=sha256:607d4f4eb2e1bb9ab04436745202814b8557d01c1670f64343bdb184b758a559

Observation d4b68599-789b-435c-8a86-cbb9ed050131 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.858208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.858208Z digest=sha256:ab22d944d51fba0bef59b673cb38c77a455bb5ad79419843e449f53eb0a13dd5

Observation 49a59eb3-7bd5-4e50-998f-ff7c12f7cd3d · outbound

This paper cites Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , year=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.863309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.863309Z digest=sha256:f2880f2b101275b232e44ceb2d3278d47554dca5abce73cdf1a623e12b7d6ed2

Observation 70374c0c-db4f-431b-b5c2-f166eb60edff · outbound

This paper cites RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.867599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.867599Z digest=sha256:a87b633a572d844c54f180ba47b5e2362cb2162066442054692dc0fca9ab3dd0

Observation 2b78b768-697d-45d1-b113-e0137f379788 · outbound

This paper cites Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.872098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.872098Z digest=sha256:6265d6b8b603825c9bb6d628d429f4f76322bc374c93fd910a47810777fe0549

Observation ca890e42-18c7-42ef-88d0-6cd076503697 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Transactions of the Association for Computational Linguistics , volume=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.876103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.876103Z digest=sha256:555ea151a827e26432dcd9189be5ab7f8a0e1f773feef8a2e07d4e542569d878

Observation 175eca42-47d8-4753-877c-3e004cb7779d · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.880269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.880269Z digest=sha256:45d75fa15ecbe6ba183b3f93b99ebe7571e33389ff9e515c299855d10a3ff600

Observation a1396ab5-b73d-4700-8790-27b4dddf5f64 · outbound

This paper cites an unresolved cited work.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.884644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.884644Z digest=sha256:939ca741e45dcf3ab2ead414b00a762a600ca77b7599b9492f7be55d58cdabb7

Observation b75ab52b-5a9d-46b5-ab88-2bbc3d0e8fe1 · outbound

This paper cites International Conference on Learning Representations , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Learning Representations , year=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.889357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.889357Z digest=sha256:2553d281562fbdd7c3949cc34653f9cc6431b3aad9b242fc51ef742fd2db8aaf

Observation 61a4762a-5f90-4194-b209-5115cc89f4c6 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.893475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.893475Z digest=sha256:7b0c2d9d02d3102ca9377a5fa9eda5d4979b7e68a291fba05799c2579b10b232

Observation 33ebb7a5-735a-4953-9847-79c801f70974 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.897596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.897596Z digest=sha256:91e38fe740db998c4e6957b0eabe2076916b4e57b4d1486719a764aeaf43401f

Observation e17a205b-ca14-43f9-99cf-8f4373ec772b · outbound

This paper cites International Conference on Machine Learning , pages=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Machine Learning , pages=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.901925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.901925Z digest=sha256:797fe9d35219f6445477f5e629663107911a197f7d078deda8d5e6d0d1ff3e4f

Observation d84e1a8e-12d0-4a91-aacf-80ad70baba5c · outbound

This paper cites 2025 , month =.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , month =

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.906261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.906261Z digest=sha256:d55a264fbbf5b00aa67073efd0fef75959218dbfb8ab85a1106a111912e23dad

Observation fe91f08b-68b0-443e-9035-f008b32777f9 · outbound

This paper cites The eleventh international conference on learning representations , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The eleventh international conference on learning representations , year=

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.910350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.910350Z digest=sha256:67a1d9d4aee219744ce3ee06fc54cf01d3ad67a30ea49b2051899ecd432c73bf

Observation 861899d9-528d-4ee6-846c-ef0544b13a53 · outbound

This paper cites The eleventh international conference on learning representations , year=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The eleventh international conference on learning representations , year=

Reference 102

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.914493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.914493Z digest=sha256:bf6d184041dbd9de7b36c3ae1723e815501b371066cd5304ed77fcf516fe063f

Observation 4bc5a407-1aa8-4c21-8077-3807b207ef9e · outbound

This paper cites Transactions on Machine Learning Research , issn=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Transactions on Machine Learning Research , issn=

Reference 103

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.918914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.918914Z digest=sha256:91a14914e242f3345ca4ccd713aff777f32addc71abe25976aadd80860d69246

Observation fe172b63-e2f3-44e8-a740-76bd65b24e7e · outbound

This paper cites 2021 , eprint=.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2021 , eprint=

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.924418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.924418Z digest=sha256:ed400b27997d38fd5e9ed78815c3e2947c86f14a9f43473a548a2d023f512b33

Observation f4865f98-4fd9-45cb-b06f-b5a092efe053 · outbound

This paper cites Program Synthesis with Large Language Models.

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Program Synthesis with Large Language Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-07-30T15:06:50.929024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T15:06:50.929024Z digest=sha256:28ac577c0a5c2cabea90bd7a89c37fecc782c8f6e9d4b781724459f6f3493752

Pith citing papers

No inbound Pith citation observations are available.