Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-30T15:06:50.929024Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 100 of 153 outbound references and 0 inbound Pith citation observations for arXiv:2607.23722.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-30T15:06:50.929024Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 153 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d23089d4-61f0-4a7b-bbdb-fe94e923f2a2 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The claude 3 model family: A new standard for intelligence, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6b8ca4-b165-4f04-9d0d-af0f6d52fe4a · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41670fb2-1f50-4766-b240-27ec4199b402 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios VitaBench : Benchmarking LLM agents with versatile interactive tasks in real-world applications
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1611af1d-2d1e-4c57-ae80-d623a95ffcc4 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Measuring massive multitask language understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baaf7c9-c91d-4cb2-82de-31141964e9c8 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 436ef764-3d88-4ead-8003-b91e06f282c6 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Agentbench: Evaluating llms as agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faf4ce9-fe33-44b6-8e92-0c9a08061033 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2da1654-1f38-42bc-b4af-0eceeb491e6d · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gaia: a benchmark for general ai assistants
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 749906e3-22e1-4593-86d5-f25ef8f9a211 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios LiveMCPBench : Can agents navigate an ocean of MCP tools? arXiv preprint arXiv:2508.01780, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ecc0af-2ffc-4c5c-8908-5d515d2a7732 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gorilla: Large language model connected with massive apis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7287b1-9d9f-4d33-85bf-9880568fbfa4 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gonzalez
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08aa23dd-41a1-44cc-990a-026ceebb4a5b · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Tool LLM : Facilitating large language models to master 16000+ real-world API s
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83dddffe-a3f9-4319-a24b-1bd4c218947e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4ae660-82e9-4357-833d-4fc9b015d405 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Appworld: A controllable world of apps and people for benchmarking interactive coding agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a51692f-8694-48a0-8fb4-0a9e6671e89a · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCP -bench: Benchmarking tool-using LLM agents with complex real-world tasks via MCP servers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4556963c-f17e-475e-bf66-26d72c930d26 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MCPMark : A benchmark for stress-testing realistic and comprehensive MCP use
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c59c88-6112-4111-9690-e715c6280ce2 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13800e47-32a7-4fbd-bf4d-4e79a626ef32 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios React: Synergizing reasoning and acting in language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b2c2f2-d770-4007-ac4f-d7ee608520ca · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bfeb167-351e-4ac1-960f-08be979f9906 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , month = dec, note =
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f00fb57-a8d4-482e-ae93-fe4f7f129307 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd35665-6d68-4fa9-98de-d00e66f5caea · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a604f60e-9eb3-4895-9865-55edc3c049b6 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Kimi K2: Open Agentic Intelligence
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663ee2bb-2bb6-4082-b032-e71ac3cff926 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios OpenAI GPT-5 System Card
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c449061c-aa8b-4dcc-81fc-d62d550b3feb · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gemini: A Family of Highly Capable Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe9c79f-a62d-4bb6-8ce3-eb7fe2a23cfc · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb0b3b0-f6ae-40e2-b0ee-2630d6a15095 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Qwen3 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e483a5-4a28-4113-a7a6-3c8b6099ecb0 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Nature , volume=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bab084d-e0ec-41f0-bee1-7510612efe14 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Learning Representations , year=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9234a514-2ac9-414e-af0b-cd5de2157a31 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaa14e9d-6605-4f5e-8ffd-a9b9581794f6 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , eprint=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa87c48-6d32-4642-ad70-85ea410183e4 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , url=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393b61d8-457c-4100-94ef-a4e40bda92a5 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International conference on data intelligence and cognitive informatics , pages=
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c1ba00c-dfbf-4acb-ac63-f8327b034434 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca60aef-8939-4105-9801-dd3572331f9c · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios ACM computing surveys , volume=
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c3fdd2-6314-4444-a1d4-c82a8e2f7e83 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c946090-8afb-40e4-9133-4177b74e51be · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7251f17f-8f50-45cc-80e6-1614488b5b91 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2024 conference on empirical methods in natural language processing , pages=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e36905d-20f9-4b70-8fb2-2e35175f3742 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2022 , journal=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a5b5e9-2be1-4739-b10c-f7420d29410b · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Learning Representations , year=
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb504571-9572-4c4f-9bd9-0fe85975bf0e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a15518f-fd94-4a7f-a107-2c879c6434d5 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Machine Learning , pages=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d671a6-b954-4dc0-bf68-3c63b21c6738 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Memory in the Age of AI Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe566331-bc70-4d1f-b9dc-51469b636247 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48d34f3-b585-4117-8a88-71cbb1297464 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios A Survey of Context Engineering for Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5bf318-db56-4cb4-a594-d6be71eeae3e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffbb6e0c-0559-4ca2-9dab-95044fb4ce8a · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios MemGPT: Towards LLMs as Operating Systems
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8299ba-918b-491b-ae34-85297261db5a · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , eprint=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21d6de1-cd3f-4f4b-9ae3-f47fe13ccb80 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , eprint=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6423f3-5828-457a-9d94-7e868d908ee4 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , url=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f83c223-930f-448d-8718-dd0c8db160ab · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios GPT-4 Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c79bf0-9144-4042-9a74-ddc86e93b0c1 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f367cb90-dd75-499e-9549-9c6145770500 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d45abc-89b5-431f-b5f9-ed8d395d7e18 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The Llama 3 Herd of Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e32090a-12d8-4798-bc4b-421caf897f2d · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Findings of EMNLP , year=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a67e815-4097-455c-9aca-8374aae3a6c8 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce1d9751-1236-4328-b462-3c4cbd4ad505 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9486c4ca-82c5-4719-a6a4-d9ef725c0c17 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406c0b95-5214-4750-984d-4ebdcdbd4249 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77af0014-4020-42b7-8390-5fb5d7ca707a · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb91d6cb-16a4-428f-b9b2-63c5b50d4734 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2024 , url=
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed6ac1d-fda7-4d4a-9110-b6dbfa8945d8 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 282e8ede-bdcc-413a-911b-d59fe69cb688 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Transactions of the Association for Computational Linguistics , volume=
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 675b1cd0-4f2c-4c40-8ab1-15a7a90acc31 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b474b89-7286-437b-b831-70e4656ac215 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1621d39-f659-46ca-bc26-366eb2c5dd1d · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing , pages=
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7528bc7f-a0f9-4df1-849e-436e67b079dc · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ab2634-dbab-408f-aacc-94ed92c26b89 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0d518f-abed-445a-92b8-ad7e3f72a157 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e0c5d8-a6ea-4edc-823a-ea00f21c9de1 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2023 , howpublished=
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8afbb7c1-3b38-46ca-ba0d-27d0bc9faff1 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8859dca0-bb2a-4149-a0d5-6097e69d4df5 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 31st International Conference on Computational Linguistics , pages=
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e156d0fb-af87-4346-8ef1-6d2152d6efbb · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Findings of the Association for Computational Linguistics: ACL 2025 , pages=
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f87baf-dc03-4aed-ad53-d6bed3fdc143 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381ba74f-8ae1-4531-a890-2d81d47b7e2e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Instruction-Following Evaluation for Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490982e3-58d5-4116-9428-249c276dca2e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 12th International Conference on Learning Representations, ICLR 2024 , year=
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66aabc91-d026-49a9-9114-13af427fd794 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c62427a9-6ed0-4651-98ae-77923081505d · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Findings of the Association for Computational Linguistics: ACL 2024 , year=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 782e5bcf-6efd-4c10-af8f-53dfd08f4ac6 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72afd11d-82b9-420c-8bb8-3589a8d170bb · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 12th International Conference on Learning Representations, ICLR 2024 , year=
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f36f000-37f3-4b01-93b6-e936a14d57d9 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543eb8a8-7eae-49d6-88b4-e0cbab2b4dfc · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aa57ca6-37ac-430d-84a2-c058854ff408 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios arXiv preprint arXiv:2511.10507 , year=
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b68599-789b-435c-8a86-cbb9ed050131 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , pages=
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a59eb3-7bd5-4e50-998f-ff7c12f7cd3d · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , year=
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70374c0c-db4f-431b-b5c2-f166eb60edff · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b78b768-697d-45d1-b113-e0137f379788 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca890e42-18c7-42ef-88d0-6cd076503697 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Transactions of the Association for Computational Linguistics , volume=
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 175eca42-47d8-4753-877c-3e004cb7779d · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Proceedings of the 28th International Conference on Computational Linguistics , pages=
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1396ab5-b73d-4700-8790-27b4dddf5f64 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Unresolved cited work
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75ab52b-5a9d-46b5-ab88-2bbc3d0e8fe1 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Learning Representations , year=
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a4762a-5f90-4194-b209-5115cc89f4c6 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ebb7a5-735a-4953-9847-79c801f70974 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Advances in Neural Information Processing Systems , volume=
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e17a205b-ca14-43f9-99cf-8f4373ec772b · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios International Conference on Machine Learning , pages=
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d84e1a8e-12d0-4a91-aacf-80ad70baba5c · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2025 , month =
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe91f08b-68b0-443e-9035-f008b32777f9 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The eleventh international conference on learning representations , year=
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861899d9-528d-4ee6-846c-ef0544b13a53 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios The eleventh international conference on learning representations , year=
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc5a407-1aa8-4c21-8077-3807b207ef9e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Transactions on Machine Learning Research , issn=
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe172b63-e2f3-44e8-a740-76bd65b24e7e · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios 2021 , eprint=
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4865f98-4fd9-45cb-b06f-b5a092efe053 · outbound
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Program Synthesis with Large Language Models
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.