Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:34:24.662157Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2605.07247.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:34:24.662157Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0dc5fe6c-9b8f-4f02-8066-4ff477331a20 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Yu, and Ming Zhang
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 82aea30d-b6bc-48cc-90a0-c5dffaefe32b · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation τ-bench: A benchmark for tool- agent-user interaction in real-world domains
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 603afbac-cb92-4d12-a2fd-ec52ad5306da · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Userbench: An interactive gym environment for user-centric agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ce585e3-b4a3-4c82-9900-71bc1af714df · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f45c943-dc36-424b-84be-f4fe69daac79 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19b971c7-73f8-4444-a06e-7d86a577bdbb · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72ae48bb-def6-4078-8b98-3993ccdc1bb9 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1db7ba02-1d12-4b64-ac6c-5b64d57f84f2 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Are: Scaling up agent environments and evaluations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 974e9eaf-633f-4b06-b7fb-bbe85201c1cf · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78321067-5c46-4c9b-ab6f-1a6b3a394862 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Webarena: A realistic web environment for building autonomous agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 528e2ee6-228a-4353-ac17-1614dd277a94 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation {ALFW}orld: Aligning text and embodied environments for interactive learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 413d3639-68bf-4462-8596-f42ce2a3219b · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Simulating environments with reasoning models for agent training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b76d50ab-9c5b-4f6d-acc4-cc8e0bc06433 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff9012a4-18be-472a-95e9-44a4dc59cd1e · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Survey of hallucination in natural language generation.ACM Computing Surveys, 55(12):1–38, March 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 49589e0b-565d-4d37-a5db-d63e3aa07ff3 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Gorilla: Large Language Model Connected with Massive APIs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29ced906-5e7c-4a8a-8932-cfb2502f6fa8 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Siren’s song in the ai ocean: A survey on hallucination in large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d36844c-9fa8-49b0-9b9b-e7dc6070fdd2 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Maddison, and Tatsunori Hashimoto
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65df21cc-0c48-49ae-a0ab-e9b72a46e349 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Measuring and improving consistency in pretrained language models.Transactions of the Association for Computational Linguistics, 9:1012–1031
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fdf8c5bd-6ab7-440c-b44b-8c8134f8f6c9 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Apigen-mt: Agentic pipeline for multi-turn data generation via simulated agent-human interplay
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89f1a842-2651-405d-9cb5-490caecbb8b3 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Springer
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3bbdfb63-c875-4bff-8532-9db2d9941f93 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Interactive fiction games: A colossal adventure
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce1f80fe-575a-4ef8-91c9-560a5045bb1b · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14771c3f-1536-4200-af16-31b536c98692 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08cb15e6-43a6-4dcd-b7ce-5136aeacce23 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e65e8568-95b5-494d-a32a-5680206b9638 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Agentbench: Evaluating llms as agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f477d23f-871c-4b31-bba6-54a088ea2bfd · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Mind2web: Towards a generalist agent for the web
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0c9f1c6-2f57-44ae-90aa-3d91be73327b · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Agenttuning: Enabling generalized agent abilities for llms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 542d7999-d460-4431-9779-cadb60b8884d · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Toolllm: Facilitating large language models to master 16000+ real-world apis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b70674dd-665a-4e59-8298-952bfe7ed487 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Metatool benchmark for large language models: Deciding whether to use tools and which to use
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 857ba90c-b839-4eb3-baee-b61a6e4ad69e · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation τ 2-bench: Evaluating conversational agents in a dual-control environment
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86265249-c411-447d-a122-e9141e267d12 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation APIGen-MT: Agentic pipeline for multi-turn data generation via simulated agent-human interplay
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dcba389-cced-464e-a5bc-79d389089aa9 · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation LlamaFactory: Unified efficient fine-tuning of 100+ language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d573ee2-74af-4992-9d0c-bfc37c4dd26e · outbound
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation add blocked entries for 2025-05-01 through 2025-05-10
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.