Pith. sign in

Paper Citation Record · LEDGER

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

As of 5 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2605.07247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07247 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:34:24.662157Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0dc5fe6c-9b8f-4f02-8066-4ff477331a20 · outbound

This paper cites Yu, and Ming Zhang.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Yu, and Ming Zhang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.964740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:6138d0591781d1851bbc3364b0c359ae9c824efc5e137ca58e6b8d894f2c457d

Observation 82aea30d-b6bc-48cc-90a0-c5dffaefe32b · outbound

This paper cites τ-bench: A benchmark for tool- agent-user interaction in real-world domains.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation τ-bench: A benchmark for tool- agent-user interaction in real-world domains

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.946181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:4fa1a28f4bcc7b50f9b9a1cac0817ed79e4f64875c668afa6f9a0c69ff1eb972

Observation 603afbac-cb92-4d12-a2fd-ec52ad5306da · outbound

This paper cites Userbench: An interactive gym environment for user-centric agents.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Userbench: An interactive gym environment for user-centric agents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.970976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:26e71a6c022e537eebd40f38c90d256c02225c59e6a9e5c7cdbf0160023fdd9f

Observation 6ce585e3-b4a3-4c82-9900-71bc1af714df · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.942341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:b451969ae46134d2727ab286aa0c8b8320248779d2bc8785f061ea4604e9be02

Observation 0f45c943-dc36-424b-84be-f4fe69daac79 · outbound

This paper cites Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.960267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:5b1e204e22b6f9000d3a132046ed17bf363ae7cfde810451c0b7e2fb94718e5e

Observation 19b971c7-73f8-4444-a06e-7d86a577bdbb · outbound

This paper cites an unresolved cited work.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-14T16:11:57.978207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:a1f5b7ded33294e5921d2efd56b0f8888f345b4ad769a4914ede212256e9623b

Observation 72ae48bb-def6-4078-8b98-3993ccdc1bb9 · outbound

This paper cites an unresolved cited work.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-14T16:11:57.966764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:4bb884df4967d1656b0358b8a30db11f655f0b717b4e58dbccd40ed8c461ea1f

Observation 1db7ba02-1d12-4b64-ac6c-5b64d57f84f2 · outbound

This paper cites Are: Scaling up agent environments and evaluations.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Are: Scaling up agent environments and evaluations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.976560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:1346046c32d82e1b978aaff30db05c2b253f9dfd15db5dfb1196cce00160e32b

Observation 974e9eaf-633f-4b06-b7fb-bbe85201c1cf · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.950340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:82cc4574d570195cf5443553758ddb58b916e451c81c724d72084095e096ea8c

Observation 78321067-5c46-4c9b-ab6f-1a6b3a394862 · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Webarena: A realistic web environment for building autonomous agents

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.972748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:25c0ed2e404a9c84304ee3cc3c596641b6166c4342b7cec56b637fc9295d61f5

Observation 528e2ee6-228a-4353-ac17-1614dd277a94 · outbound

This paper cites {ALFW}orld: Aligning text and embodied environments for interactive learning.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation {ALFW}orld: Aligning text and embodied environments for interactive learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.974710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:200d5f69e5eea8321bb69cd3e057e3ce20b89bd5b7aa34261a32d7c2553615ff

Observation 413d3639-68bf-4462-8596-f42ce2a3219b · outbound

This paper cites Simulating environments with reasoning models for agent training.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Simulating environments with reasoning models for agent training

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.956364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:a163b0b56ae9a1c29bdef145a7b38b83b161454a25079f60b23b665ceeb5d722

Observation b76d50ab-9c5b-4f6d-acc4-cc8e0bc06433 · outbound

This paper cites Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.944189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:d63cc510686e58dfee5d6ef3ef232cc623983aef1080587a2b973beb3b76fd5d

Observation ff9012a4-18be-472a-95e9-44a4dc59cd1e · outbound

This paper cites Survey of hallucination in natural language generation.ACM Computing Surveys, 55(12):1–38, March 2023.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Survey of hallucination in natural language generation.ACM Computing Surveys, 55(12):1–38, March 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.954444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:209eb36343e9abf758c4ec5b2438a279d4430ded99f82da86f8e345738f50941

Observation 49589e0b-565d-4d37-a5db-d63e3aa07ff3 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Gorilla: Large Language Model Connected with Massive APIs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:22:17.488448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:d3f95b562b6e57229bb86db0f8f4c1ce156d040718ab97f0f097bfe946aecc24

Observation 29ced906-5e7c-4a8a-8932-cfb2502f6fa8 · outbound

This paper cites Siren’s song in the ai ocean: A survey on hallucination in large language models.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Siren’s song in the ai ocean: A survey on hallucination in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.958335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:2aa0f4d09ceaf0ac047c51a8a620e9240967927cfc0ab55c48233c4b6dee2fef

Observation 4d36844c-9fa8-49b0-9b9b-e7dc6070fdd2 · outbound

This paper cites Maddison, and Tatsunori Hashimoto.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Maddison, and Tatsunori Hashimoto

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.962924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:239c183dbb4b321792856cb27c60cb9d5ba3a8df030b94a1dc2325c8c39828d0

Observation 65df21cc-0c48-49ae-a0ab-e9b72a46e349 · outbound

This paper cites Measuring and improving consistency in pretrained language models.Transactions of the Association for Computational Linguistics, 9:1012–1031.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Measuring and improving consistency in pretrained language models.Transactions of the Association for Computational Linguistics, 9:1012–1031

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.952450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:c75e716a32aa2830ea62067d22bad839e1c8ae4339643564e3c1502750dda655

Observation fdf8c5bd-6ab7-440c-b44b-8c8134f8f6c9 · outbound

This paper cites Apigen-mt: Agentic pipeline for multi-turn data generation via simulated agent-human interplay.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Apigen-mt: Agentic pipeline for multi-turn data generation via simulated agent-human interplay

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.968767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:28da0b67aa30d0541ba26bd50c186a615507873e988307945cad979b39658afc

Observation 89f1a842-2651-405d-9cb5-490caecbb8b3 · outbound

This paper cites Springer.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Springer

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.940448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:d19ada79392d72f0058a6a5409460e83110201469484c258293daacd96b2f062

Observation 3bbdfb63-c875-4bff-8532-9db2d9941f93 · outbound

This paper cites Interactive fiction games: A colossal adventure.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Interactive fiction games: A colossal adventure

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.948014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:b2fc2987d5a49ed6f2f6d773fd410cc72aec63d1e793b4cc1536b81d717e9511

Observation ce1f80fe-575a-4ef8-91c9-560a5045bb1b · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.938448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:a55de18062a25f8fbcbe6a8f809c85edc751175ff7c2a093c1aba5dc29597cb6

Observation 14771c3f-1536-4200-af16-31b536c98692 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.926119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:84e5420a77da91295dd47aa4ca9b7c4ce20fdfec5c995961d6adbcacb1790027

Observation 08cb15e6-43a6-4dcd-b7ce-5136aeacce23 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.934653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:061c053eaee88555e5e33484ea5daf973af105052cb9c49bb0ff001f35c1265b

Observation e65e8568-95b5-494d-a32a-5680206b9638 · outbound

This paper cites Agentbench: Evaluating llms as agents.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Agentbench: Evaluating llms as agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.930119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:f6b133ed0959d28272bcd31f082373ea046d0fbe4e33dd6b8e039b36e78cdd8b

Observation f477d23f-871c-4b31-bba6-54a088ea2bfd · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Mind2web: Towards a generalist agent for the web

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.918500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:94d507dabd2cebcc95d60dc7d03b62843e7bcb486c74623c9e630d37f45a3137

Observation d0c9f1c6-2f57-44ae-90aa-3d91be73327b · outbound

This paper cites Agenttuning: Enabling generalized agent abilities for llms.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Agenttuning: Enabling generalized agent abilities for llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.915997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:e108f9cf2c67b9edcf7efc72149dba12d782e16cc8b0de42e1cb75cdd6d48064

Observation 542d7999-d460-4431-9779-cadb60b8884d · outbound

This paper cites Toolllm: Facilitating large language models to master 16000+ real-world apis.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Toolllm: Facilitating large language models to master 16000+ real-world apis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.936601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:ed9aa209af83aa79116f8f35c6e82293865338fae9ed5b643cc918d55afbda01

Observation b70674dd-665a-4e59-8298-952bfe7ed487 · outbound

This paper cites Metatool benchmark for large language models: Deciding whether to use tools and which to use.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation Metatool benchmark for large language models: Deciding whether to use tools and which to use

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.932274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:c23ad179a87fed4624cf306923ede04f1254d9b31a64876215fad4d00150f4e0

Observation 857ba90c-b839-4eb3-baee-b61a6e4ad69e · outbound

This paper cites τ 2-bench: Evaluating conversational agents in a dual-control environment.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation τ 2-bench: Evaluating conversational agents in a dual-control environment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.928118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:c2a983e6286afea291b7c2901f4a9c3889245f3ee994dfa4ebac9c976cac5f55

Observation 86265249-c411-447d-a122-e9141e267d12 · outbound

This paper cites APIGen-MT: Agentic pipeline for multi-turn data generation via simulated agent-human interplay.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation APIGen-MT: Agentic pipeline for multi-turn data generation via simulated agent-human interplay

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.922340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:6bfa5d5b76a3aad13ac03a8346ad2275ab1a240a14551e35866c309e69e5547d

Observation 1dcba389-cced-464e-a5bc-79d389089aa9 · outbound

This paper cites LlamaFactory: Unified efficient fine-tuning of 100+ language models.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation LlamaFactory: Unified efficient fine-tuning of 100+ language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T16:11:57.920419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:b7ddc1c0f02370d155c63bf66aeaa07e5819791f23f1d64d654a59febee53628

Observation 1d573ee2-74af-4992-9d0c-bfc37c4dd26e · outbound

This paper cites add blocked entries for 2025-05-01 through 2025-05-10.

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation add blocked entries for 2025-05-01 through 2025-05-10

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-05-14T16:11:57.924221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:34:24.662157Z digest=sha256:b65fbf111a46b8aa9d73a701f5d306f15be8ff0a9b5804a91c755cf6d077b5fb

Pith citing papers

No inbound Pith citation observations are available.