Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:32:49.575837Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.02372.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:32:49.575837Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 504de7a3-c67a-47a9-8bcc-9fc7f327b4a1 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Claude Haiku 4.5 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6651767-5d6b-4e87-9e57-7be3754334cf · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Claude Opus 4.7 System Card
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb68e588-dc12-4889-a02c-ae1a043ca196 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61c0bf5-ac1a-4ccb-8fc9-43d31a4b96a6 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8cba22b-efea-4878-94e8-fe83925c57f7 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timebench: A comprehensive evaluation of temporal reasoning abilities in large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fef0fec-1a1b-4150-a7fc-761f7621088e · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timer: Temporal instruction modeling and evaluation for longitudinal clinical records.npj Digital Medicine, 8(1):577, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ce5431-deca-4e9f-a547-6dd8fb684c4c · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cca429d-0fc0-4c08-a55c-1c4a31f629a2 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Mind2web: Towards a generalist agent for the web
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce20688-edbe-45a7-9bb0-5f444207b151 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.1: A consol- idated multi-domain dialogue dataset with state corrections and state tracking baselines
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ffb4de-1d67-4d89-8d48-91273adfc39b · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3, 1950
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a40c3e-78a1-42fe-ac92-2c25ac880ac6 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gemini 3 Flash Model Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bf996f-f749-4ae0-859c-fc4493a44e5b · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gemini 3.1 Pro Model Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b067a85c-c5c5-46a1-8cb9-b206d13a74bb · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Weinberger
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f4bd87-66ad-4a6f-a044-5613a91275f4 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Multiwoz 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830417f9-cd75-43b6-afd5-a1ca3fe3b38c · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Towards explainable temporal reasoning in large language models: A structure-aware generative framework
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08f3d5c5-c7f0-42cc-bb66-d06d2a73ae06 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284f6816-524c-4673-b9d1-b91056d127c5 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Beyond perfect apis: A comprehensive evaluation of llm agents under real-world api complexity, 2026
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e060976-f4d5-4e0a-952c-fb79fe3dd40c · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Counterfactual-consistency prompting for relative tem- poral understanding in large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84088aa6-4cc8-4cc2-ac77-616a30bf0104 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Open university learning analytics dataset
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a899b7-ad39-479d-b78e-e5d8038d5f50 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Prefix: Understand and adapt to user preference in human-agent interaction, 2026
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66332d06-a20d-41c4-afb5-9cb5b9bed7b4 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Ministral 3
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255100c5-8742-4475-b396-e1400a260400 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e86b4c-fcf0-4065-85c7-3749efa01303 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Mistral-Small-24B-Instruct-2501
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71686b62-e7c6-459a-bd70-7a6636cb7f99 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Time is encoded in the weights of finetuned language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a706882c-f725-4e17-bd48-abea13598413 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise GPT-4o mini: Advancing cost-efficient intelligence
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72914927-0152-422b-bfff-28cb4effaeb2 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Introducing GPT-5.4 mini and nano
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40f4512-4bf9-4399-97a5-9c030786c193 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise GPT-5.5 System Card
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f37d3c-eed8-47bf-8043-126e87bb8ea7 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Patil, Tianjun Zhang, Xin Wang, and Joseph E
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e01096d4-b7c7-4943-bd83-118cb1a920f0 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gonzalez
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c0ab94-5970-4b58-ad47-ffba373e0c5e · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise LLMD: A Large Language Model for Interpreting Longitudinal Medical Records
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d9be73-7da3-4619-8016-bb71a2b0e0db · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ToolLLM: Facilitating large language models to master 16000+ real-world APIs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10d19e3-2d72-4e75-9188-f1ad4d526830 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Qwen3.5: Towards native multimodal agents, February 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b448c8-e92b-47fb-af8c-57308fb56607 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8689–8696, Apr
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40a3819-c6ba-4e12-a8cd-b8fdf4d9aa35 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Appropriate reliance on ai advice: Conceptualization and the effect of explanations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47de76a1-0fa2-4b7c-942f-79974f0ab4a2 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timo: Towards better temporal reasoning for language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d6230b-03b7-4c1a-80fd-b2298631812a · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Paladin: Self-correcting language model agents to cure tool-failure cases, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e98ff6-7aba-4a92-b24f-eb9de741ed8f · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Agentnoisebench: Benchmarking robustness of tool-using llm agents under noisy condition, 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b08e8d73-0612-4b52-9e9b-f23719076039 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Butterfly effects in toolchains: A comprehensive analysis of failed parameter filling in LLM tool-agent systems
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1028e53-6f80-4a14-af3b-5d0589a0e7ac · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Reducing Tool Hallucination via Reliability Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a016a723-7571-439b-88df-58fd08e72c43 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44df9c9e-d438-4cc8-9a43-8a7c522550a6 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise τ-bench: A benchmark for Tool-Agent-User interaction in real-world domains
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d09bbd-ee5a-40d9-ade5-c71ec30779f5 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.4: A multi-domain task-oriented dialogue dataset with essential annotation corrections to improve state track- ing evaluation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff6ca94-4e10-49a4-95c4-83d46d8e6326 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee22073-1f93-49c2-bb72-19f0b9bcd373 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise From allies to adversaries: Manipulating LLM tool-calling through adversarial injection
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26473219-2caa-4716-8338-478878987eef · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b893b0a0-8e23-4c5e-b412-4f229cb1024f · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7205e0c-2a3a-410a-b1b7-6f0cb03fe58f · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce43522e-37e3-4a92-be93-0966cec3ac4a · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 101634f9-4680-4160-80c2-014b0440f502 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 998d2d4c-86fd-4029-90a9-bc50bc628c88 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3363725e-05f2-415f-aa54-fe35f671204a · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69870187-8700-4e82-b1ee-fbb3c57269df · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e5948d-c24b-4d20-a769-f1996a9946d4 · outbound
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Failure Risk: Critical
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.