Pith. sign in

Paper Citation Record · LEDGER

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.02372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02372 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:32:49.575837Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 504de7a3-c67a-47a9-8bcc-9fc7f327b4a1 · outbound

This paper cites Claude Haiku 4.5 System Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Claude Haiku 4.5 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:43.743272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:43.743272Z digest=sha256:d7669354e8e55ea1e559ff22a7ebafffc486e42b010b8295d3af01444a9c387a

Observation d6651767-5d6b-4e87-9e57-7be3754334cf · outbound

This paper cites Claude Opus 4.7 System Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Claude Opus 4.7 System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:43.873998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:43.873998Z digest=sha256:9151683eb45af08085c91899946365b327c4d601e900d904ee9459faf7343ef3

Observation cb68e588-dc12-4889-a02c-ae1a043ca196 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:43.971429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:43.971429Z digest=sha256:b3d8ecde574903c3e8ea0f6ef7d9b35c99d360db349de2e92794360d0a84a5d5

Observation c61c0bf5-ac1a-4ccb-8fc9-43d31a4b96a6 · outbound

This paper cites MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:44.072039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.072039Z digest=sha256:dc81bdf635e135bcdaff3dd1f00cc7aea7739ab9bcfe629bc8fdfe378f43e70b

Observation e8cba22b-efea-4878-94e8-fe83925c57f7 · outbound

This paper cites Timebench: A comprehensive evaluation of temporal reasoning abilities in large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timebench: A comprehensive evaluation of temporal reasoning abilities in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.169178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.169178Z digest=sha256:8a592a74c5e31aa253214fb97abae7fdcd754499407e72dd0c298c8ca2834c7c

Observation 8fef0fec-1a1b-4150-a7fc-761f7621088e · outbound

This paper cites Timer: Temporal instruction modeling and evaluation for longitudinal clinical records.npj Digital Medicine, 8(1):577, 2025.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timer: Temporal instruction modeling and evaluation for longitudinal clinical records.npj Digital Medicine, 8(1):577, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.284896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.284896Z digest=sha256:d0040944f94962aff1d39de095fe72454eb41ebbb668658ffc02b225718778ff

Observation b4ce5431-deca-4e9f-a547-6dd8fb684c4c · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.458523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.458523Z digest=sha256:6b083aceea4c55c73d2b5800edb3ccbdec0aff58ecf6d057f15ce51b5a5adbc0

Observation 3cca429d-0fc0-4c08-a55c-1c4a31f629a2 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Mind2web: Towards a generalist agent for the web

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.629181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.629181Z digest=sha256:d274f4c45523efa1af5a2893243b7946bd360e3a99ea4e18427a61692918cfb8

Observation 8ce20688-edbe-45a7-9bb0-5f444207b151 · outbound

This paper cites MultiWOZ 2.1: A consol- idated multi-domain dialogue dataset with state corrections and state tracking baselines.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.1: A consol- idated multi-domain dialogue dataset with state corrections and state tracking baselines

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.770225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.770225Z digest=sha256:f91827b27f7a4ced55c6f708cac927902eb94b50e81e06eef052aa317b8b1a29

Observation 72ffb4de-1d67-4d89-8d48-91273adfc39b · outbound

This paper cites Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3, 1950.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Verification of forecasts expressed in terms of probability.Monthly weather review, 78(1):1–3, 1950

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:44.861482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:44.861482Z digest=sha256:37f81e13ed6909e0fe97593915be894a8fb6f77eb2437eb191f57febf07bdfbe

Observation 00a40c3e-78a1-42fe-ac92-2c25ac880ac6 · outbound

This paper cites Gemini 3 Flash Model Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gemini 3 Flash Model Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.014107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.014107Z digest=sha256:15fcd02f6496dbb5aab2e6f5c89a55327f0d829e1c136aed7daeac861b204ebe

Observation a3bf996f-f749-4ae0-859c-fc4493a44e5b · outbound

This paper cites Gemini 3.1 Pro Model Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gemini 3.1 Pro Model Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.136049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.136049Z digest=sha256:ac58f50f1d8d1b32fd34b6a2f882c46e9be79748eaf1039b55d3145cc60db338

Observation b067a85c-c5c5-46a1-8cb9-b206d13a74bb · outbound

This paper cites Weinberger.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Weinberger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.233219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.233219Z digest=sha256:35b98629e7303ba7dfac84897a65aad9c81b0d209f3bb10a3254ece9e1938c7b

Observation 07f4bd87-66ad-4a6f-a044-5613a91275f4 · outbound

This paper cites Multiwoz 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Multiwoz 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.318824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.318824Z digest=sha256:713c419e8abb3e1d3df9ccc36e2932d5f0f73d20bf13104b64b62dcb292dbcfb

Observation 830417f9-cd75-43b6-afd5-a1ca3fe3b38c · outbound

This paper cites Towards explainable temporal reasoning in large language models: A structure-aware generative framework.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Towards explainable temporal reasoning in large language models: A structure-aware generative framework

Reference 15

Resolution
verified exact
doi, observed 2026-08-04T08:33:23.343722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T08:32:45.449837Z digest=sha256:55b7f445bdb69b92725770fdd509ef973e0244607a75a474cb8392d77f11fd43

Observation 08f3d5c5-c7f0-42cc-bb66-d06d2a73ae06 · outbound

This paper cites Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.556145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.556145Z digest=sha256:68ada2fae9751fa4f5ae8331b7f6d40568664d2609221eacd5f7a0fc35b00134

Observation 284f6816-524c-4673-b9d1-b91056d127c5 · outbound

This paper cites Beyond perfect apis: A comprehensive evaluation of llm agents under real-world api complexity, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Beyond perfect apis: A comprehensive evaluation of llm agents under real-world api complexity, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.656984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.656984Z digest=sha256:b1610143f97af995db4d694e3c3dc47a5caf97532162fd4ad95f5e5c702b0f48

Observation 2e060976-f4d5-4e0a-952c-fb79fe3dd40c · outbound

This paper cites Counterfactual-consistency prompting for relative tem- poral understanding in large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Counterfactual-consistency prompting for relative tem- poral understanding in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.761300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.761300Z digest=sha256:2ba4e6ee94910f27f2096b489dd10313129e4c88f6b1d1be9dba0e0259ea7a9e

Observation 84088aa6-4cc8-4cc2-ac77-616a30bf0104 · outbound

This paper cites Open university learning analytics dataset.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Open university learning analytics dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.875768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.875768Z digest=sha256:0ec5ff911481a7edb8dd227e72c29b7c5c466e713af38e081edce6f74cb600d9

Observation 26a899b7-ad39-479d-b78e-e5d8038d5f50 · outbound

This paper cites Prefix: Understand and adapt to user preference in human-agent interaction, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Prefix: Understand and adapt to user preference in human-agent interaction, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:45.975785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:45.975785Z digest=sha256:2adb3d68a61e91e8678da79010d9b5d47a782fd9cf9b63b757037eef318a80fa

Observation 66332d06-a20d-41c4-afb5-9cb5b9bed7b4 · outbound

This paper cites Ministral 3.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Ministral 3

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.082440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.082440Z digest=sha256:a137954a5d26092529775dcb61edf579535a9a4ae3ea1417d9b41f5bda557148

Observation 255100c5-8742-4475-b396-e1400a260400 · outbound

This paper cites ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.218986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.218986Z digest=sha256:17733156b8984ed9a233be05f0560cf4a5fb0c38f5f889fe74e4126fb801a38b

Observation b1e86b4c-fcf0-4065-85c7-3749efa01303 · outbound

This paper cites Mistral-Small-24B-Instruct-2501.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Mistral-Small-24B-Instruct-2501

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.348786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.348786Z digest=sha256:ab331d57189b4078ce64cb407119a33f505fc0b9d0c3d31cf5ee34fa4151337a

Observation 71686b62-e7c6-459a-bd70-7a6636cb7f99 · outbound

This paper cites Time is encoded in the weights of finetuned language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Time is encoded in the weights of finetuned language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.464972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.464972Z digest=sha256:de2b79aa1616f2e53084edac802a8901b6a475028b6577c44c564a1023f62c4f

Observation a706882c-f725-4e17-bd48-abea13598413 · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise GPT-4o mini: Advancing cost-efficient intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.582534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.582534Z digest=sha256:ab55691ddd18f5d06e60c9515a02768e31911842be6eb98db4c3b8905cf8212c

Observation 72914927-0152-422b-bfff-28cb4effaeb2 · outbound

This paper cites Introducing GPT-5.4 mini and nano.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Introducing GPT-5.4 mini and nano

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.659446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.659446Z digest=sha256:81f32294350ae1c26b39042a41f035b15dbfb2ab23c21ccf1f6a089c10da78ae

Observation a40f4512-4bf9-4399-97a5-9c030786c193 · outbound

This paper cites GPT-5.5 System Card.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise GPT-5.5 System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.815954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.815954Z digest=sha256:572be39dca0a160ddd3546c429b6f4323c5aa48f80b3de1dc2e2dbe5b96afe8a

Observation d3f37d3c-eed8-47bf-8043-126e87bb8ea7 · outbound

This paper cites Patil, Tianjun Zhang, Xin Wang, and Joseph E.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Patil, Tianjun Zhang, Xin Wang, and Joseph E

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.878354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.878354Z digest=sha256:75ad07368014ec47f2e985f79b3aa718dc8864d1ecee1d0e07bb1200f4635b39

Observation e01096d4-b7c7-4943-bd83-118cb1a920f0 · outbound

This paper cites Gonzalez.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Gonzalez

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.926362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.926362Z digest=sha256:c8722c8576f924cf38c922b89d674240d4941954e0b270021cb2fa8715738e1f

Observation f5c0ab94-5970-4b58-ad47-ffba373e0c5e · outbound

This paper cites LLMD: A Large Language Model for Interpreting Longitudinal Medical Records.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise LLMD: A Large Language Model for Interpreting Longitudinal Medical Records

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:46.996990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:46.996990Z digest=sha256:f4189363a095db40eb84025be0c30f21425c7d51f70c0b962e8432621478ad77

Observation 32d9be73-7da3-4619-8016-bb71a2b0e0db · outbound

This paper cites ToolLLM: Facilitating large language models to master 16000+ real-world APIs.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ToolLLM: Facilitating large language models to master 16000+ real-world APIs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.081849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.081849Z digest=sha256:65eee16d605374cc68fa228b84c3c7442765d09d95927773aa68886c7170d293

Observation c10d19e3-2d72-4e75-9188-f1ad4d526830 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Qwen3.5: Towards native multimodal agents, February 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.213183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.213183Z digest=sha256:3fefa9463c26aeba0756818ad3e5a439c4da803c4ceccb17c55d01788d2d1e3d

Observation 84b448c8-e92b-47fb-af8c-57308fb56607 · outbound

This paper cites Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8689–8696, Apr.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8689–8696, Apr

Reference 33

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:47.333976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.333976Z digest=sha256:ca62e14ad869e17772d3de5f854c9c021c1b3402e7e6e6f5e39bb015b206b506

Observation e40a3819-c6ba-4e12-a8cd-b8fdf4d9aa35 · outbound

This paper cites Appropriate reliance on ai advice: Conceptualization and the effect of explanations.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Appropriate reliance on ai advice: Conceptualization and the effect of explanations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.394889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.394889Z digest=sha256:a9e791dacc4450896abdf3607d076e20d708545d9295c2969dd3e0755f912920

Observation 47de76a1-0fa2-4b7c-942f-79974f0ab4a2 · outbound

This paper cites Timo: Towards better temporal reasoning for language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Timo: Towards better temporal reasoning for language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.507523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.507523Z digest=sha256:fb7dff672b3ca83a4ba95fe21ce44ee72e83c01ecf8b204172ed317bb237433b

Observation 89d6230b-03b7-4c1a-80fd-b2298631812a · outbound

This paper cites Paladin: Self-correcting language model agents to cure tool-failure cases, 2025.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Paladin: Self-correcting language model agents to cure tool-failure cases, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.600137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.600137Z digest=sha256:f72e8f7d58808903328a6cedfe0e1abd23b1c42540447319c5ab1e9ce0c5d85a

Observation 47e98ff6-7aba-4a92-b24f-eb9de741ed8f · outbound

This paper cites Agentnoisebench: Benchmarking robustness of tool-using llm agents under noisy condition, 2026.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Agentnoisebench: Benchmarking robustness of tool-using llm agents under noisy condition, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.762337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.762337Z digest=sha256:72f49c23ba606edbc884374fd9acaf05b4e028c430433e12ca5fd8914a0fb5cc

Observation b08e8d73-0612-4b52-9e9b-f23719076039 · outbound

This paper cites Butterfly effects in toolchains: A comprehensive analysis of failed parameter filling in LLM tool-agent systems.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Butterfly effects in toolchains: A comprehensive analysis of failed parameter filling in LLM tool-agent systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:47.874010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:47.874010Z digest=sha256:a46a75566a44e3d53b237780e59d5d29fbbe0f273878e9b0a8c280ae9d626a6c

Observation a1028e53-6f80-4a14-af3b-5d0589a0e7ac · outbound

This paper cites Reducing Tool Hallucination via Reliability Alignment.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Reducing Tool Hallucination via Reliability Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.011938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.011938Z digest=sha256:02121fd8846aafbca30acd24be0170574340a679bc1b4dab8a648f37378e1071

Observation a016a723-7571-439b-88df-58fd08e72c43 · outbound

This paper cites Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.154098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.154098Z digest=sha256:fd2d0aa35e1fafbe5bde7fa09cc370fc19e1b2057fdbc1568c78f345a615d5f7

Observation 44df9c9e-d438-4cc8-9a43-8a7c522550a6 · outbound

This paper cites τ-bench: A benchmark for Tool-Agent-User interaction in real-world domains.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise τ-bench: A benchmark for Tool-Agent-User interaction in real-world domains

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.264525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.264525Z digest=sha256:6e44cacb6e459f1213de6946f39832bcb53d5a23a3ffb70f6011e5c462f74ac5

Observation d1d09bbd-ee5a-40d9-ade5-c71ec30779f5 · outbound

This paper cites MultiWOZ 2.4: A multi-domain task-oriented dialogue dataset with essential annotation corrections to improve state track- ing evaluation.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.4: A multi-domain task-oriented dialogue dataset with essential annotation corrections to improve state track- ing evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.388211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.388211Z digest=sha256:873151fb3d543d2fdb7d5adecc00968e174aa2d075fa615f0a5d66a66c27fed6

Observation 8ff6ca94-4e10-49a4-95c4-83d46d8e6326 · outbound

This paper cites MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.511445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.511445Z digest=sha256:773871df9e29fe9263b25dbebc7e20f16ad7079a02a2f66dcfb1e3e73f212c79

Observation 0ee22073-1f93-49c2-bb72-19f0b9bcd373 · outbound

This paper cites From allies to adversaries: Manipulating LLM tool-calling through adversarial injection.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise From allies to adversaries: Manipulating LLM tool-calling through adversarial injection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.612316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.612316Z digest=sha256:d7f56f9742c719cbdcc358df6086dc1041fddb0544038494a2603364c97d7542

Observation 26473219-2caa-4716-8338-478878987eef · outbound

This paper cites CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:48.760073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.760073Z digest=sha256:abfa0e231699e49c5f1781c0cdaec106c5040afd54296bf86eb54f376c55ebda

Observation b893b0a0-8e23-4c5e-b412-4f229cb1024f · outbound

This paper cites ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:48.911714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:48.911714Z digest=sha256:7c5b2181e58d4c1e69cbb0829f70e3ec537e53399844c70885496854b850163d

Observation b7205e0c-2a3a-410a-b1b7-6f0cb03fe58f · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.005953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.005953Z digest=sha256:53d84a8c4694b5c70eb519f9ba4a512f29e6f32d55014f1ab7d6d706fc8910ca

Observation ce43522e-37e3-4a92-be93-0966cec3ac4a · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.105357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.105357Z digest=sha256:923870c4c276997c2117b6e5e17eb430066a1ae6ab84fc14954df1a567f450cf

Observation 101634f9-4680-4160-80c2-014b0440f502 · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.193660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.193660Z digest=sha256:15ed4001bba0e81f6bb0f5274c3053b2c314b145645351ed4affadf313059c8d

Observation 998d2d4c-86fd-4029-90a9-bc50bc628c88 · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.300953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.300953Z digest=sha256:c79664e63354fbac0c020d479a37ea44fb2e48b51f0eca1f1b9ffb576701d7be

Observation 3363725e-05f2-415f-aa54-fe35f671204a · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.405786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.405786Z digest=sha256:03b0f3423ca4e364838971f3929078d1701b8ce025710d4ad2dbc06982eacff1

Observation 69870187-8700-4e82-b1ee-fbb3c57269df · outbound

This paper cites an unresolved cited work.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T08:32:49.476122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.476122Z digest=sha256:13ef814168d11bdde86a2c730a2dede58a69ddccfd0370dbdcbce13ce5d4d9c7

Observation e5e5948d-c24b-4d20-a769-f1996a9946d4 · outbound

This paper cites Failure Risk: Critical.

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Failure Risk: Critical

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-04T08:32:49.575837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:32:49.575837Z digest=sha256:78567cc94548c3767c3b9c1ebad5d660bb0520b33b106c0e2b67ac85ac2211f9

Pith citing papers

No inbound Pith citation observations are available.