Pith. sign in

Paper Citation Record · LEDGER

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

As of 22 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.10042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10042 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:07.485591Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact8
  • verified fuzzy13
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0f803dc-53ec-4aa4-86b0-bd2376cd6061 · outbound

This paper cites Tuesday Retail Notes.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Tuesday Retail Notes

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:18:07.934741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.440337Z digest=sha256:61c539c259c7d307b9ecb3343ca09f284d1d82f659f9cd76b9494a97b97dfd93

Observation db2e51a4-4383-45c8-bf7e-ea8884e0005d · outbound

This paper cites Personalized Language Modeling from Personalized Human Feedback.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Personalized Language Modeling from Personalized Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:06.672677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:06.672677Z digest=sha256:a1904d92da98995013786bc4d14bb08ec0a28f8652593fb04f1b4043de49bd6d

Observation 517e448d-53ec-4278-a7ac-fb773b0f178b · outbound

This paper cites Aligning LLMs by Predicting Preferences from User Writing Samples.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Aligning LLMs by Predicting Preferences from User Writing Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:18:08.623168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:06.724740Z digest=sha256:3d8743a2e5f17edf7c412b02becc540d7b941b7217c324f4bd78c10a824582f1

Observation a7b0ef19-e581-4183-987f-e346d5b353c1 · outbound

This paper cites API-bank: A comprehensive benchmark for tool-augmented LLMs.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs API-bank: A comprehensive benchmark for tool-augmented LLMs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.874925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:06.774765Z digest=sha256:5556bb29e0e9a246ac6b8763c448ef55d45c5c23c246df55e888a3bce2a841ed

Observation ad1d714d-76d2-469d-b3a6-25bc860e3ea8 · outbound

This paper cites Benchmarking LLM Tool-Use in the Wild.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Benchmarking LLM Tool-Use in the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:06.864763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:06.864763Z digest=sha256:891f88f4c28f21a4f4a7f96e10d88c1e34c6a6574e68ed0f6d7f37fde011cdd3

Observation 1dd23d7c-8bac-4165-8e68-e4205f2b0193 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:06.914755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:06.914755Z digest=sha256:fcf973327b9b3c81a949d65fd4f07824a932b7013808e3b889f620e6e740ffc6

Observation c6b40833-3471-441e-8ff9-c6fa643a70ea · outbound

This paper cites Advancing and Benchmarking Personalized Tool Invocation for LLMs.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Advancing and Benchmarking Personalized Tool Invocation for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:06.995679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:06.995679Z digest=sha256:ffa117ee01058bce5b6fc670075e157f8fb4319571d44aa93db744d118d7be4e

Observation a72393e1-6011-4523-b149-226db075730a · outbound

This paper cites Tool- spectrum: Towards personalized tool utilization for large language models.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Tool- spectrum: Towards personalized tool utilization for large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.655761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.054748Z digest=sha256:cdb976bbbb74630ba1eb78fcea31bb8e25b1537bea2c4a6b38041369dd20d2ba

Observation c3dbc6e2-2c05-4cd2-b0fd-6ffb7a7af9b1 · outbound

This paper cites doi:10.18653/v1/2026.acl-long.370.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs doi:10.18653/v1/2026.acl-long.370

Reference 12

Resolution
verified exact
doi, observed 2026-08-14T04:18:07.767881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.084830Z digest=sha256:d24a95ad2aed0986f52fdcdb4e9ac1de6b1c790498de44d7f1ea2fe96894afc5

Observation 7ee59956-cf73-4f46-b6a2-0b4c852d1d37 · outbound

This paper cites Fingertip 20k: A benchmark for proactive and personalized mobile llm agents.arXiv preprint arXiv:2507.21071,.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Fingertip 20k: A benchmark for proactive and personalized mobile llm agents.arXiv preprint arXiv:2507.21071,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:07.135485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:07.135485Z digest=sha256:b4eebcaecc3f0b190fe20012585635380f18454e7e783b9d5b9b2dc9da98c512

Observation 979fbe54-b07c-42de-b2c6-13e77d830895 · outbound

This paper cites Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:07.145058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:07.145058Z digest=sha256:767d6a014542d04cefe346251c885c1e08e5905fe6e7a9e6987212db1d7555a5

Observation 3d496b7a-8c64-40ac-b6e0-1805b5055a90 · outbound

This paper cites Me-agent: A personalized mobile agent with two-level user habit learning for enhanced interaction.arXiv preprint arXiv:2601.20162, 2026a.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Me-agent: A personalized mobile agent with two-level user habit learning for enhanced interaction.arXiv preprint arXiv:2601.20162, 2026a

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:18:08.401262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.150631Z digest=sha256:220824a3b988b8300d926eb391cf32fef745e88ab8f3b446ec9475c8531e1aa4

Observation f8c2895f-6e01-46f0-a2c0-07fe6acf8e51 · outbound

This paper cites ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:18:08.286973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.165600Z digest=sha256:31930a901ec1ea4a2f30d56293e699b7e588878cd9046100d74e4b555ae18700

Observation 68011b94-2c97-44dd-8284-0a887c959a75 · outbound

This paper cites Shopsimulator: Evaluating and exploring rl-driven llm agent for shopping assistants.arXiv preprint arXiv:2601.18225, 2026b.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Shopsimulator: Evaluating and exploring rl-driven llm agent for shopping assistants.arXiv preprint arXiv:2601.18225, 2026b

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:07.173493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:07.173493Z digest=sha256:6f37698592a4ca0fa898d7165535d48d0627ed429ba6c54984176c80a1ed23f7

Observation 6932b77f-8bd6-4fed-9041-7eb330819605 · outbound

This paper cites PersonaLLM: Investigating the abil- ity of large language models to express personality traits.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs PersonaLLM: Investigating the abil- ity of large language models to express personality traits

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.587460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.186039Z digest=sha256:07f676e6ac53b06fdb008595e5bde77ed535f056f20a67e639edc95472e6e535

Observation 5f065543-cae1-4658-a75d-66bf879e6adf · outbound

This paper cites URLhttps://aclanthology.org/2024.findings-naacl.229/.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs URLhttps://aclanthology.org/2024.findings-naacl.229/

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-14T04:18:07.214750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:07.214750Z digest=sha256:bf351ea571c6736e65b12326344caaf0a57c4a3226c10523410dedbba0d049ab

Observation 368143bb-c301-43a6-96f4-1519c277ccdc · outbound

This paper cites an unresolved cited work.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:09.469404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.223364Z digest=sha256:f54eb07809b0335b83273a3863e7a550c34f0cf2ce491a79dc56ef352fc655c6

Observation 8aa8e873-aeb4-409c-8cd7-21377d555d8b · outbound

This paper cites Qwen Team.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Qwen Team

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.406526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.231328Z digest=sha256:550f294df38aae418d86cbe4fb7d5e767af31981ccdcdca14350956afec19a50

Observation 3d6bb34a-2c46-4c8a-8d39-3e5595bf4294 · outbound

This paper cites DeepSeek-AI.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs DeepSeek-AI

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.341136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.249330Z digest=sha256:98c1ce7e1323dbdc20e32af5704aa72464698f0b547d2dacf84eaafe0adcf905

Observation f026ab4a-d16c-4f84-9b91-b675ac11f5f2 · outbound

This paper cites Google DeepMind.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Google DeepMind

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.236384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.255437Z digest=sha256:d95a2f5cc68b55c18eff71578149dbb527ddd1fc095fa0db206c3344b6e38853

Observation a8f7a612-f0fd-4e71-a953-bb34be7ba50a · outbound

This paper cites an unresolved cited work.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:09.104749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.261202Z digest=sha256:57d76bb3a1bc9551144771dec64a481160046e53e6466e09a15bf92b4c8aff08

Observation f26e52b9-39c8-4f0e-8d05-9a286aaab7e4 · outbound

This paper cites Qiqiang Lin, Muning Wen, Qiuying Peng, Guanyu Nie, Junwei Liao, Jun Wang, Xiaoyun Mo, Jiamu Zhou, Cheng Cheng, Yin Zhao, et al.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Qiqiang Lin, Muning Wen, Qiuying Peng, Guanyu Nie, Junwei Liao, Jun Wang, Xiaoyun Mo, Jiamu Zhou, Cheng Cheng, Yin Zhao, et al

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:07.268543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:07.268543Z digest=sha256:4a7401939a8df2a2bc3ca9bd90c10b9f9612d8ebfd774e05ed1bedf29bebcbbd

Observation 14769fae-8a0c-44f1-8d5c-1eb8751057c8 · outbound

This paper cites Accessed: 2026-05-25.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Accessed: 2026-05-25

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.026654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.287222Z digest=sha256:d312714daf5e830da1e2699964f9721914645967d4e4d5d772ce6930cefd7303

Observation c47f70df-02f7-498f-823e-b26079cea3a4 · outbound

This paper cites co/Team-ACE/ToolACE-2.5-Llama-3.1-8B.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs co/Team-ACE/ToolACE-2.5-Llama-3.1-8B

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:08.944279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.296085Z digest=sha256:65a8f3376e33101d9993756fe6bd06faf88158d7c1902e813a2aac80b71cfce1

Observation 105f1510-6014-43e5-a1a1-c0022f71cd55 · outbound

This paper cites Accessed: 2026-05-25.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Accessed: 2026-05-25

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:08.869221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.319485Z digest=sha256:71571b62eca28db71f4c2f1d061fd1ec7a73be34ce1cd7bcda744ad85c49c860

Observation 99a4a60f-b52c-4512-8e2c-3f3a6665f622 · outbound

This paper cites Direct Multi-Turn Preference Optimization for Language Agents.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Direct Multi-Turn Preference Optimization for Language Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:07.327158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:07.327158Z digest=sha256:a6eaa29cbfb0444de2e4094df0646f6a50dae6b386eff055455f3c27d62dc83e

Observation 06c927ac-55a2-43a5-84b3-8487cd956336 · outbound

This paper cites URL https://aclanthology.org/2026.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs URL https://aclanthology.org/2026

Reference 30

Resolution
verified exact
doi, observed 2026-08-14T04:18:07.720918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.359764Z digest=sha256:c450d657c15ee04e372b6f95abe4e441aa26a8fee80770d750b8f47c79fce83a

Observation 32b744bc-020d-46c1-8c25-765f86d01d3a · outbound

This paper cites URL https://aclanthology.org/2026.findings-acl.1080/.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs URL https://aclanthology.org/2026.findings-acl.1080/

Reference 31

Resolution
verified exact
doi, observed 2026-08-14T04:18:07.689459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.395956Z digest=sha256:c19b25e781c003f68e62c1c54ec0dec1815d4332a68c24771ea7a1ddcf0cce04

Observation 51e3e151-5882-4d0f-bbe5-9ab7da35cd4c · outbound

This paper cites 36Kr GreenLeaf retail rising star.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs 36Kr GreenLeaf retail rising star

Reference 32

Resolution
verified exact
doi, observed 2026-08-14T04:18:07.595153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.427066Z digest=sha256:883d5df6f70331bdf7e036165cda478fab204cd37d00ecd9cad275775a641b91

Observation 72c0a1b2-46ad-4482-844b-ef89a96ff2d5 · outbound

This paper cites B.3 Quantitative Persona Diversity Audit We complement the field-coverage analysis with a quantitative audit of the ten selected profiles.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs B.3 Quantitative Persona Diversity Audit We complement the field-coverage analysis with a quantitative audit of the ten selected profiles

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:08.832172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.458135Z digest=sha256:0191b69ada3063d111b474c8e027ffa990763fa9f3354226a49e525c4c53f9de

Observation 07fe565e-04c8-4585-a2c9-2d88c98ca057 · outbound

This paper cites Left: token Jaccard similarity.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Left: token Jaccard similarity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:08.777001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:07.485591Z digest=sha256:f61d59670611f484b1533275c803af54371573ff881b31670787dd32e0218a5c

Observation 0f9fa615-9808-413b-88f4-21cf915a913d · outbound

This paper cites doi:10.18653/v1/2023.emnlp-main.187.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs doi:10.18653/v1/2023.emnlp-main.187

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:06.818994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:06.818994Z digest=sha256:934b9dcfe0e3574930f73a038317d4b1ddf18e7bac93ff712a74e3feca4690cd

Observation 2f9c32b5-d519-4eb7-9de1-e782eedbb5e4 · outbound

This paper cites Know me, respond to me: Benchmarking llms for dynamic user profiling and personalized responses at scale.arXiv preprint arXiv:2504.14225,.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Know me, respond to me: Benchmarking llms for dynamic user profiling and personalized responses at scale.arXiv preprint arXiv:2504.14225,

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:06.605781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:06.605781Z digest=sha256:e5c3e72a529524cf6acc976e5b1f3d97d186fe6ef6c23491a9b437c8cad3368a

Observation 27ae0a1f-ea3c-470c-9f0b-b490c0c9ee14 · outbound

This paper cites Personalens: A benchmark for personalization evaluation in conversational ai assistants.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Personalens: A benchmark for personalization evaluation in conversational ai assistants

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:10.014822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:06.644752Z digest=sha256:4a73a06452cd18adf6816248ca168b31bcc909d3c3643e3124e059a68588edd8

Observation 1ef11c0d-b1c0-41ea-bc04-2227aea61bed · outbound

This paper cites Petoolllm: Towards personalized tool learning in large language models.

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs Petoolllm: Towards personalized tool learning in large language models

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:09.802674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T04:18:06.948152Z digest=sha256:26390f0e0a44e66fc0d603234e99285eff73bf579897698c2ef3119f4cb1f476

Pith citing papers

No inbound Pith citation observations are available.