Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:23.733585Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2505.14106.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:23.733585Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:26.701496Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T09:54:34.994786Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e9861fdb-99de-45c2-8df0-515c31a77adb · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Conversational Health Agents: A Personalized LLM-Powered Agent Framework
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00bf5c9-bc0c-47b8-b2f7-cd2eb10dc9d3 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persobench: Benchmarking personalized response generation in large language models, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03167942-a7e7-48c2-bf1e-e6ae533ea020 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Claude 3.5 sonnet
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 237bf522-348b-4651-a528-b532151c0119 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Little Human Data Goes A Long Way
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3e01726-ba8e-4cf9-b027-26cfc2176d36 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Graph-Based Retrieval for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03232c27-8236-4033-b7d8-f3d2f54d19af · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a903f797-26e0-48dc-b27f-1e35298c32d9 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations LoRe: Personalizing LLMs via Low-Rank Reward Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8722fdf-21c2-468e-b36a-a02ebb7c345c · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74b080f6-b2ed-42c0-a046-72636d8e93b4 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Beyond prompts: Dy- namic conversational benchmarking of large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3e6934d-15cb-4155-9c60-ad8fd8a2b49f · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ef064e0-e58f-4565-9410-fc52a2f1045c · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cc7d673-0b61-4547-aa25-30019abbd5a8 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations REALM: A Dataset of Real-World LLM Use Cases
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e070871-557d-43a0-adea-d59997ee4ab3 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b06db013-27d0-4816-9cb6-68ce7003e4d0 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88126155-e729-4f8f-8fc5-f140559fa286 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c48a18f2-6a9e-40d5-b11d-dd9a31c50a36 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations RedCaps: web-curated image-text data created by the people, for the people
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a06c4d6-7fac-4a12-a0aa-9fc9e604ac27 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The llama 3 herd of models, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850ceb35-110c-4dc9-9b9a-a27a373e84d2 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ruddit: Norms of Offensiveness for English Reddit Comments
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 194d8e5f-4214-4e7e-bfc7-0a4e1ee518eb · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4f4b3bd-d184-45bd-9946-d970bf281d4c · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 570f0968-3ae8-4466-8fbe-45f11d61adc4 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6daae78-9455-4ebb-ac01-1c48db55eb46 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A framework for building adaptive intelligent virtual assistants
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77ecc1ea-0684-4054-b46c-9b666cf43b3c · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Teach LLMs to Personalize -- An Approach inspired by Writing Education
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1945642b-fd5f-49bf-a9bd-77d49420dab2 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62133eab-494d-409c-9d3b-59c8a7b733aa · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 955fdb45-569b-4517-b7be-6a5169debecc · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rouge: A package for automatic evaluation of summaries
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 646fc09e-b860-47e2-a683-92459fac4bb5 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persona-sq: A personalized suggested question generation framework for real-world documents, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9533affa-216e-4632-8ac6-64befd694cc9 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a054c5ee-571e-44a1-bd79-40f20b53da15 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Gpt-4o mini: advancing cost-efficient intelligence
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ac76a55-64bc-42c0-96ca-c0cfd0f0339c · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Introducing gpt-4.1 in the api
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d69a277-3a48-4798-855c-a90eab1feb4e · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Bleu: a method for automatic evaluation of machine translation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e016f6-564b-4426-8085-d6a390ff8535 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fed4877-ccbd-4a02-9aee-37490e77aec9 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e44488dd-ecf2-4aaa-83fb-0f4bdac730a5 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Sentence-bert: Sentence embeddings using siamese bert- networks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08ac11f-6f2a-4ef2-bcd4-b722c05ab059 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Lamp: When large language models meet personalization, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 639f9236-0fca-4285-9c56-1765c052d4e5 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf2d392-76b7-4210-bacf-45ab18e9a5db · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Democra- tizing large language models via personalized parameter-efficient fine-tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d87413e-c6d6-43ab-9c83-59a6cddc2c4e · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b15c0d2-b97b-46a3-abd4-7495a23729a5 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Position: Will we run out of data? limits of llm scaling based on human-generated data
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d03cefe-d484-4bf3-8557-0007abc5c064 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Multimodal Large Language Models: A Survey
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c64fbc1-cfc5-4224-ba00-99c35c647e1c · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e06eb9-0d06-48ae-bf85-deb4ba1386b5 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations xdial-eval: A multilingual open-domain dialogue evaluation benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca8f1e77-8842-440d-a565-21323465c189 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e23151f7-29ff-40fe-9dbb-d849f8ae67e2 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalization of Large Language Models: A Survey
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e526ef-fff8-4464-832b-32f5844900cb · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations DiQAD: A benchmark dataset for open-domain dialogue quality assessment
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 791e650c-e121-426e-87d2-3373f14446e5 · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c921583-99e9-4d70-8560-649b95daa7cc · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Best Response
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c5c23a3-3e7f-4e88-9235-16ffc7f9185b · outbound
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · inbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a226cc-6e86-4bcb-af97-65ad555a9fe3 · inbound
Cat-DPO: Category-Adaptive Safety Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0113f25-a5ff-4824-a94c-1ed1365b7eda · inbound
A Survey on LLM-based Conversational User Simulation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18109681-da7c-4dc8-9741-66ce23d3569d · inbound
TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7364afbf-13de-463a-8464-87a822720ec4 · inbound
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfa57893-7f3d-4d52-96d5-3d73151bf6d4 · inbound
Agent Safety Is Action Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7a42867-3b8e-4f87-b3a1-afb6ee7d0ab8 · inbound
Benchmarking the Personalization Capabilities of Large Language Models A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.