Pith. sign in

Paper Citation Record · LEDGER

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2505.14106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14106 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:23.733585Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:26.701496Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:34.994786Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved19
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9861fdb-99de-45c2-8df0-515c31a77adb · outbound

This paper cites Conversational Health Agents: A Personalized LLM-Powered Agent Framework.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Conversational Health Agents: A Personalized LLM-Powered Agent Framework

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:18.519445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:18.519445Z digest=sha256:9651965bc7112b735bbe57356867a0f36f13386c7c10887afb183a1755a0617f

Observation e00bf5c9-bc0c-47b8-b2f7-cd2eb10dc9d3 · outbound

This paper cites Persobench: Benchmarking personalized response generation in large language models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persobench: Benchmarking personalized response generation in large language models, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.054834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:18.625291Z digest=sha256:feb6be4225759d255c2610116d19e49b6ef040e0ae02faa42ffafd4ca908c441

Observation 03167942-a7e7-48c2-bf1e-e6ae533ea020 · outbound

This paper cites Claude 3.5 sonnet.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.909324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:18.729628Z digest=sha256:2031e5d71d45a56fc3a051e9a7f9d187008922097f5c270ebe6a0f722eab1707

Observation 237bf522-348b-4651-a528-b532151c0119 · outbound

This paper cites A Little Human Data Goes A Long Way.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Little Human Data Goes A Long Way

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:25.276669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:18.993776Z digest=sha256:da47cb1edebacb87a179a8742c888191dd53931baeaaf1606b2c64af52312e69

Observation c3e01726-ba8e-4cf9-b027-26cfc2176d36 · outbound

This paper cites Personalized Graph-Based Retrieval for Large Language Models.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Graph-Based Retrieval for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.113061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.113061Z digest=sha256:cdbb8ce43e1a39349d826adc2dfc1e65a87fcf11cd2a1fde96ca27f1c06663e1

Observation 03232c27-8236-4033-b7d8-f3d2f54d19af · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.244719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.244719Z digest=sha256:e276cfa21d3d4cc22ecc9f357f8cf078737b1ae3a473ca7b02a23307658d58d9

Observation a903f797-26e0-48dc-b27f-1e35298c32d9 · outbound

This paper cites LoRe: Personalizing LLMs via Low-Rank Reward Modeling.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations LoRe: Personalizing LLMs via Low-Rank Reward Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.367275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.367275Z digest=sha256:79a1cbdbe995d275daf75011cd630da842bb2552c5cbcb204c2706e12f4cbf14

Observation a8722fdf-21c2-468e-b36a-a02ebb7c345c · outbound

This paper cites Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Optimal classifier for imbalanced data using matthews correlation coefficient metric.PloS one, 12(6):e0177678, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.490084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:19.513223Z digest=sha256:c73d21bdaa5e66c4a8b10e5626bad65b89926e46c46699adc4cd17ee93ca2e83

Observation 74b080f6-b2ed-42c0-a046-72636d8e93b4 · outbound

This paper cites Beyond prompts: Dy- namic conversational benchmarking of large language models.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Beyond prompts: Dy- namic conversational benchmarking of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.268054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:19.655393Z digest=sha256:8666b7edd820c1106ff565de682f46501e1597635b7b2e01c28fceac9f8f3f23

Observation f3e6934d-15cb-4155-9c60-ad8fd8a2b49f · outbound

This paper cites Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae).Geoscientific model development discussions, 7(1):1525–1534, 2014

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:34.075713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:19.807578Z digest=sha256:21b57eaa6ca0c5988f3d9b4660d890650850311cd9e0d57f90673c4b82b9cfaa

Observation 2ef064e0-e58f-4565-9410-fc52a2f1045c · outbound

This paper cites When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations When large language models meet personalization: Perspectives of challenges and opportunities.World Wide Web, 27(4):42, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:19.894818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:19.894818Z digest=sha256:e83a234c9577aa5274a2ce1becbb19842e572e450bc374ff36253d47b49bfad7

Observation 3cc7d673-0b61-4547-aa25-30019abbd5a8 · outbound

This paper cites REALM: A Dataset of Real-World LLM Use Cases.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations REALM: A Dataset of Real-World LLM Use Cases

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:24.811302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.033370Z digest=sha256:23e813092bb7641dc4b955599ea7691d9e84e3e89629209070b2df278b3433b4

Observation 4e070871-557d-43a0-adea-d59997ee4ab3 · outbound

This paper cites The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation.BMC genomics, 21:1–13, 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.909265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.144923Z digest=sha256:bfbff71ba6969de39d404c5a9b72b4ede2ebc50927061fec47a5937399902dad

Observation b06db013-27d0-4816-9cb6-68ce7003e4d0 · outbound

This paper cites The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.761772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.324986Z digest=sha256:4fd21e7a5b5779617e2e3d2b7deca3aa42eb091d6029d3f4ad25267e9ef1dd2a

Observation 88126155-e729-4f8f-8fc5-f140559fa286 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.595175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.465431Z digest=sha256:b448265ec962ee2588e6a934356578e2445ff6ab06bc0f43a651527b50b636dd

Observation c48a18f2-6a9e-40d5-b11d-dd9a31c50a36 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations RedCaps: web-curated image-text data created by the people, for the people

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:20.580260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:20.580260Z digest=sha256:44606bb8d84a3b503b6477199e50c7bf903a3f50ad816565e727e6ef55f95fd9

Observation 9a06c4d6-7fac-4a12-a0aa-9fc9e604ac27 · outbound

This paper cites The llama 3 herd of models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations The llama 3 herd of models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:20.662819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:20.662819Z digest=sha256:4833a023471a84ae9b4895fd2a489dd2ee133fcc9212e148c67197a896ea8d65

Observation 850ceb35-110c-4dc9-9b9a-a27a373e84d2 · outbound

This paper cites Ruddit: Norms of Offensiveness for English Reddit Comments.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ruddit: Norms of Offensiveness for English Reddit Comments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:24.552148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.731343Z digest=sha256:dcf59fdee33826d05df9063f1f16b9e38a9f3852c432ae78e023c41901789976

Observation 194d8e5f-4214-4e7e-bfc7-0a4e1ee518eb · outbound

This paper cites Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Root mean square error (rmse) or mean absolute error (mae): When to use them or not.Geoscientific Model Development Discussions, 2022:1–10, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.382963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.867591Z digest=sha256:306eb35a321c5edb6c037302a124a9737fc3f0b6d1aa940b59dca7d84bb89add

Observation e4f4b3bd-d184-45bd-9946-d970bf281d4c · outbound

This paper cites Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:33.165335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:20.984363Z digest=sha256:e56ca3d06648fee3f0c710c1db0b15c651378b7e23a65d5cb58698c110e34086

Observation 570f0968-3ae8-4466-8fbe-45f11d61adc4 · outbound

This paper cites Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Mt-eval: A multi-turn capabilities evaluation benchmark for large language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.936888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.117526Z digest=sha256:792733cbf79ac867aa18b62b2a851e487e2bcb1aec7f84a86fc3b6d4674b37ca

Observation b6daae78-9455-4ebb-ac01-1c48db55eb46 · outbound

This paper cites A framework for building adaptive intelligent virtual assistants.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A framework for building adaptive intelligent virtual assistants

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.672196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.243210Z digest=sha256:a0acfe5d129ca462e08d8e198b2fde9f570d3bd589e3d7b3e2e02d0aca33c457

Observation 77ecc1ea-0684-4054-b46c-9b666cf43b3c · outbound

This paper cites Teach LLMs to Personalize -- An Approach inspired by Writing Education.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Teach LLMs to Personalize -- An Approach inspired by Writing Education

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.349336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.349336Z digest=sha256:72a1d0ac26bc5093c58dd527a0b366f41be426bc6f34e184ce15f956cc907837

Observation 1945642b-fd5f-49bf-a9bd-77d49420dab2 · outbound

This paper cites Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Panoptic scene graph generation with semantics-prototype learning.Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3145–3153, Mar

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.437751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.487031Z digest=sha256:16cc3691509b498af0fe03c6a3c9af53ab2f0d378f1623823a806b9e51451c03

Observation 62133eab-494d-409c-9d3b-59c8a7b733aa · outbound

This paper cites Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Artificial intelligence in intelligent tutor- ing systems toward sustainable education: a systematic review.Smart Learning Environments, 10(1):41, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:32.240242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.584797Z digest=sha256:376c58d4c56ca0834a9080553dd0a27cddff1e60bdd3e85abc5b27b2a0e26936

Observation 955fdb45-569b-4517-b7be-6a5169debecc · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Rouge: A package for automatic evaluation of summaries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.651056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.651056Z digest=sha256:18d1a2a7c32fe5991f413bdd042c405f58c979bd885dea11f621635c7473e100

Observation 646fc09e-b860-47e2-a683-92459fac4bb5 · outbound

This paper cites Persona-sq: A personalized suggested question generation framework for real-world documents, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Persona-sq: A personalized suggested question generation framework for real-world documents, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.971000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.760648Z digest=sha256:f41b17d0d1a0baa44619b469b0c7e6285bef7a9d8764a96b63b47a45858e3c17

Observation 9533affa-216e-4632-8ac6-64befd694cc9 · outbound

This paper cites Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Soda-eval: Open-domain dialogue evaluation in the age of llms, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.829533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.906459Z digest=sha256:c1c6045b191a360a80ea92ef5043f4e1c9edb5621b9d28cc1c19f423b560d832

Observation a054c5ee-571e-44a1-bd79-40f20b53da15 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Gpt-4o mini: advancing cost-efficient intelligence

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.643229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:21.975869Z digest=sha256:5d9dcdd9674debcb7f6b2fcad129f5c6b51d0badc1a3f91bdcf5a5ad1181870a

Observation 6ac76a55-64bc-42c0-96ca-c0cfd0f0339c · outbound

This paper cites Introducing gpt-4.1 in the api.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Introducing gpt-4.1 in the api

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:31.480886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:22.043964Z digest=sha256:9166d45b4b14eeb132409097058171cf67b023f79950b647e24a40e9380fad39

Observation 9d69a277-3a48-4798-855c-a90eab1feb4e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Bleu: a method for automatic evaluation of machine translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.172514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.172514Z digest=sha256:faa5dcd9fab322fdaf3960421479a71554082fa80313f68deb7b83e88e337375

Observation 65e016f6-564b-4426-8085-d6a390ff8535 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:31.250804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:22.264169Z digest=sha256:be1e833d99ed7a26cd46dc8e260c6d592157256b07184bf830f42cca432d4255

Observation 0fed4877-ccbd-4a02-9aee-37490e77aec9 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:30.988064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:22.394518Z digest=sha256:c4ba56acd0945e7b1c94e8346378c70488b6d39cf2542951329f2b932186f645

Observation e44488dd-ecf2-4aaa-83fb-0f4bdac730a5 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.486570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.486570Z digest=sha256:1b55370a149d8d714dc47d7f5e329d687e97036f1172178da78577d1a0268d80

Observation e08ac11f-6f2a-4ef2-bcd4-b722c05ab059 · outbound

This paper cites Lamp: When large language models meet personalization, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Lamp: When large language models meet personalization, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.780537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:22.581458Z digest=sha256:3d72812d2cb882859880a75ab38a77fb951d231370fce521eaa4aa2e81bb12aa

Observation 639f9236-0fca-4285-9c56-1765c052d4e5 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.649086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.649086Z digest=sha256:cd58a162f45f70782edb92a22b9e4ee27e6a791341ab4efc2ed6177a77950400

Observation dbf2d392-76b7-4210-bacf-45ab18e9a5db · outbound

This paper cites Democra- tizing large language models via personalized parameter-efficient fine-tuning.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Democra- tizing large language models via personalized parameter-efficient fine-tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.623584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:22.762396Z digest=sha256:39c7684b5f87919fbffb9a211de71e13d111af9f08e708aedb86a690bdfff79b

Observation 2d87413e-c6d6-43ab-9c83-59a6cddc2c4e · outbound

This paper cites An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations An ai-based decision support system for predicting mental health disorders.Information Systems Frontiers, 25(3):1261–1276, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.404821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:22.887964Z digest=sha256:7e966c44655ca513bd9bef897cd04b908381ce857d6a62da7469ca7d9caecb61

Observation 6b15c0d2-b97b-46a3-abd4-7495a23729a5 · outbound

This paper cites Position: Will we run out of data? limits of llm scaling based on human-generated data.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Position: Will we run out of data? limits of llm scaling based on human-generated data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.943347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.943347Z digest=sha256:26dc892f2906ef18562086950f3123d88da485ada9d0176e9f89db253d8e1483

Observation 2d03cefe-d484-4bf3-8557-0007abc5c064 · outbound

This paper cites Personalized Multimodal Large Language Models: A Survey.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalized Multimodal Large Language Models: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.001354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.001354Z digest=sha256:3df9d18673fd7d60929f624646709bef663e57f9125665ef6f5af8a2a5f5fbce

Observation 0c64fbc1-cfc5-4224-ba00-99c35c647e1c · outbound

This paper cites A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.118282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.118282Z digest=sha256:7daede8aa4260af6b6b2d2903dda9e6a3f217d80cc5753f7bf40d3619bb62616

Observation f5e06eb9-0d06-48ae-bf85-deb4ba1386b5 · outbound

This paper cites xdial-eval: A multilingual open-domain dialogue evaluation benchmark.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations xdial-eval: A multilingual open-domain dialogue evaluation benchmark

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:30.185319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:23.196806Z digest=sha256:a607b1bd80ec4a96d6d5179f3a85f5d5e5818c0d411dc04444a410e4df399f3b

Observation ca8f1e77-8842-440d-a565-21323465c189 · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:43:25.986776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:23.314664Z digest=sha256:b3712af02c727ac36bcc0b2e9ba886490345cec012258188c0ac919f65918e43

Observation e23151f7-29ff-40fe-9dbb-d849f8ae67e2 · outbound

This paper cites Personalization of Large Language Models: A Survey.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Personalization of Large Language Models: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:23.443293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:23.443293Z digest=sha256:94ebdfe006c24e2de3c9eab4bae2f70b2e7f630f2ff413b71950a3981944faaf

Observation a9e526ef-fff8-4464-832b-32f5844900cb · outbound

This paper cites DiQAD: A benchmark dataset for open-domain dialogue quality assessment.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations DiQAD: A benchmark dataset for open-domain dialogue quality assessment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:25.739429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:23.527996Z digest=sha256:072488dd4a69181270ca9245b2de87c004fc7c374d241a391d400b76105d8e21

Observation 791e650c-e121-426e-87d2-3373f14446e5 · outbound

This paper cites Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:25.537365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:23.662353Z digest=sha256:56e4ca2679c66a5f179c81c962c20e13ea926684e2409fdf66039e62e3182a7e

Observation 7c921583-99e9-4d70-8560-649b95daa7cc · outbound

This paper cites Best Response.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Best Response

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:43:24.106045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:23.733585Z digest=sha256:bc167206804f54622a7be75f856a4de77b8361bc5c3430ed55afa424c5c611e0

Observation 3c5c23a3-3e7f-4e88-9235-16ffc7f9185b · outbound

This paper cites an unresolved cited work.

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:43:34.736747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:43:18.862204Z digest=sha256:13bebb5437cc03dcd2a1359cc7a4601eb6cc82672e5f468dab3785e38941accf

Pith citing papers

Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.701496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.701496Z digest=sha256:a548c666990630c9958e06d0e0178fc4545d406c2a40af6ef45407abf43824e1

Observation 84a226cc-6e86-4bcb-af97-65ad555a9fe3 · inbound

Cat-DPO: Category-Adaptive Safety Alignment cites this paper.

Cat-DPO: Category-Adaptive Safety Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.741205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:33:32.642379Z digest=sha256:3052ae75473babe573249ec1a67c9f490c989ec4c8406ef10327210ce24eeec6

Observation e0113f25-a5ff-4824-a94c-1ed1365b7eda · inbound

A Survey on LLM-based Conversational User Simulation cites this paper.

A Survey on LLM-based Conversational User Simulation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:11:12.091734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:21:09.118243Z digest=sha256:ead46371aedcbfb9bb00da86f3def123e3623f12aa99a0908494d473ad87b602

Observation 18109681-da7c-4dc8-9741-66ce23d3569d · inbound

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation cites this paper.

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.182424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T12:54:56.401350Z digest=sha256:ffa672563f3c60f28fd371ce3a22e396ae05150de5a72d630feb5bd90243c5d0

Observation 7364afbf-13de-463a-8464-87a822720ec4 · inbound

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking cites this paper.

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:52:52.347650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T19:51:45.657658Z digest=sha256:8919530b2bd9893df2e4b138e8584cb1db77ccfc310191fc83c549ff2432fcb0

Observation bfa57893-7f3d-4d52-96d5-3d73151bf6d4 · inbound

Agent Safety Is Action Alignment cites this paper.

Agent Safety Is Action Alignment A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.996301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:50:45.759936Z digest=sha256:c4bb341ec5b58e757063350dfd4d07284e6089f9891730214717e3d67a20b8d6

Observation b7a42867-3b8e-4f87-b3a1-afb6ee7d0ab8 · inbound

Benchmarking the Personalization Capabilities of Large Language Models cites this paper.

Benchmarking the Personalization Capabilities of Large Language Models A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T13:17:09.892437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:17:09.892437Z digest=sha256:661bfab97bbc4aa6f66e65eb39791de94946be3bacf9a8e6c54f3d712050c821