Pith. sign in

Paper Citation Record · LEDGER

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

As of 12 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2412.00207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00207 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:41:03.002125Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:01:27.883371Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:45.977108Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36787a03-ef3a-44ab-909e-4b1b1c06b2c6 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Evaluating Large Language Models in Theory of Mind Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.904636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.904636Z digest=sha256:11bfd739e050a27c16b2c4446be3aff168a90d03254d1a4bb8423f53a1c9f0b9

Observation fe614e7a-2fcb-4573-911e-ad99908723ca · outbound

This paper cites UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:41:03.320070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.931581Z digest=sha256:62a6aab55657c163b5515e3b0ca7365a2741297d948c08cb408d770cec2a97f5

Observation 6f410a0d-28a9-46fb-b33e-89c72df2270b · outbound

This paper cites Rethinking Model Evaluation as Narrowing the Socio-Technical Gap.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.936689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.936689Z digest=sha256:1f62025d1b8b9333abb535599ff6cbc2f3931bb9e18343bae5d2b54d04d2bdd4

Observation 3756c180-8dff-4676-9723-cefb7ac9ec5a · outbound

This paper cites Personality-adapted multimodal dialogue system.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Personality-adapted multimodal dialogue system

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:41:03.282707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.942932Z digest=sha256:3b9c8e578d5493f179d78f073feda51bc87df0d9ed75013c7c596f9ef88b3166

Observation f7139515-c948-4f04-846f-9a29eec06e4c · outbound

This paper cites Jeongeon Park, Bryan Min, Xiaojuan Ma, and Juho Kim.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Jeongeon Park, Bryan Min, Xiaojuan Ma, and Juho Kim

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.948359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.948359Z digest=sha256:24bb9e27699794e9d43847a531b42675f7c570b041b760bccbd13253ab36d3b3

Observation 3d8c7d7c-a47c-453f-8bbb-6e0bde2641bd · outbound

This paper cites Personality Traits in Large Language Models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Personality Traits in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.962924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.962924Z digest=sha256:30e39beffc8f47a356896d9aa50965da5886ff1db8ecfa5c7f775b9c78701c21

Observation b80a8dae-7c40-47d5-a487-93b688367bd2 · outbound

This paper cites The next big five inventory (bfi-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots The next big five inventory (bfi-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.505038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.967851Z digest=sha256:3ad36970ec8f3d20875604c7a7e28e0f0895db18a0b4172284f7819718d09058

Observation 123dfebd-e3ba-417f-80ed-5aba894cfef8 · outbound

This paper cites Will the Real Linda Please Stand up...to Large Language Models? Examining the Representativeness Heuristic in LLMs.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Will the Real Linda Please Stand up...to Large Language Models? Examining the Representativeness Heuristic in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.977982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.977982Z digest=sha256:060acfdae36603cde7ebf725d0b62b30aebc49d0da98790aba53bdd8b639ec15

Observation 068dec37-78a4-4941-911a-32bd60fbea95 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.676.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots doi: 10.18653/v1/2023.emnlp-main.676

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.987172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.987172Z digest=sha256:5979c168c469bb384a96cdb6822d7d2a8fdcd33d13a6566b7d2e543bc493ada4

Observation 97c02d10-830d-4325-9bfe-c0ac901ff5d8 · outbound

This paper cites {personality description}.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots {personality description}

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T05:41:03.479806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.991910Z digest=sha256:e98b65ad604968464dbf9596130c4da5bd9c0e45ff4dc27fbceb355c5f53813b

Observation 0856b444-a09e-4266-a009-1c7e3bccac63 · outbound

This paper cites disagree strongly.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots disagree strongly

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.464981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.996740Z digest=sha256:82c87adc82fea8082de824271f25342593ad0a07ed33eed1d8cbd3e40acf6468

Observation 824a227b-6615-4f42-a0ea-f21fcd49f22f · outbound

This paper cites Table J details the specific instructions used, where transcript refers to the human- chatbot conversational scripts.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Table J details the specific instructions used, where transcript refers to the human- chatbot conversational scripts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.450407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:03.002125Z digest=sha256:7cab2e186772487c6c26ade76f481fa3677294486023a4655857e4c72793eaec

Observation ef3c1e2b-5243-4376-95c2-8e4db1206ddf · outbound

This paper cites Talebrush: Sketching stories with generative pretrained language models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Talebrush: Sketching stories with generative pretrained language models

Reference 1959

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.586869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.871741Z digest=sha256:e02928b92be5269e8b1afd8381a070e7b1cb4a190faa4daaa4c2ccd6541263fe

Observation 726dfc18-df73-44fb-8ec0-6a9437665e9d · outbound

This paper cites Vera Liao.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Vera Liao

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.982597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.982597Z digest=sha256:776ce6d0769f198aef7f313ecbf3f56c2ad92946bf92744861e3e650ed75cdb6

Observation c6c5a87a-51e0-48b1-9a76-8663a61615e0 · outbound

This paper cites Lewis R Goldberg et al.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Lewis R Goldberg et al

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.877406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.877406Z digest=sha256:238e206036be92946c5306d3e70a64cf99a04ede5181ecca5208d874354924dc

Observation 556b9a44-2a19-40b0-bb2f-b91b5d54c2c5 · outbound

This paper cites Who will go the extra mile? selecting organizational citizens with a personality-based structured job interview.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Who will go the extra mile? selecting organizational citizens with a personality-based structured job interview

Reference 1999

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.570617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.882534Z digest=sha256:11d560b668345e9668b4ef6ea315dd30105c8555720c29bae5eb1dc2fc3b9605

Observation cfb047bb-545f-4c18-8857-196d6351039b · outbound

This paper cites Behavioral change and consistency across contexts.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Behavioral change and consistency across contexts

Reference 2001

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.532629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.958572Z digest=sha256:b909d90b0282f424e7f54d097ddae772be147f8011108f3bf83ca32d399c8ebd

Observation 3d05eea0-ff1f-47f1-acc5-a3a534dadc71 · outbound

This paper cites Revisiting the Reliability of Psychological Scales on Large Language Models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Revisiting the Reliability of Psychological Scales on Large Language Models

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.893879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.893879Z digest=sha256:6cbc21254983c0f04b6e919dfb2b2987247788299612184d46045dbaa8c32d67

Observation 333533d2-29a9-4bb5-8291-ef8b93cd3023 · outbound

This paper cites Evaluating Human-Language Model Interaction.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Evaluating Human-Language Model Interaction

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.920337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.920337Z digest=sha256:e89aa7b6ee88435fb7310350bee22992e2563c383c4fbc58b3a41e751062e30f

Observation 1dad5904-dd25-4010-a710-ed5c86e2af45 · outbound

This paper cites Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.953254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.953254Z digest=sha256:456c673cc3bb57215a267ce748e3c4f4b995ce28dc9227be9c8f7f862a4352e5

Observation c8e103a3-09d0-442e-856a-27b66376ee4a · outbound

This paper cites CharacterChat: Learning towards Conversational AI with Personalized Social Support.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots CharacterChat: Learning towards Conversational AI with Personalized Social Support

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.972979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.972979Z digest=sha256:6def8d8846659b30480b87cdcc32acf8a4309657fa00282d0212f696b10bd4bb

Observation 9795cd0e-bb90-4f2b-b1b9-124eb638a0da · outbound

This paper cites AI-TA: Towards an Intelligent Question-Answer Teaching Assistant using Open-Source LLMs.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots AI-TA: Towards an Intelligent Question-Answer Teaching Assistant using Open-Source LLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.887989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.887989Z digest=sha256:a257de7279ea14b0397cacfc6c5bcff0b0b1552a60e92c1711b57bb5a6faaf03

Observation 705000ef-f3fb-4196-8a70-f7aa53ac8687 · outbound

This paper cites Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.925722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.925722Z digest=sha256:9666f4277520e974e7ea9b457951ec2d96f84a8d91f87b83cd6a3bcc9edd0881

Observation 8b744194-7571-4dfe-b54d-3ffc60e6f067 · outbound

This paper cites Psy-LLM: Scaling up Global Mental Health Psychological Services with AI-based Large Language Models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Psy-LLM: Scaling up Global Mental Health Psychological Services with AI-based Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.910069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.910069Z digest=sha256:935be3d0b39e5152e00fd62db6dc7356e68e647da99aa454a3d24bad5fbf66ca

Observation fb89bc62-8fc3-4c31-9bb0-6e213ab858e1 · outbound

This paper cites Construction and evaluation of a user experience questionnaire.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Construction and evaluation of a user experience questionnaire

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.550364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:41:02.915354Z digest=sha256:bcd78b883ec38f4b4ebc37e6a80b81108a8b30ef00518418b4c5311a0a6a78c8

Observation d3aa69a5-462c-448c-9694-4529bc229783 · outbound

This paper cites PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.899394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.899394Z digest=sha256:22b84a032d4867896e842d703fc6567144d6988e7bd85a145841c79ef9549352

Pith citing papers

Observation 06b9d794-77d8-46e7-a462-11c3cedf38c1 · inbound

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs cites this paper.

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.599671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T10:01:27.883371Z digest=sha256:3917a11ce65013bb8e17d04a081384eb94c20c5db7b19705145f6d8bc6fd2339

Observation e9a963d0-dc86-49a6-bf57-422a0464f753 · inbound

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior cites this paper.

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.736197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T09:36:36.821058Z digest=sha256:28406e7a1a25a7b2f2242886825b5130ad0a79d9128b9fc373a4826f876f2dd5

Observation bc28a1b3-6bd4-4e66-8b59-1a1c5909be0d · inbound

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure cites this paper.

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:45.978691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T08:10:46.067374Z digest=sha256:953c5867de7b6f660b9a2c92fff91919e6f6eeb5cabd994d454c7e35abb91514