Pith. sign in

Paper Citation Record · LEDGER

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

As of 5 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2601.18061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18061 v3

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T11:46:09.108129Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:18:54.595022Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T04:18:56.704885Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy57
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89705895-d623-41a1-9534-e24224803783 · outbound

This paper cites Clinician-Rated Severity of Nonsuicidal Self-Injury.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Clinician-Rated Severity of Nonsuicidal Self-Injury

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.788739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:71820e721a88dcb587a0fa7e1cd3abbb7a18dafaf72f94ec894d6be5ecf24934

Observation 0adde6b6-9e29-4ee8-8ce8-dee16e65a5ba · outbound

This paper cites DSM-5 Clinician-Rated Dimensions of Psychosis Symptom Severity.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing DSM-5 Clinician-Rated Dimensions of Psychosis Symptom Severity

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.802673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5e2b25b26e86d58c98378f44398a941cd961c04e78e0c27bd884d4551fa1a327

Observation 1e4bec4c-d8b2-4dc8-90c3-a88efc156446 · outbound

This paper cites DICES Dataset: Diversity in Conversational AI Evaluation for Safety.Advances in Neural Information Processing Systems, 36:53330–53342.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing DICES Dataset: Diversity in Conversational AI Evaluation for Safety.Advances in Neural Information Processing Systems, 36:53330–53342

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.797026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:23b23464ce1feb75210319a9a9bc268e5414a111510f2b3d545f09326b8b7f0d

Observation 05f8dcd9-93a8-4a4b-b3a1-d00b890fe36c · outbound

This paper cites Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation.AI Magazine, 36(1):15–24.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation.AI Magazine, 36(1):15–24

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.799885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:fc888bbb608313c1b5e05c5bbe25fa368ccc096389f9ddec897bcc4e2d63bf5c

Observation e9922f6e-93e3-40d3-b196-7528be7f4d82 · outbound

This paper cites Crowd Truth: Harnessing Disagreement in Crowdsourcing a Relation Extraction Gold Standard.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Crowd Truth: Harnessing Disagreement in Crowdsourcing a Relation Extraction Gold Standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.777957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:dced17fe3ed1a8444f4bf15c179491c8de4869aefd2f83343382835187a29ad3

Observation 4e6e9d3c-fcdb-419a-8271-a99241af5c35 · outbound

This paper cites Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.775559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:c0f1eedd878b9894e08c011a08f91309b0f2c45f0dbb28350e5255b7441bc403

Observation f2e61751-f922-45ee-9b08-979a55795f70 · outbound

This paper cites Bech.Clinical Psychometrics.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Bech.Clinical Psychometrics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.794378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5643649b787fc3ca0c56a7e7d897d73e4a45e8fef564356286afd165e6696942

Observation 0d376012-b765-4928-8354-102339370c21 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.791417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:be290d97f4de9639011d04b886d610c24f0354a1fd1bed7fe83ac713c39d2675

Observation 55997058-543a-4706-a361-f2c18a69573d · outbound

This paper cites Consensus report of the apa work group on neuroimaging markers of psychiatric disorders.Am Psychiatr Assoc.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Consensus report of the apa work group on neuroimaging markers of psychiatric disorders.Am Psychiatr Assoc

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.771018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:11dd69f0cea4a8dc801d05065832afdd3826826fced1b25b72d7433f6fb5a74c

Observation 940f40b0-d336-49c5-9f1a-d098a446facc · outbound

This paper cites Using Thematic Analysis in Psychology.Qualitative Research in Psychology, 3(2):77–101.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Using Thematic Analysis in Psychology.Qualitative Research in Psychology, 3(2):77–101

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.785997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:b75aede0957f45c2dd4eb972090c2d5766ac61852e71f8e76db5731c3e792821

Observation e77b00b4-e418-4355-9ed3-960f0ca0f5cd · outbound

This paper cites Minton, Abigail Lott, and Jinho D.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Minton, Abigail Lott, and Jinho D

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.780735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5830a74a99d9146c10f5096dce4829a1108f32611814c1d2890c5728728c54de

Observation 5c35b276-3251-4a8b-8e82-3ed483f4f883 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.692465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:6e0b246ad7d59699d4b364a7514f720517b1bbeb8bc1a3a2a2100f8adff6cf7f

Observation d5270066-a82f-4590-87f7-3912b967b006 · outbound

This paper cites How people use chatgpt.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing How people use chatgpt

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.641532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:01e03a7fee17ad611c3ea8239f81f1a58d5a128c93b8d9a197d1a618586517f8

Observation e7ba2d63-d3bf-49e4-ad8a-1ce5804a2c25 · outbound

This paper cites Predicting Depression via Social Media.International AAAI Conference on Web and Social Media, 7(1):128–137.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Predicting Depression via Social Media.International AAAI Conference on Web and Social Media, 7(1):128–137

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.653085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e8e3b857e6c0b915cb70f72261eb6928d13f902c6e45a9cb8e2ebfc8acc36ac9

Observation 28ab72d8-fa4f-4ac1-9997-2d9a277ac1fd · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Deep Reinforcement Learning from Human Preferences

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.650228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5b8920c8bcdc394874749b13fff93f104e95df8ffce8b3ae0dfec2f7087da99a

Observation 4e60b9a8-85f4-4021-bc97-fbed1d805f31 · outbound

This paper cites Cicchetti.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Cicchetti

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.661356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:1f3e9b5c1e721ea672e3186f0950b6ca09a2545e0d8316f3463a96cfbd3df7d5

Observation 5f765e66-41fa-4f04-b2a7-08f9b4020713 · outbound

This paper cites Hashimoto.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Hashimoto

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.658470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e54e041bb781cac9988fbf6a6ef62fad563a6a8b947b618a86d39934614e872c

Observation 793adfb1-7bd6-409d-8b4c-9fad441839d2 · outbound

This paper cites Diagnostic and statistical manual of mental disorders.Am Psychiatric Assoc, 21(21):591–643.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Diagnostic and statistical manual of mental disorders.Am Psychiatric Assoc, 21(21):591–643

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.735587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:6afcbaff75525a6834b182fe822c18f1dc899357dcb45863681419f2066bfd07

Observation 792b627f-be36-4f30-ae29-b1e0dac0468b · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-05-16T11:47:49.282324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:402d1c287adc5ce85a6a8bbad8c86db85411552252cef4578bb6cd74b1661481

Observation 212b9a9d-18d0-4aca-8802-7d3bbb48e4a0 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.624465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:358f0d07613ffa7af3d57eb04874128b8a07d1d4870099a0db02933859aae7e7

Observation 598aa3da-0aeb-4af6-9d2c-1e2b4133470f · outbound

This paper cites Can AI relate: Testing large language model response for mental health support.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Can AI relate: Testing large language model response for mental health support

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.633249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:68fbe3576adedf5894ecb5e4591b93978f01344ea4f8ed29ca8b1748fc021f8f

Observation b3d75980-8685-4f1f-b133-7e00f5295efc · outbound

This paper cites Impact of preference noise on the alignment performance of generative language models.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Impact of preference noise on the alignment performance of generative language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.704854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:360d3e045d85634a1067d6cb74ccab5dba8750058fac1d990c9980a158a6e3d2

Observation 942edd00-ba09-4dff-8d65-bcc7949576ee · outbound

This paper cites Blind spots and biases: Exploring the role of annotator cognitive biases in NLP.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Blind spots and biases: Exploring the role of annotator cognitive biases in NLP

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.619087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:2104020fbd7219726e24c95f16f57e58d318801777c0674a24389e301a24f438

Observation 26db7dd9-5be6-4967-b51f-cec76322a12d · outbound

This paper cites Goodman, Lawrence H.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Goodman, Lawrence H

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.639057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:1d29c3626b6904b54c7f1487599b96ff81ba22603bced7bf8d74a138d9f22055

Observation 027577f7-8577-4fcc-9c3a-e8e95fbf536f · outbound

This paper cites Gordon, Michelle S.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Gordon, Michelle S

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.723443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:05ac17a1810ea4c53c3a184b4a5a7ddfd5a8b50f288062d38532084a95be0b1c

Observation 38e14316-48b9-4210-b0e8-dcd5e5fc5685 · outbound

This paper cites Risks from language models for automated mental healthcare: Ethics and structure for implementation.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Risks from language models for automated mental healthcare: Ethics and structure for implementation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.768121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:0f97165c4aa1e223630b3af5e947e76e5e7096251ee58b3da2e557d395885ae1

Observation f73d4089-4f83-4d33-ab31-a800c9d1bae7 · outbound

This paper cites Human Feedback is not Gold Standard.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Human Feedback is not Gold Standard

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.689018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:375097cc2d189fb8e8c8992d02c6971b65eb0854075e12b3ae575ee107b37e23

Observation 98fac7e5-2c51-4c5d-a5ba-31b4f65ba2b9 · outbound

This paper cites How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.783275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:aff051c7780864b5c60d6a2ec22ef97cc1542e7dd95cbc36c8fba702e5239ace

Observation 33cf71bc-79df-4819-ada5-1bedf766dccc · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.670029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e0f7f3aaf5e32d9d61c59e26a5d238665127b04558d15a0a0cce46282cd83eb6

Observation 428a34c9-6494-48b3-bc1d-ece9f3437232 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.701487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:0b595fc666239977ac0be084bedad00a41f05d2fec30ba052c66f23bfe967875

Observation f48788c3-4db0-498f-addc-942c10481929 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.683843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:108579088c70ddd3948e3d574a540e7337a549fa93ae527b09ef01f93dc63cf6

Observation 65cf5188-ab62-47c4-a72b-a9555075c5e4 · outbound

This paper cites Reliability in Content Analysis: Some Common Misconceptions and Recommendations.Human Communication Research, 30(3):411–433.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Reliability in Content Analysis: Some Common Misconceptions and Recommendations.Human Communication Research, 30(3):411–433

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.655706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ed8227ae2d77d3976ad89b54a976f4c211d3c10b652010886e86d8b4a6d2bf20

Observation 4f915837-2483-460e-bffe-798b41f4a387 · outbound

This paper cites Kunstman, Aaron Lulla, Monika Drummond Roots, Manu Sharma, Aryan Shrivastava, Nina Vasan, and Colleen Waickman.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Kunstman, Aaron Lulla, Monika Drummond Roots, Manu Sharma, Aryan Shrivastava, Nina Vasan, and Colleen Waickman

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.698419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:3a758c1b5f774608435ec809a04143aea0f4e0f7eb2ea4680f596292a4881998

Observation 257c5123-393b-483e-abc0-32495fe4c450 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.636158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:fb317773b5d4ec51c2b09695874b7dc40d0e3a1529141bd79190c7b11043744b

Observation 8d28453c-fef4-4410-a4be-afee8f1fde6f · outbound

This paper cites Hashimoto.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Hashimoto

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.664401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d33655a80e05b0e91d99ef7168057835828666fe96499ea637ad4f29d9e28c4b

Observation 4731ff42-453d-448c-b60a-aea94481133b · outbound

This paper cites Bunyi, Adam C.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Bunyi, Adam C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.627248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ec54f78b70a50938e7179afe55ccace176e65bba08c854c79ea7bac841b43852

Observation c6897c9c-874c-422a-8468-6affd4168880 · outbound

This paper cites Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study.Journal of Medical Internet Research AI, 3:e52095.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study.Journal of Medical Internet Research AI, 3:e52095

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.681405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5d3b20729fd3ea9bc5b1541cd4ed14fa13569bc52a6ccb84ae8d7c5823ebf6c4

Observation bfae3754-10ce-452d-9296-72d1725222a7 · outbound

This paper cites A diagnostic meta-analysis of the patient health questionnaire-9 (phq-9) algorithm scoring method as a screen for depression.General hospital psychiatry, 37(1):67–75.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing A diagnostic meta-analysis of the patient health questionnaire-9 (phq-9) algorithm scoring method as a screen for depression.General hospital psychiatry, 37(1):67–75

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.616428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:2cbcae8767d533f9a776befc05112e3f0cb02792285f7a4b9703d0be5e1af844

Observation fb3125eb-1f42-49fc-884e-8dea0af2abc7 · outbound

This paper cites McGraw and S.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing McGraw and S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.678285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:db9d4c3038cf3dec5bfcbd589524867f25406b27fba39cb3744f470a8b12efb1

Observation fc5bb067-0f13-4b44-9a31-6fe731312cf2 · outbound

This paper cites Ong, and Nick Haber.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Ong, and Nick Haber

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.686452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:7cd166efcc3a31ce4b92ec9dbfe64ee86c2fa4beaf8096a5b4930f15f33fa897

Observation 4d351919-2dd3-4888-99f0-6bac4c34a7b2 · outbound

This paper cites Moyers, Lauren N.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Moyers, Lauren N

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.647305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:f810981fa938658718ae0f56a74d98e26c397de4d57d9e013e6913d0ff568e22

Observation 59346976-8a19-4392-a242-d53521eaf689 · outbound

This paper cites Department of Veterans Affairs.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Department of Veterans Affairs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.739563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:82ca4f0effd10c6bd4fbfe24b4a8dc4dd3aafdfb6bdd24b405fa3d08c94c852e

Observation 7b07b7ee-fa44-4bde-be7d-1099d2d3f527 · outbound

This paper cites NICHQ Vanderbilt Assessment Scales.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing NICHQ Vanderbilt Assessment Scales

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.720304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:18c288b8c16fd09de1b7c1a9c94c07d1678ec31c90cc821e1ac47ba9366e4658

Observation 98ca2986-71ec-40ee-8239-f70dbe2cf385 · outbound

This paper cites Depression.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Depression

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.695200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:3f49fd4f370a072eb2f9f8a4bbd53b15b65572de7255e8226bec8db3b5a090b4

Observation fff6dc70-4901-4d49-8b6a-d778979de163 · outbound

This paper cites Enhancing mental health with artificial intelligence: Current trends and future prospects.Journal of medicine, surgery, and public health, 3:100099.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Enhancing mental health with artificial intelligence: Current trends and future prospects.Journal of medicine, surgery, and public health, 3:100099

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.732269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ebb49def3d0755c2df5f6ca5c1bc1ce5785055d47b796b9c62d160f724e98d30

Observation 00113e04-ee5f-4e6f-ab52-b7563cfa53ba · outbound

This paper cites Christiano, Jan Leike, and Ryan Lowe.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Christiano, Jan Leike, and Ryan Lowe

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.675515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:00de95edd6009d5dc72a6b173f5c36dff27575c7f8b185253f57f7fabec642ba

Observation d4b3fcbe-96fb-417e-9b84-6d5f98ee1736 · outbound

This paper cites Inherent Disagreements in Human Textual Inferences.Transactions of the Association for Computational Linguistics, 7:677–694.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Inherent Disagreements in Human Textual Inferences.Transactions of the Association for Computational Linguistics, 7:677–694

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.644545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:aa55a16458b03e79262713ef54762464c43aae8886e460cdc553dbdf0f4c7333

Observation 70a2be2c-e2e5-49c6-b176-fb9157ec16cd · outbound

This paper cites Red Teaming Language Models with Language Models.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Red Teaming Language Models with Language Models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.672570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:7013b40118d990ca1e5c8814f393dbb343ac0ce8899d174335a7af88996d609e

Observation 7ad7fdb4-d134-4f39-a202-dfbe0a93e440 · outbound

This paper cites Posner, D.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Posner, D

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.630361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d27233946fd594f5012cb670183908edb922140a48c4333bc0355bd013d2195f

Observation ba9f5582-3a8e-4bf6-ba85-08a669a79415 · outbound

This paper cites Prochaska, Erin A.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Prochaska, Erin A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.729386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5686d8adb7574527e790e251893f26f3a5a5ae736a941aaed8ba39bc473a0298

Observation e65ef78c-0a0e-4f46-a5cf-4677acdf2c8a · outbound

This paper cites Manning, and Chelsea Finn.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Manning, and Chelsea Finn

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.621831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:2073428facbcdd3709bb9935704c1214f19fb7078ec8202ea3c556ec4305d932

Observation b76de69a-9338-4b1e-b092-3f5d05798da4 · outbound

This paper cites Regier, William E.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Regier, William E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.742505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:9b0f37527a3acbfdbba208d342dabde685b2c361085590e05d51fd4077612266

Observation 04ef73ce-2388-4902-93be-9bdb52f0d312 · outbound

This paper cites Large language models as mental health resources: Patterns of use in the united states.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Large language models as mental health resources: Patterns of use in the united states

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.711024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:cb28962cbec686913970c22ce6e96b448a5677af4a71f41d147fb2c5867ea44e

Observation 98307c84-6c72-4814-8ecf-597ffa8652cd · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.745311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:af94d19db3ca347fd5fb0106b743548e8028880487bd55a2622ee37a18fa8fef

Observation 50d4e86b-a636-44a0-a658-f91e3e379ab1 · outbound

This paper cites Lin, Adam S.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Lin, Adam S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.762597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:98ab24bca17dcbcb54dbf802229b9fcc8da23f78c7d421eeea4ece3313fb8017

Observation db133a62-4faa-41a7-82f7-d4b4065e615f · outbound

This paper cites A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.667265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:91e377fd023d9c0d0dc3cf914cbdc59876fd40ff7bc2c747803ed0c3236d9e39

Observation 752050e8-e4d8-4045-8ed3-b4cbd6a81428 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.707700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:4a0912ec1823d5f3d970b1da0949edf1a394f999a814a95573b8416637f86e9d

Observation 54393a32-3d9c-4bc0-aca0-aa974f33a79c · outbound

This paper cites Clinical Practice Guidelines on using artificial intelligence and gadgets for mental health and well-being.Indian Journal of Psychiatry, 66(Suppl 2):S414–S419.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Clinical Practice Guidelines on using artificial intelligence and gadgets for mental health and well-being.Indian Journal of Psychiatry, 66(Suppl 2):S414–S419

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.726521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e14362799955a9f1dca21ec27ad761735fa448bc7a24b2217ab1d35a723d6de0

Observation 6ab45c2b-2947-46ac-9713-ec7a1e4f1c37 · outbound

This paper cites Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.757209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:325dcef84bf26c407cf1214b8e72a21d6069e6b62fdf635038eb5ec5d6486064

Observation f99fe2f6-d891-48eb-958a-008a368e67a1 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.717393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:0e8e4d46a1e2bbfb8c9d76e9d0ad0d3525e57bb72a5959a33461f911852533ef

Observation b4f3a450-f6f0-4bea-ae9a-071ef5658955 · outbound

This paper cites A Practical Guide to Fine-Tuning Language Models with Limited Data.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing A Practical Guide to Fine-Tuning Language Models with Limited Data

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.714235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d59bb6c929616b36fe4843704c42a7f87b548bc3c4270470bd76376c5659f227

Observation d00a5e5b-2d99-4f8a-987b-5806e4dab1c6 · outbound

This paper cites Lukoff, Keith Nuechterlein, R.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Lukoff, Keith Nuechterlein, R

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.754158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:a25005a6c1c7a5cfa50084dc5a403edce772b862d98a67968cc91a13752e9549

Observation 0d9e7534-5b1d-422a-a024-16efb4896549 · outbound

This paper cites Wang, Patricia Berglund, Mark Olfson, Harold A.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Wang, Patricia Berglund, Mark Olfson, Harold A

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.759924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ee2991ad81e9cb1d5ef93b8382646b0875fe7b1026e9e8676bb5921908304137

Observation 750f9970-3b04-4c32-a787-ab37a9b73f95 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.765353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:7b3ce62d2413f91e752e19f7877e5e9a2086ee38dd0c8d3e504818d4cec1d0ac

Observation e29dff42-8874-474f-8a6e-480b1e3c836b · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Xing, Hao Zhang, Joseph E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.748254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:5e6cc4aad30842ee0fe4de8770b9a2731c8335dd5a329f2dd0c3f3c9f92df3ca

Observation 6c6a7069-3c34-4aa4-80cd-17125693c22f · outbound

This paper cites Cold plunges cure psychosis—stop your medication.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Cold plunges cure psychosis—stop your medication

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.751327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:671da6696cc86eeb864a96eef689ec027f9f9dbced4cd606efb4e7acc66394d2

Pith citing papers

Observation cdf42a1e-2a34-4555-8b2b-3949d90c5286 · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:18:56.830854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:18:54.595022Z digest=sha256:5de2b5fc55f5c8fbbd77625d6cdebc1ecf7ef383e09a7c8a956883ee284a2168