Pith. sign in

Paper Citation Record · LEDGER

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2505.17095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17095 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:22.701010Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 545e9168-9c1c-4fa1-a6d6-2160340b64ce · outbound

This paper cites online" 'onlinestring :=.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.545896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.545896Z digest=sha256:596e54f9bc5e61203465e3cd6f3faed581b090276b7dfb64efe63ebfb7dfe3d1

Observation 3e199913-0c7e-493e-b552-1f73826b8e57 · outbound

This paper cites write newline.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.676307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.676307Z digest=sha256:8814b4f671cdd8e61cd4eeddb37562bd6505b47629eca74b37eccd90d65de681

Observation 6f031bff-9e47-4538-8380-c782f87652aa · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Non-Determinism of "Deterministic" LLM Settings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.784104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.784104Z digest=sha256:a87921fd5558a8e36b96ba56f9a02b7c57e2047cd3de8a3f13f6a0e46eb20f18

Observation cb749c51-d812-4432-8d9b-66baad453fd4 · outbound

This paper cites Rahmani, and Youlin Li.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Rahmani, and Youlin Li

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T15:28:26.006630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:18.916189Z digest=sha256:00b47eff767986a1dc184905aea6fa83097f6f1c4db9b5ce227b929dc5699834

Observation 35ad88a8-a0a5-4022-af59-67259a56e4ab · outbound

This paper cites Sebire, Saleh Khalil, Elham Asgari, Christopher Tan, Andrew Taylor, and Dominic Pimenta.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Sebire, Saleh Khalil, Elham Asgari, Christopher Tan, Andrew Taylor, and Dominic Pimenta

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.025921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.025921Z digest=sha256:6cedfb26ed97900dfc3fabd8d3e7501087387e8ce315861495bc0e6e6d274ebb

Observation f0354ff8-c2e3-4b61-9a13-cabf543b041e · outbound

This paper cites Prompt Stability Scoring for Text Annotation with Large Language Models.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Prompt Stability Scoring for Text Annotation with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.116860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.116860Z digest=sha256:5f9375107624901b8a2ff4797bab117b82fef16ebdb7f94ba973226d4e2ddea0

Observation 03e9cc4a-01bd-4cd9-973f-6860447e380e · outbound

This paper cites Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.270196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.270196Z digest=sha256:5f08b7ff78ce0097fd5a843813cfbfdc84cd90f0c16e47f075be7f8a8ccd075a

Observation b5c38088-02c5-4497-a132-c5fa265f4781 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.789786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.350669Z digest=sha256:43a3f083876eb9064a9b44dee0813a5f367b092bb89c41f59224feb24d00ef63

Observation 4762bb75-b517-48c9-a52f-b6a177d21600 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.419917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.419917Z digest=sha256:0e6a3a22d93d18b59d9db92b4378d7a2d87bd95d3cbcc808447ea4c69d93c044

Observation 016cd5c8-f737-4163-8f59-2481fb8ccbc1 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.474288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.474288Z digest=sha256:1b5f7c61a647816eab9005c786b92aeb2ad6d44b3e537d3b29530716b1c45f35

Observation db66a195-658c-4d4d-ad70-a76d807b50de · outbound

This paper cites Nambudiri.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Nambudiri

Reference 11

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.516563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.553401Z digest=sha256:d1652d13dd4994cd3a7d98842fe966bef6a1c2c31a4e812c60e4e8e71303a638

Observation d218f442-d4aa-4f91-8526-093e25e86295 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.211566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.699109Z digest=sha256:8ba3abc926372543b3a5c889b149a621af4993f7131d44f03c1ee2d21fefbfe9

Observation 8d3c30a9-5ce3-4dfe-aea0-71832e21c7cd · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:24.911183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.835466Z digest=sha256:2373a6d7b5c1402bd30763bef424c366bcbcac1758108aa87608bde3f543f4f2

Observation b4f155b0-a035-4adf-bc0b-acfbe676b719 · outbound

This paper cites Gold, and Vishnu Mohan.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Gold, and Vishnu Mohan

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:24.646083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.997095Z digest=sha256:ac26d1f77170876be3c7bfed73cff61e67d918cfb46c644b5cd7be9e7b9f1bc2

Observation 0b00c702-a108-430f-b331-d05736248e93 · outbound

This paper cites Akdogan, Jessica Atkins, Mohamed B.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Akdogan, Jessica Atkins, Mohamed B

Reference 15

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:28:26.928726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.118227Z digest=sha256:4ba158211c32dca980da8a92e73c6fa14a491cc8db468516beb66712bae20915

Observation 360da593-0879-41ae-b68d-d180032a3f19 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.709683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.257178Z digest=sha256:7d1214282807935d5f150ccf32440798c9490d0640175dc76e50f8545dee96d5

Observation 471f0add-3d1d-4d98-b9e8-5ada45aac9cc · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:28:26.715873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.356613Z digest=sha256:52b8653a6e2bc14038a478f8052d9500df23ae7878f01f6badfb067247c619ce

Observation 66b77b6c-3cb5-46a4-bb30-34e4f5e0ae11 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:28:26.543053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.457250Z digest=sha256:8aee02b05db27dbfb568586b83b91c59729abc6769412926299bf3c264de071b

Observation f0bda5ea-ae83-4587-9ad0-18d2cf1f5e63 · outbound

This paper cites McCoy, Faye Yu Ci Ng, Christopher M.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation McCoy, Faye Yu Ci Ng, Christopher M

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T15:28:24.346060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.573090Z digest=sha256:1b1e43acbdb0542078e7893d08fb18ad8c202abff365aa07431f1b203e4d0030

Observation 9c0539b9-67c4-4346-b6ec-8b8f6dc119f9 · outbound

This paper cites Pennathur.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Pennathur

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T15:28:24.065742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.719002Z digest=sha256:684f875bc0ee2d4219c498a15fd9cd51fa0ccb673216ffaaf3161a547e2d2db0

Observation 7076c52c-d4d9-4717-86ae-8babcc286d6f · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:20.840862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:20.840862Z digest=sha256:4510a6b18b3b78e2777bfade03d0624cafd13422ec141302f6f8418f7a69a60b

Observation 3dd95755-3335-47c8-90c9-2a4a316c76e1 · outbound

This paper cites Quiroz, Liliana Laranjo, Ahmet Baki Kocaballi, Shlomo Berkovsky, Dana Rezazadegan, and Enrico Coiera.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Quiroz, Liliana Laranjo, Ahmet Baki Kocaballi, Shlomo Berkovsky, Dana Rezazadegan, and Enrico Coiera

Reference 22

Resolution
verified exact
doi, observed 2026-08-07T15:28:23.800870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.944217Z digest=sha256:042c6bcc1c9f0821eb5e51bd780eee5d033aa733f76f4219b2b390923e4c5ae2

Observation dcd1a8d0-97bd-4d35-a5ad-19a631f90eba · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.045394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.045394Z digest=sha256:c417717c1b50fa88b939c080e33f1f75aeba3885e35b5727b45231695c39f344

Observation 8bcd07dc-8050-4cc9-8e32-0f41c6a7c13a · outbound

This paper cites Evaluating Consistency and Reasoning Capabilities of Large Language Models.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Evaluating Consistency and Reasoning Capabilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.131043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.131043Z digest=sha256:7fb7b84d45592ffd6e54874506fbf6abcd464585fccc128bb5528332fb70488b

Observation a08155d6-4545-4c81-90bb-908ddb7edd6c · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.571936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.251803Z digest=sha256:4b047b88411edd07669a4472028cf11f62feffba207261fcb11e03307087bd1c

Observation a31d8ed8-6421-4745-91a0-bf2c0f81b4a1 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:23.505634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.397661Z digest=sha256:2a92c66103ddad6b0d3196b02cc8571a0745d843cb51bcada96da985d954d1ea

Observation 43ed2592-f723-4b4e-8681-1e7894b5f171 · outbound

This paper cites Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:26.272163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.525302Z digest=sha256:b980d2897bbfbb8af9a0166ad0f662202853899a8c681c5be33f43a300393820

Observation a6267c53-5da6-4ecf-a5a7-6a05736bf0a9 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:28:21.711043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.711043Z digest=sha256:938c3f149f02ce3f05a1259c0ceac9c015e53778638978f02d73fca53f397bd3

Observation 52e6195b-9c5f-4d53-82b7-fd7a8ccc746e · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.819970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.819970Z digest=sha256:d7293a9a35c227961fa96874154081b07fa5d4670c83d953e4d5fdb73d81fca0

Observation 524d9d08-9753-4281-bfed-e5bb647aad90 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:23.238463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.958266Z digest=sha256:ccd41c791fa70bdc85d23a019cc552df91e78cc186b87c16f7e8b113801c5593

Observation 3359e3e9-8a9c-4779-8484-ce2658534173 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:22.130493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:22.130493Z digest=sha256:e1054ebc18c7029cb6a641d718993515556001466ca200b2ae7c5400b11cc3b3

Observation 224e5fb2-952e-4f34-9811-f7240bb54e4a · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.413515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.229663Z digest=sha256:f81296b3dd704e7cab8b7d7ecebd1c56467334164547441ea26d1081ed5194ba

Observation 2dd1728c-1d31-487b-b5fc-577f9164bcd2 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation BERTScore: Evaluating Text Generation with BERT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:22.370716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:22.370716Z digest=sha256:4e94a6525101cf9c44791b8f63d50bcefb3e36a45495847c552d50470f19dbe2

Observation 37bc608e-0626-4d53-ba53-32c804110b48 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T15:28:22.995871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.540412Z digest=sha256:f14bc7cd8412298b01c81f9e12cd050c15cd274b949d707e56d6e16fbeacc669

Observation 42334b87-d979-4a62-9709-4d43b0c01f1d · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.263988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.701010Z digest=sha256:fb9e02eaef52aceb8ca4338456eab7d168efa15791bfb01786bc34f1f690c324

Pith citing papers

No inbound Pith citation observations are available.