Pith. sign in

Paper Citation Record · LEDGER

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

As of 2 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2604.20441.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20441 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:23:11.924161Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:13:05.342181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2847d94c-c931-4c37-a0d7-30ecd0f1aec1 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:20:05.782711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:6057c623b4c0659658455c7caab7df15c666a4331c1fbe04425935f87f8d3673

Observation f58c9f30-9fa4-49ea-aafa-a2f8016efb60 · outbound

This paper cites Available: https://arxiv.org/abs/2603.04448.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Available: https://arxiv.org/abs/2603.04448

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:46.851570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:b441f61e3e05c6f41645dde2b4f4717aac7dab56eced4af78e8885df697fc001

Observation 387dfcc1-819c-40dd-a578-28cfb4369da3 · outbound

This paper cites Artificial hallucinations in ChatGPT: implications in scientific writ- ing.Cureus.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Artificial hallucinations in ChatGPT: implications in scientific writ- ing.Cureus

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.442383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:32cd40061b9a284e61b102e4f1fa6c64ce75c253696109d406bf03321210e24a

Observation bdad854f-bcd1-457c-845f-26ebfa7c95e0 · outbound

This paper cites Evaluating large language models and agents in healthcare: key challenges in clinical applications.Intelligent Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Evaluating large language models and agents in healthcare: key challenges in clinical applications.Intelligent Medicine

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.430105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:d5d7b6e59dd6f277c8cad06466519e9c2dcd69ee51f3e64a54612aeb8e7c124c

Observation 9408c621-496b-4684-bbb0-a26fdaac6da7 · outbound

This paper cites Survey of hallucination in natural language generation.ACM Computing Surveys.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Survey of hallucination in natural language generation.ACM Computing Surveys

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.426975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:7f32fd4eadc1c600832e0a3ebfa1cdc25145c769dcda3433b9f8e0ca660c40a3

Observation ad98ef44-2e97-4ca7-a20d-7c4655b1de5b · outbound

This paper cites Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.432830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:d1db5d4e7c4d68cd22da4047dc1f858af98c59c99ddb6f9fbd70d1d77a618677

Observation daf00cea-982c-4247-b2c7-5a2a64badf19 · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Capabilities of GPT-4 on Medical Challenge Problems

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:43:34.900774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:3b90315031ece24e07fd0e6d86e5adaa35d5106428623d9896c8970164084950

Observation 0c63dbd3-3003-4cd2-ab30-81e2ace8e169 · outbound

This paper cites Toward expert-level medical question answering with large language models.Nature Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Toward expert-level medical question answering with large language models.Nature Medicine

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.421182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:006faad0eb6c00676a725191b157d00a9f0f905dcf7c243b7283dfb109cf59b7

Observation fd1967b1-fdc5-4b47-bb65-25ad8f361079 · outbound

This paper cites A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains.npj Digital Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains.npj Digital Medicine

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.418331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:422277c673f7df01f71ac64a0c5819648299a6c6ab773fbc558224d61cd63eb1

Observation 005e241e-5a16-42f2-bf25-2288a53cdda0 · outbound

This paper cites Large language model agents for biomedicine: a comprehensive review of methods, evaluations, challenges, and future directions.Information.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Large language model agents for biomedicine: a comprehensive review of methods, evaluations, challenges, and future directions.Information

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.424241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:ed54f0ca994b636d2b40e27880a2720c85c6ec5413a4c8ef6a0d9d3cce778e27

Observation 4d08f078-c778-4bac-a497-96a05758cf77 · outbound

This paper cites MedAgentBench: a virtual EHR environment to benchmark medical LLM agents.NEJM AI.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills MedAgentBench: a virtual EHR environment to benchmark medical LLM agents.NEJM AI

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.435941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:2140da04668e1c96f64310469cb97df80d02523372852430218c34a0bbdf71b5

Observation a22f7ee9-8407-44b0-9ea0-80cc2561f871 · outbound

This paper cites The clinicians’ guide to large language models: a general perspective with a focus on hallucinations.Interactive Journal of Medical Research.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills The clinicians’ guide to large language models: a general perspective with a focus on hallucinations.Interactive Journal of Medical Research

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.439189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:786cf3db11dca5c6b6f891b7b0c8a2fcf00d8d35fb70739063be0c1ce80e46ca

Observation 565ed65d-653a-4a3a-b728-8822b37d76af · outbound

This paper cites Human researchers are superior to large language models in writing a medical systematic review in a comparative multitask assessment.Scientific Reports.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Human researchers are superior to large language models in writing a medical systematic review in a comparative multitask assessment.Scientific Reports

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.408720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:627b2fa84b2783db7f9d75b5ed29ebd5b8463f0892c80af77e2a9ab0311522f1

Observation 53a6b705-4d3a-4897-b309-2cb9beaf5c3e · outbound

This paper cites Citation integrity in the age of AI: evaluating the risks of reference hallucination in maxillofacial literature.Journal of Cranio-Maxillofacial Surgery.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Citation integrity in the age of AI: evaluating the risks of reference hallucination in maxillofacial literature.Journal of Cranio-Maxillofacial Surgery

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.405836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:26272026cc1150cbb5b7f67e4227a7aada187b5f82b76895eb2e023fbb42f193

Observation 025bbcfd-554f-446a-bb02-500322b85120 · outbound

This paper cites 2025;12:e80371.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills 2025;12:e80371

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.445786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:440c0a1fea1d5d4e04dec7fc0aa7f9fd7097a763cbe5d2659ea50fce49062ee6

Observation 561e8192-e414-458d-9353-44f1a726abf8 · outbound

This paper cites Systems and software engineering — Systems and software Quality Re- quirements and Evaluation (SQuaRE) — System and software quality models.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Systems and software engineering — Systems and software Quality Re- quirements and Evaluation (SQuaRE) — System and software quality models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.411454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:866081411d49d32618ea40a87514a2c4f12383787edbd71f0f7dd4e4e35abce6

Observation 4da631c9-e38e-4451-b8c3-71e12503e2cc · outbound

This paper cites Data structures for statistical computing in Python.Proceedings of the 9th Python in Science Conference.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Data structures for statistical computing in Python.Proceedings of the 9th Python in Science Conference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.414472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:278f3c69cb2ebc315bc79ec4a5e0a3a002af4734f463a11a51ebc9732a02d0fc

Observation 63a37b86-0588-457a-aff9-f5ad6307b955 · outbound

This paper cites Pingouin: statistics in Python.Journal of Open Source Software.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Pingouin: statistics in Python.Journal of Open Source Software

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.462501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:db903a3643404f00d7f23705ce7dd592fffc292b94456ac29297d336c3e13be5

Observation a1fecdbd-0f00-43f0-b644-3fdd6af9e029 · outbound

This paper cites SciPy 1.0: fundamental algorithms for scientific computing in Python.Nature Methods.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills SciPy 1.0: fundamental algorithms for scientific computing in Python.Nature Methods

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.458971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:cb6fa1db91df7d148ea1b13053997d5e1b1647baaa8d74356d4d02df8a795f29

Observation 5b40463b-23a6-4d72-b3a0-629a133577a9 · outbound

This paper cites Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.455243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:3813c27426117351b2ed66df828f7af87469e81b7dec9fb1def752367b78969a

Observation b95c98ec-d528-42ba-bfa1-afe077dfc496 · outbound

This paper cites A guideline of selecting and reporting intraclass correlation coefficients for reliability research.Journal of Chiropractic Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills A guideline of selecting and reporting intraclass correlation coefficients for reliability research.Journal of Chiropractic Medicine

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.449377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:09035a7cb5860e40495f3a8e47d4227643d24c03fde41183df55e005dd17c89f

Observation 5408ba97-481c-4d44-9529-a39d5275c84b · outbound

This paper cites Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit.Psychological Bulletin.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit.Psychological Bulletin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.452309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:0d859d077e0cb0979d362b36055c52084327debf962b6842376c05a89be3aad2

Observation 11af07d3-5b69-4f4b-a94a-f244269f03fe · outbound

This paper cites Statistical methods for assessing agreement between two methods of clinical measurement.The Lancet.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Statistical methods for assessing agreement between two methods of clinical measurement.The Lancet

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.466281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:cde77b781a9be8cd099ec652c6fc45c3fd5a46ab65ce320617a3a5d52c5483f5

Pith citing papers

Observation 976039e1-885a-41ad-9f78-d3851b1bacd5 · inbound

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries cites this paper.

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T14:13:05.342181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:13:05.342181Z digest=sha256:57ee1e7f5cb2f780b41b083b8943015437cdfaa4d5c81fcb08c1403ab074bde7