Pith. sign in

Paper Citation Record · LEDGER

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations

As of 20 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2506.11114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11114 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:39:42.506184Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.443456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T23:21:14.448985Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 060c174e-c991-4b8e-87d0-2019120b2618 · outbound

This paper cites an unresolved cited work.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:45.354572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:40.369014Z digest=sha256:4750b6156f1cc29ae9e4e3ade071e27e5310e1b9eaf100208ec80000a31dcbd2

Observation 27856649-cae8-4e56-911f-f832df18cab1 · outbound

This paper cites Computer software (2024),https://acrobat.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Computer software (2024),https://acrobat

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:45.340053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:40.467025Z digest=sha256:55565e2cec02692f94a35afb86d20bdfaf26557ed78bac09c8ca0bdb98ad5ef3

Observation 96fda058-952c-40bf-8a73-30b661501932 · outbound

This paper cites Clinical Anatomy38(2), 186–199 (2025).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Clinical Anatomy38(2), 186–199 (2025)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:45.323735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:40.638781Z digest=sha256:4177c2a62b0317538c12f2026ab0b935b82d79d42a7d2612d9557bd2a768c501

Observation 3795bcb1-f27c-41a8-91e8-033cfa954105 · outbound

This paper cites an unresolved cited work.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:45.307733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:40.809202Z digest=sha256:bf8f4072a2db533371f93e7648c1855845d7eb207c67e4b713adac9f98db5f5e

Observation b5464ab6-122e-485f-b741-99726d5abbfe · outbound

This paper cites ishiyaku-dental.jp/, accessed: 2025-06-05.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations ishiyaku-dental.jp/, accessed: 2025-06-05

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:45.292741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.025291Z digest=sha256:bacaf199545095f7b531ec9b716b8c8c508dececca4a024a3c0d7e52a78e0f8e

Observation 464a702b-addd-466d-9fb8-246070c98a74 · outbound

This paper cites an unresolved cited work.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:45.088649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.204662Z digest=sha256:43bccf17e8dfcbea445b1adb29806ef7345a1017a27da8dbc17ec8ba4a366759

Observation 94858ee3-b033-4a99-ad26-05bfe244faa4 · outbound

This paper cites mynavi.jp/conts/kokushi_kouryaku/, accessed: 2025-06-05 KokushiMD-10 9.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations mynavi.jp/conts/kokushi_kouryaku/, accessed: 2025-06-05 KokushiMD-10 9

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:44.891693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.344767Z digest=sha256:3041321ed29f70a55bd481f1e78d279ca0f96ab4f7e8b719389ef3569c790c0f

Observation 9b2af865-a51f-4637-9c95-60be196d77d3 · outbound

This paper cites Applied Sciences11(14), 6421 (2021).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Applied Sciences11(14), 6421 (2021)

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:41.416728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:41.416728Z digest=sha256:c6d6b8eb5cf65f071a221e994abc01d7ca1e50ffe03ce608720eae64b68365a2

Observation 15ea280b-ca5b-4968-bfa2-489e37eabb39 · outbound

This paper cites Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:41.561909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:41.561909Z digest=sha256:a5d00914677e367718cc62d70dea2ab1d3d5122423eaa297bf3f6bdd7d8198fb

Observation 6eb9a93d-b5df-4203-8c67-3eb414012bdd · outbound

This paper cites Advances in Neural Information Processing Systems 36, 52430–52452 (2023).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Advances in Neural Information Processing Systems 36, 52430–52452 (2023)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:44.516979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.728068Z digest=sha256:ae68f5f6b5268c609a11b4803f5cfe49597478fc84afe2026e411581a2052fef

Observation 25469ee5-777c-4eff-81d7-19685c55b6de · outbound

This paper cites an unresolved cited work.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:44.269056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.844522Z digest=sha256:91bf8effff5ff3f200e6984204f6583c78faffd5858df00f908e786ebc66522f

Observation 71c0fcb2-ec5c-4f2f-a345-4377efcd3715 · outbound

This paper cites html, accessed 2025-06-07.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations html, accessed 2025-06-07

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:44.016572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.879923Z digest=sha256:0ca24521edab606846a917926881e16f59715b28575ea23c21d72c6b4d784820

Observation 88b648ca-0929-4668-a34f-202c24122b8a · outbound

This paper cites an unresolved cited work.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:43.756452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.927605Z digest=sha256:b58cb823c9a14b1284e2d2641e7b2baf18a697a401d286266d1eb5a8d8c17532

Observation fed79d2e-52f3-4942-8150-f4e0437c148a · outbound

This paper cites JJDEA40, 3–10 (2024).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations JJDEA40, 3–10 (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:43.589654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:41.989459Z digest=sha256:ad5907ce569591dab41a93ca547e501e85d2362fbbcc5cedb0ac6de1932637cf

Observation 7494be47-5966-49ea-8042-2d749708eb89 · outbound

This paper cites an unresolved cited work.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:43.479864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.034777Z digest=sha256:e6b4d6ed10c6dbf34ea0e77ec478e15c11969342a6efb687309177ddba7d2e97

Observation 13117dba-ec30-4cf4-aeca-4b095e4f3f0c · outbound

This paper cites In: Conference on health, inference, and learning.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations In: Conference on health, inference, and learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:43.328755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.087521Z digest=sha256:ccdccf574bf28f793b55fb53b3c1be1b4d8734c63517fbcbfb55ec7899f4a62e

Observation 46ab7c81-32c5-40c9-a13c-d56e4dcf76c1 · outbound

This paper cites guppy.jp/, accessed: 2025-06-05.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations guppy.jp/, accessed: 2025-06-05

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:43.154140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.148285Z digest=sha256:4b36c7928f44fd8ba5208dc8b60c06cf0ec7f0c6a6b6ae454e87cd9452639a51

Observation 79da25e9-5225-4eaa-afb2-ec965edc9699 · outbound

This paper cites Bell System Technical Journal27(3), 379–423 (1948).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Bell System Technical Journal27(3), 379–423 (1948)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:42.193816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:42.193816Z digest=sha256:4a07525d73203502bfd65d6a12ba3c69f4187e4951d132ebe67cc8648118006a

Observation fc12455e-3dac-4895-a750-c48eaca57604 · outbound

This paper cites Scientific Reports14(1), 9330 (2024).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Scientific Reports14(1), 9330 (2024)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:43.040530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.253208Z digest=sha256:656be29cd09be40a1430676f6ac6c52ab86c70acf6118db3ee3ffc28bc6cc359

Observation 7798c562-97c6-46ae-a591-2a1f0aa362b5 · outbound

This paper cites A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T05:39:42.647730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.317690Z digest=sha256:9663b34d9ad626eb62daecbdbaf5ea6a52d9847741ae45451fbc1c97cd620306

Observation f3a30e7f-1d35-4b32-93e0-0085db6287a8 · outbound

This paper cites JMIR medical education9(1), e48002 (2023).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations JMIR medical education9(1), e48002 (2023)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:42.891688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.396218Z digest=sha256:cb2bb133446d847c865db401667ea9118802b7c47461a94ebfaed289126fb89e

Observation dceb8435-7d19-409c-bbb2-7283f6464edb · outbound

This paper cites Advances in Neural Information Processing Systems35, 24824–24837 (2022).

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Advances in Neural Information Processing Systems35, 24824–24837 (2022)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:42.767195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:39:42.448581Z digest=sha256:1d372918040a155471f08854377301ac10f4391bd0c6b976bc481a7c70f02b00

Observation 4c4cdb70-7b00-4997-831d-c99d919ca950 · outbound

This paper cites MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine.

KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:42.506184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:42.506184Z digest=sha256:dd717502427d024f44d7cf07fe84d0c638aa00a624c6a4db742f3dd021ae69c1

Pith citing papers

Observation be2dbc15-0c43-4ea1-85ab-a6312380732f · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations

Reference 292

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:21:14.454515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:21:13.443456Z digest=sha256:7434a457d78486fbe82cd35b165c72e1655e53e73e99a4b51878cb7192c80a24