Pith. sign in

Paper Citation Record · LEDGER

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

As of 19 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 3 inbound Pith citation observations for arXiv:2507.07988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07988 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:31:03.576676Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:11:20.481528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T10:38:36.176606Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fea62828-9d94-41fa-8edb-b65d3f7ba34e · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.602141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:57.538335Z digest=sha256:e48a47db37d621bd81a47d6270933f6e0eccc7ea9a621de7f70a6786cc6d0923

Observation 82c86c53-cff5-4b11-9e22-becdd6e7bb11 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.584943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:57.605075Z digest=sha256:6d09efbd4227dcb935ea4c17fd00275c43bb83f26f18029c656457bbafb2d2c9

Observation d024cad4-2a9f-4d7e-8650-c479bdefc270 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.569216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:57.699546Z digest=sha256:f10f9eb2cf3aac580cf187b0e588329bb32ac8281e91c93046b828002062065a

Observation b1fa0974-30b9-4dc5-8a72-2a64f4e7e327 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.550385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:57.775519Z digest=sha256:4fb80a2076f031f2c6007c9bc8bbb1362a522fe0f71c873d389a9a3857d6f2a0

Observation 9938ec45-ea59-4fd7-9f6a-7656b396dd7d · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.532446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:57.849681Z digest=sha256:a13d6f9f0f515c1f1a095a207674515525b4b1f9bc7eb8e4198db4503e9d3e9c

Observation 2a1ea739-43f7-4404-acc5-266adaefde6a · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.510289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:57.926277Z digest=sha256:acab34cfc773cac7556256bed31128aee4bd7bb19c80dd027601d51228f04ad5

Observation cf93f74a-3642-4262-a184-22013c439aa7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.490531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.030842Z digest=sha256:c025e19442aefad13e4080e80d93bffc0f84a39b837c98994a00cc8df964dc4c

Observation b8ea57cd-4e6e-42c7-a217-dee22dfa13c9 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.469292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.126704Z digest=sha256:84e3b96ddb4a670771a128e4f0edd8761bfa987d58385e525952057f028841c6

Observation 33371637-d575-458a-aaa0-b0ad6ed691c1 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.454065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.206001Z digest=sha256:e9b2c23d770fb1da64c560c125cb0ffeee80887ccaaad95528b64c2cd6d0a63e

Observation 78f6a361-d8f3-4917-b1aa-a90e6b6535bb · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.435785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.279514Z digest=sha256:002467e133b3db4fa61da0a22a0da1b17a5bdda45b44ffc1ae33a38e2c2701ab

Observation 8201bc19-11ec-40d2-987a-27f7010a8402 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.410700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.351768Z digest=sha256:991edad316e010de089ce6aa0c8a83b8a59a0512758489b4ce75bdfe0b5527a8

Observation 94d06134-d812-4eea-8618-3082cdef62c0 · outbound

This paper cites Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:58.427520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:58.427520Z digest=sha256:189624d9552a57b6c0c8b5f0e2686e196fc3de47aad275a2e0cc05f5a6255ed8

Observation 68a926b1-31da-48cb-ae54-ad162f7f71c5 · outbound

This paper cites & Ranisch, R.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Ranisch, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.387708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.512040Z digest=sha256:e4d0cd61dfae169ee99fa1dd713d1b3402433a177f323fd4d6f55a56b2d05ebf

Observation 0ee57d6b-f59d-4a04-ae2f-9d27b508b54c · outbound

This paper cites & Chen, J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Chen, J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.360873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.585270Z digest=sha256:295bed59f0fabe2f76e55476ef54f1448a4a52fb90ba2f597629f5af13376dad

Observation c410b272-f65c-4388-928f-d10e1bfbef9c · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.343527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.645309Z digest=sha256:9674802957b76a61e6bf62b22b36fed71352cbeb771c55b61174bc9b5a992abd

Observation df2e0a39-cbd9-401b-9a47-813ca40e9eb5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.315728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.703906Z digest=sha256:927afaee09b36193794d9bfcaaa9899f6048d218cbf25a70394f8f12220f9503

Observation 2b2ebbac-7dc1-4ffb-989e-a8cd2c496ec7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.296147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.787928Z digest=sha256:daeed03e739ce7082c890ae23fd99bcc19f56b02075adbc790edd5d850cbbdc6

Observation b6563a44-3ff1-455f-99b7-d1138c6e2e84 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.275469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.867942Z digest=sha256:d93b05ff9394326f097a2c653cb040ea1ae67062ae55ce55c98313aaecc357cf

Observation 9780ca1c-fb59-4ce2-a6fd-3f9d7771df50 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:31:03.894469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:58.912490Z digest=sha256:8889bf2f613be19464cc4abf03993345a095a7e73595f3d9e2145271398fc9fe

Observation 96bcb4f4-9993-4925-8232-c60a00c68819 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:58.963784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:58.963784Z digest=sha256:471590e8f7a8083b3529a9654596108cdcc1ff11e46a10b8401be600e6010f76

Observation 0edfd9dc-f811-4d8b-b2cb-163399547ac4 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.239650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.028935Z digest=sha256:ef02a799575b022197843ec001e594e398eff76f6a67d6bd78f2fa0917188552

Observation ce73f86f-a492-49a9-9699-29eccc6f86d7 · outbound

This paper cites E., Motzfeldt, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models E., Motzfeldt, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.223653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.088944Z digest=sha256:014396d457738759530a143bf2199b29a3e0108b3ff6e856e0a632c29368c987

Observation 4badec55-5a92-445a-beef-5119c443f365 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.199441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.135672Z digest=sha256:a527c18080df99bfd4325930cd4df07e5481d63c1f96737c18d03be2ec5451cf

Observation 7920dcd3-97d4-46be-82ea-5322ae45aee7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.155223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.176330Z digest=sha256:b378174fa79296ba85cd4fd07b056b670ccb3d7434fc074602d65d808d291154

Observation 9715e056-8792-4869-8056-d83e11810222 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.135472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.223086Z digest=sha256:6f4b0083d8d238e6869292259e0f28b8fd8213a4bb9efb9053086353e8535aaf

Observation 15b7a50e-c4b9-4033-9ae9-2ee3170bf2e5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:59.270249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:59.270249Z digest=sha256:cf4568ed21848cb0bce2a78f45e38fca90b9695f85ef063f72701cba9b72f9f1

Observation 654eef0e-e4ea-44f0-818e-9ab4b5b51d0a · outbound

This paper cites & Zhang, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Zhang, M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.095502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.323111Z digest=sha256:f95e651b1948fdc4081fdc8b051884c76dc317ec6d13c42d15aa6ae4161bd52e

Observation 595d3d64-676c-47f4-a650-f243b5b26c23 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.077073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.398355Z digest=sha256:46d42fda5a4d33503e0f7aa902dc2067c63c48d60b49c1d2fb7dc2ece250519d

Observation 513cfb26-1ebe-446f-8783-d749c16559e5 · outbound

This paper cites & Zhu, W.-J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Zhu, W.-J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:59.448696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:30:59.448696Z digest=sha256:17946306f2c7c7cd5dd514a94b61b04bc095b59899bb30bd6f3b7c4af110ea1a

Observation 32742a83-d522-476b-aada-9cafc088dc44 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models ROUGE: A package for automatic evaluation of summaries

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:09.056931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.526033Z digest=sha256:aaed2fc8f874eb41004431fdfc9758101f198894dee0397f983d02045603e6d8

Observation cad86f9f-c74a-42b3-8a7d-996eca6af49e · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.032771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.591080Z digest=sha256:67e67a3ed69f4674f271c4c286ab676972a7d2edb81b5f2a000db9a53099bc71

Observation f41df4a2-84cd-4fbf-9d46-856873202ddc · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:09.011481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.647833Z digest=sha256:39b7840895e7f4968d7ee0f152771f581cb1dc4362b7d0a612cc91a7c7b6c8cd

Observation 5325ec03-3582-46f6-96ef-30aba848691c · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.992397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.731276Z digest=sha256:b20624c2d366a935d8a6f01cb6c4f76adaf577172dba2474513f613afc6c2486

Observation 2b2bfa33-9464-4c8b-9c7e-74a7161d602b · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.970046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.816765Z digest=sha256:d8f641f75bc49bee97aa0a26831a54674e75aab02c9badafa20efac35b2450d7

Observation 0924487e-7e62-4d87-bc51-a69a859c37ff · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.948488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.927540Z digest=sha256:483362275837644762b4fee2e5b73bbb6c97afc8feea3db1c9682b3302de7516

Observation 7d75a361-2d51-4ec0-b04e-42f68e294ab7 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.933069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:30:59.986127Z digest=sha256:a79c9ded7d037a949adb549997be60a12445adb1bd23685d2eb8d03b142d645e

Observation 35697907-0a56-492d-9a62-502def21f83f · outbound

This paper cites & Yuksel, D.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Yuksel, D

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.063157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.063157Z digest=sha256:f9519830c6ff6db9ea5f4b3404a7dabe18e0897358e6739d196f8c4b9504c332

Observation 5104ce6d-cde6-4ee8-8a4e-d50de8cbda29 · outbound

This paper cites & Dredze, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Dredze, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:08.901236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.143843Z digest=sha256:ae1ec926a9331ecdbbc2adc3a1723d465cadb7d8e54496964ab088b216dc9703

Observation 9c211c11-1538-4043-be55-3e18d6957f65 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.711102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.227387Z digest=sha256:8b7a02d206baf550f9500f60101bb9fffcec48a544b35b86e8308491f3b28c91

Observation d77a8f62-3d2a-4da2-8746-98299c88fbc3 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.297754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.297754Z digest=sha256:bf9b9f0ad70008b8bc8d492f123c27ec42618a5a15cef92a718e5b47c11644af

Observation 7a873045-b05b-4cae-b63d-989c30b9dfd5 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:08.436810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.380994Z digest=sha256:f685eb310628526f3cc505d25a6320c368058e6e1fe18f72349e42893d51687f

Observation fc03a2dc-81fd-4c6e-8802-3e8c43caeb5c · outbound

This paper cites F., Goel, R., Wen, Z., Martel, J.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models F., Goel, R., Wen, Z., Martel, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:08.300854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.463207Z digest=sha256:91b79677b043fe74d5298778ce02e0c342169ea975eb2aa5ce29d4b5d4de9da7

Observation dba215f0-da18-488e-a5ef-b838c43c7cee · outbound

This paper cites & Dredze, M.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Dredze, M

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:07.990827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.540262Z digest=sha256:3f0860e5f6d6747d64a7b00698cac983880c9798b0d0e660c4accd4b4ff2e03f

Observation e012493f-0dd8-46c4-ba8b-5a9496c6af2f · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.630372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.630372Z digest=sha256:6aa5fc4eec3a65730edd2e78b117e81cd38622ac78f953e2f053ffce2fa0e5b0

Observation 62b7f041-cf32-42c2-aa67-16cd97084df2 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.848943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.760960Z digest=sha256:d3353091e08cc65a874c77bc68cae45905047ad097418631dad86878135d622a

Observation 7d9c4fd0-ba39-41b4-85de-4610ba30e2ff · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.530259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:00.867715Z digest=sha256:efc3dc6ad2db8db803e622d23c2e962304ae0d30e49c82c3aee9306189fe1a6a

Observation 8c4aa8a8-8ab3-47c2-9ff8-51943695fa21 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:00.962024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:00.962024Z digest=sha256:2b01df335626578ee33d001736277caae1e0413638ba20d6886a49fe91834286

Observation 3e22655b-e182-4afb-b696-df244c612cce · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:07.176482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.052752Z digest=sha256:469a04ce382c18c7ad3cec76f2f653768e69b6b48328ba2c4fc29f507b45abe7

Observation a0da6c83-e3ed-4dec-bf93-7866be91e26f · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:06.997332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.146944Z digest=sha256:f5bb6e3d257be9fecbb4da0ec4095a5f22e6194f71a4db6591dd6be7ca7b195a

Observation fc27fef7-a5bf-4b47-9e17-25612dec8a60 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Measuring Massive Multitask Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:01.239886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:01.239886Z digest=sha256:cfb7c6986ceb9bb69d12daf9e6424c3510c08a9d56feb1168cb59f0314ca1222

Observation d8abe900-b029-4332-ac3d-dda2921d18d6 · outbound

This paper cites & Gómez-Rodríguez, C.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Gómez-Rodríguez, C

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.835297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.341500Z digest=sha256:ab1a73ef08cb08e63503934dc33b3402a4921d6eddd43514a4e17d434c4534ba

Observation 1cbece59-209a-468e-95be-d351c0166c44 · outbound

This paper cites & Lavie, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Lavie, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.693457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.473941Z digest=sha256:27e838979759893799dc5c6b3e32cfd3fe484f8e413ac4b9ddde4e60db377e0a

Observation bfef7063-a6f8-4e3f-9838-a6b87e8c2f54 · outbound

This paper cites & Parikh, A.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models & Parikh, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.529950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.622469Z digest=sha256:54c8d4bfab892d9d1f5afd5ea7fc6cb3ad2621df14970f70138ca4ab99521223

Observation cd95447a-adce-4700-a42b-c96546f6f4d2 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:06.332867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.723726Z digest=sha256:0e34efd487e4c813ad0bf940f204d57c882f0cb3ce10e79ac79401fd4bbaeb6e

Observation e13e57a9-178d-4d8b-943d-20462a45e7a8 · outbound

This paper cites GPT-4 Technical Report.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models GPT-4 Technical Report

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.153060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.826623Z digest=sha256:ce37b458f1ab114f28c5b25f1790b810827589b9258714ed9cd86f7baa11a339

Observation 3b6e1b6c-72f4-4171-b9eb-9aaf9216b8d0 · outbound

This paper cites GPT-4o System Card.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models GPT-4o System Card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:06.021261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:01.969972Z digest=sha256:3f0208eab24071985039deecff7637a62b277da7f9490b9d341939dfac1932bf

Observation b93086a9-ac3b-448c-83d1-74cb23b0facc · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.868662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.061257Z digest=sha256:e180cecd700f4eb823809cba476b53ace7d2e2f251cee2201916b21c22480163

Observation a0709f7d-6805-4dda-bd90-2cd96bbb9273 · outbound

This paper cites claude-3.5-sonnet.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models claude-3.5-sonnet

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:05.736069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.213379Z digest=sha256:17afb531eb92efc700506d834fdf06c8364fd8d5a7d1d6c7e50fb14816926b21

Observation ea5f7796-2527-4459-be53-fdab36e7d703 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.599003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.376098Z digest=sha256:da9a5068ec57567160c1a66e2a69b503bbc9cf9e30ba8ea38a2f5f952128795c

Observation ba711ad0-87ce-41e4-b5d4-656e21e35187 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:05.461146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.462287Z digest=sha256:9727c26fc300b3a3187dde2924c791642a33b16ab21bc52ee61fd93e86792ee7

Observation d7d83259-76cc-46e2-9e34-e33de1a424d1 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.317895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.601521Z digest=sha256:c096d11b0084fe4087b2f067550d6076afe5f846254d34ee4c5df69903336b3f

Observation 35b2db4c-e65b-448d-8c84-ab6ea47ee2fd · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.192439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.743876Z digest=sha256:4b75c8f927c20343cafaa7c4e611c08c0abd03f2382cfa59ff1b9ee68ddf8dc0

Observation 78e6ae41-eea0-4b7d-87de-8c495109db28 · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:05.019319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:02.896069Z digest=sha256:93f71eafca00900b2dcc7ad2794464c9497d3a4ac4b4edb2809016b04d56eb5c

Observation 9026ac5f-3925-46e5-ab3c-acc744688791 · outbound

This paper cites K., Raha, T., Khan, S.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models K., Raha, T., Khan, S

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:31:04.840016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:03.071493Z digest=sha256:3f377c64ee33d16fa3790ab03961f31a608c9a076419080a3e7b7e1e5115231f

Observation cb3674c9-db41-44de-9bfa-6e784ff8c288 · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:03.237211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:03.237211Z digest=sha256:b69f51fd427a66f688cd146afafdc3bbed2e7c8aed98ec44dbee9b8d014971af

Observation c02f3d50-24f3-44ac-9322-865f2fb5b65b · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:04.545657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:03.443895Z digest=sha256:e048c8b41070bc2ac0dd30f88fff4880a91c95e4d4528e7699524ba6ac214827

Observation 8245cac9-65da-4272-af6c-45b915b618dd · outbound

This paper cites an unresolved cited work.

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:31:04.317615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:31:03.576676Z digest=sha256:7faf1be9bc8ac8be287a19d1f707b3b071f011fa49442f4c64b277a8293a74b4

Pith citing papers

Observation 89338951-d8e3-47ff-a305-7f5e7bfd8d5d · inbound

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models cites this paper.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.481528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.481528Z digest=sha256:69c3981624a1f1486754c6bfb87a5339732b186cdc25a63f27cb6b79b0dc1d17

Observation 3cfe754e-cd08-4ea8-9c04-f57d9a221b4f · inbound

Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation cites this paper.

Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:38:36.213313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T10:38:31.601683Z digest=sha256:b7598cbe8fa9eb54850bcbc11833c7d443a9778a21758233a156663ba7d78ab8

Observation e3612936-29d2-4b45-b3ec-42657fc5a18b · inbound

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering cites this paper.

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:28.955923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:04:28.955923Z digest=sha256:448f70ece587749241c479fa1a976d510908dd5bafee95beca45c4314c36aaf1