Pith. sign in

Paper Citation Record · LEDGER

HIVMedQA: Benchmarking large language models for HIV medical decision support

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2507.18143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18143 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:42:18.697628Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:18:59.283834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:24:01.661042Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00093d7-01e1-48eb-bb8c-7a925d5f04f4 · outbound

This paper cites & Topol, E.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Topol, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.345475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.327944Z digest=sha256:0addb424e1110cdad48a45252503b927724aed81b78c0af4559f78f48ac5b0d7

Observation cbade836-7e62-48fc-a0c6-56b58dc1a4dc · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.324802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.333246Z digest=sha256:e4267cff2098e1d0603f76b55bd6b986d952ba7d818d2512e9b1fd99f7aeaf6d

Observation 8cb3275f-2cd7-4c1e-87aa-50b5aa0d57b9 · outbound

This paper cites & Taylor, R.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Taylor, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.301557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.338407Z digest=sha256:104e050d1f4ce3f948a5eced40670bd5ba6e7820a6e105e47ff2b080edac292b

Observation a07a9f0c-7f11-4a62-aac3-8816d3a708b3 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.279241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.348108Z digest=sha256:57445597c5899d740917f98ada620ba5331beb93a9a6679b1b490a6758b2de99

Observation 7ab591a6-d8c3-434a-a2e7-c59205dfc5b0 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.258544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.353270Z digest=sha256:d55b08665402d3dd4ba9fd555d243d974100bf04d9503125dc4e420875102b9d

Observation d1145283-b369-458e-8b44-981598025c41 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.238457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.358167Z digest=sha256:7ae83c4257b2e077df7753c428cd3325451c12cec3dfef26d05a23c2a51de669

Observation 3bfe12e4-5e2b-4491-a1f7-701e880d3e76 · outbound

This paper cites Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial.

HIVMedQA: Benchmarking large language models for HIV medical decision support Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.216975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.364769Z digest=sha256:599b5268b1283c0c7bbcefec4c3e0e856956013a4447c422638750aff50023b2

Observation 9850c560-7d20-4d05-8d5e-bf9bf6e0f933 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

HIVMedQA: Benchmarking large language models for HIV medical decision support Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.370147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.370147Z digest=sha256:96e831ef977df3655386b2564f4889ab3bc969f5ae5d0c86ca4047a518bb2476

Observation b520f252-391d-40a0-bada-e08f7a2c7906 · outbound

This paper cites & Petro, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Petro, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.195521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.376914Z digest=sha256:ea8324e61cd2476113a40e2758dc6edef4533ebc68d793c0d91d2cee54222fcf

Observation a43450ec-d55d-429d-a722-69591d5051e0 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.169555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.382624Z digest=sha256:e3ed8751c536d2ed61142c9bb70fc0eef119c4c82213aeaa0bb0f8c3666641f4

Observation b9cd8862-9e6b-4482-940d-86bda2612c4e · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.147090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.388422Z digest=sha256:eebc93842c50aab29f6ee8775bba2e8e9687e5709013db6c4fedbbf271b7c015

Observation 6d58057c-469f-4d10-a7ea-6dfd1cfea59f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.112418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.396040Z digest=sha256:c5c8e5bfaf9185ee79a8ab78136ab7eb6f8872643a23207a26bee795ced37588

Observation a0268c01-e7f3-4dd3-b561-f2d9bc5a82d7 · outbound

This paper cites & Chow, C.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Chow, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.078979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.401352Z digest=sha256:3121f60d7b6219c3de6448dfe6ec450ffd02ae1bbda7b82f278b5d20489fddd0

Observation 4258786c-ff17-458e-9d9a-3d0308c087c1 · outbound

This paper cites & Dussault, G.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Dussault, G

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:20.049365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.407255Z digest=sha256:a5a939159cca2e665c20830efab821e428a7e3ad30204a832958e2e35d93d536

Observation df319221-9d75-4762-b780-023922aa6af9 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:20.027482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.414911Z digest=sha256:2bdb1d8493477fc4b51c7201d4f50c5f93732fe212fae9eea637cc75f33fcbed

Observation 4144b980-0807-4326-83eb-902099f163dd · outbound

This paper cites S., Link, K.

HIVMedQA: Benchmarking large language models for HIV medical decision support S., Link, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.996200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.423293Z digest=sha256:4174df2c5441891963438c02b2fde65d35bc5a305de69c56a88cf51ce7de8438

Observation 8b1b808f-b918-46e3-90b2-adf5a6666d7f · outbound

This paper cites A., Lester, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support A., Lester, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.975130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.428884Z digest=sha256:d32b2768e60f38ecf573926b558eea3258aacda7d8eba75d796d79cc295aabe1

Observation e9b44823-ec67-431b-95af-2be9221ba3cb · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.951455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.433357Z digest=sha256:a787442e728c3318a8cd79e1a2d91f9f83cb4d49cfa85f7e8ff671f3c01ae81f

Observation 194980ef-8d66-47ce-885c-05672c89046a · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

HIVMedQA: Benchmarking large language models for HIV medical decision support MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.438834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.438834Z digest=sha256:0bbdd200701517a9f7d873717004c248c8ab7e501fefc05975ddd426d2499bbe

Observation 283e7acc-25df-4505-9085-9bf869211366 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.927167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.444401Z digest=sha256:8440d72886a2e35ded5c35051c7999c8be15376e2d6bf9d329925f691757bc0d

Observation b09057d3-9bd3-4014-bb7d-26f8b9c2eb2f · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

HIVMedQA: Benchmarking large language models for HIV medical decision support AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.449467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.449467Z digest=sha256:ca1881798c63b751592c212a52daf6062f5a92a52257cc6b7ab5cd212644b236

Observation bbbbf14a-1c89-4f74-8a0b-584165d72658 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.900924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.457258Z digest=sha256:cdbc7a5d1d6dfb3515235964001ee65c0922bf8f6d3eb824db19d6df0be51882

Observation 54eaaba1-7eaa-4eb1-97b4-c421d08e0230 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-06T14:42:18.753396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.466737Z digest=sha256:1917e72d4aa2989568c31c646b56115e048efae6b495ebcf418f2e5dd6768890

Observation 7100ea01-68b9-4964-9fa2-cf8de4a11d86 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.881456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.472116Z digest=sha256:3f31a73e9e3ce40164d662cea309ebad18d3804bf137be80780528e22eae7586

Observation 905063a4-eb7c-40f9-a37b-15af8cd2b7b2 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.478987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.478987Z digest=sha256:6674091b1cd2c1ac429652c5ea6fcc6a91ba15af4f8b9d1b55666f6896946d00

Observation 30d1fa25-0606-4c56-873c-e0a04c54c838 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.839199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.483722Z digest=sha256:357c68036422054bce838ce828b126412fa86814cd9d6548c64d0382d9e46c21

Observation f52043c0-b81a-4f05-9069-e414a76359e9 · outbound

This paper cites Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data.

HIVMedQA: Benchmarking large language models for HIV medical decision support Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.489506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.489506Z digest=sha256:0e973e540d1766cc098ac434fa542c1dc1d656c735b3ba0ff217d30a39859f71

Observation 32ad7521-810e-4550-8e2f-19abbdb99079 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.817052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.497292Z digest=sha256:00871018621878a14bf98370c7268ac1df11de00d917be5d68c32ec350ba3e31

Observation 127a8375-058a-45e2-ba9d-17c2f4ffb014 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.790721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.502209Z digest=sha256:c8b5cd7011f23a810439a0447e027d176387ab80073ead4af0e2dbbec83e58fb

Observation 2c305b8a-1052-421a-8afb-cbb5ccd6ec79 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.769605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.507177Z digest=sha256:440532556ce5306ea3fc0caae265f9ade688a8fabb459e225ae7e6d7032e2286

Observation af8f18bc-6c5c-4b95-84b3-8c981eaebaba · outbound

This paper cites A., Lingohr-Smith, M., Rogers, R., Lin, J.

HIVMedQA: Benchmarking large language models for HIV medical decision support A., Lingohr-Smith, M., Rogers, R., Lin, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.731208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.515205Z digest=sha256:de54469d12b962dc501bcb9f18cf8c874e096bd9279b55028f2903beef69c75b

Observation 6d6b0bb9-fb89-4bcf-85f7-e95a6da22a90 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.712854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.519604Z digest=sha256:2ad2b4f1dc3f6d6103b9377cebec8e8b09ee3ea36348679eee43c19451736a63

Observation 9b5f2c17-8b58-47fa-9b71-d29d63ef46c0 · outbound

This paper cites The Llama 3 Herd of Models.

HIVMedQA: Benchmarking large language models for HIV medical decision support The Llama 3 Herd of Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.525401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.525401Z digest=sha256:75e3b655596349788b4ae1228d39077f5839625e66b65a35369cf65bfc0600bf

Observation 1a0ddae2-c61f-4989-99e2-13ed60bfe4f6 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.693675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.531488Z digest=sha256:270a44558b4b846bea847c30028b017c9bcc6c59f18defc51f7fc71c7bcc19d7

Observation e67a81bc-4767-41e5-ae98-d45116476693 · outbound

This paper cites Med42-v2: A Suite of Clinical LLMs.

HIVMedQA: Benchmarking large language models for HIV medical decision support Med42-v2: A Suite of Clinical LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.536520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.536520Z digest=sha256:116bec0e73db8e50e36c05897c2410ed116a6792feadfbaaba79ac97981736e4

Observation b116b51f-60ef-42bf-a9f9-b39101152c97 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.542389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.542389Z digest=sha256:0991b1823c105b7d5e54fdc3248c92d10088c70809749ea65bd306c425c7cea3

Observation 91abcdb6-f3b3-422d-8517-ae17d85f68f3 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.663539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.547648Z digest=sha256:e40a19ad2756055399842cd5717e36cac64e3850990d5c5728aa033e2c79dc04

Observation 85188c7c-f502-4aad-828d-c21a2ebc916f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.553615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.553615Z digest=sha256:5242bbd9f482f193b515d49a157f9a1d1b04f046a333ffba6bbf7058afaa0d4c

Observation 238f3bf3-6cd0-4784-8f72-113d9b51bffa · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.631293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.559307Z digest=sha256:19b46a38a2bd1804d0fdb54d63cbdfd2f061f3e6afa40aa1a4e7d891ad700fe7

Observation 733c0ab9-b198-48e8-a7fd-b0d6895220af · outbound

This paper cites A Benchmark for Long-Form Medical Question Answering.

HIVMedQA: Benchmarking large language models for HIV medical decision support A Benchmark for Long-Form Medical Question Answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.565269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.565269Z digest=sha256:e57ddd5490375e388a316baa5c92b7bc126ca4e065e889324071b968aa1e4523

Observation 73429537-e222-466a-a55f-98764968841c · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.601340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.571529Z digest=sha256:24ece9bd2ac58e183675d362322e0d86acfe29669bec451687cc145d16f451c4

Observation b656ce3a-8f40-43c1-8439-db8ba960243d · outbound

This paper cites GPTScore: Evaluate as You Desire.

HIVMedQA: Benchmarking large language models for HIV medical decision support GPTScore: Evaluate as You Desire

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.577140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.577140Z digest=sha256:405f142ffc4cdab27b2b3cb1f93fd90d9d72148f94b7844d074c409a6bfab33b

Observation 7a870550-04c4-4e4a-a3c1-716001ae2c5e · outbound

This paper cites JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability.

HIVMedQA: Benchmarking large language models for HIV medical decision support JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.584585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.584585Z digest=sha256:9d2bbb69f67b387cecbee9442f753c0bf4a3d9ef91444193398f2baf2abbe031

Observation 850d7db2-28cd-4daa-97ec-fc87bfbdcac5 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.590156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.590156Z digest=sha256:607c751b7684bbca4e7b7c11462d0176bf1d2e14bbf3543d806ecfd15cdd11f7

Observation bb43128b-c32b-4dc2-85f7-612df07471ef · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

HIVMedQA: Benchmarking large language models for HIV medical decision support Rouge: A package for automatic evaluation of summaries

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.566562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.596357Z digest=sha256:d38a3ba33e04546dc011a8912eb1178781272f4f88d2a3f77473b8323c55c6e2

Observation 0cb365e8-567e-4e8f-b125-853006d854d7 · outbound

This paper cites & Zhu, W.-J.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Zhu, W.-J

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.546357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.602428Z digest=sha256:33f5fb3dad43d3b15704244d2f65bea74cbf367e60e89c99a78b97a799a69a96

Observation 2e5c27f2-ab9e-46c9-9892-4b01b10ffeb6 · outbound

This paper cites The unified medical language system (umls): integrating biomedical terminology.

HIVMedQA: Benchmarking large language models for HIV medical decision support The unified medical language system (umls): integrating biomedical terminology

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.528066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.608825Z digest=sha256:586f55b4705acb8e5dbd79967e8383963e2e3ebc7b0b959695f84840d089376c

Observation 6133ece7-f6fe-4a77-9e9b-4196eab01b6e · outbound

This paper cites ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing.

HIVMedQA: Benchmarking large language models for HIV medical decision support ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.614111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.614111Z digest=sha256:8246943be59b05f43fcaf13a32ecc77172105a568a0564e6ed9399b830c3afd3

Observation de0629ec-f942-4841-afa1-579c39959101 · outbound

This paper cites & Duclos, C.

HIVMedQA: Benchmarking large language models for HIV medical decision support & Duclos, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.509088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.619442Z digest=sha256:cbcb59924a1717112dfabd219518ca568e71cb27ea1f974dc7d42311e26628b6

Observation 88e426b7-3f18-457c-a08b-de8d3d112b1b · outbound

This paper cites Owlready: Ontology-oriented programming in python with automatic classification and high level constructs for biomedical ontologies.

HIVMedQA: Benchmarking large language models for HIV medical decision support Owlready: Ontology-oriented programming in python with automatic classification and high level constructs for biomedical ontologies

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.492646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.624809Z digest=sha256:22ad56b8a3050571f012eed571ffc12fb301e79642505d9df907256dcdd8d2e5

Observation b556e891-6910-45b7-9227-35b94625b3d3 · outbound

This paper cites Unified Medical Language System (UMLS): 2024AB Full Release Files (2024).

HIVMedQA: Benchmarking large language models for HIV medical decision support Unified Medical Language System (UMLS): 2024AB Full Release Files (2024)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.474078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.630257Z digest=sha256:62e0dddb17d4bd36f5cef4e9fdec604b912c084660eea2756f26e334e305b9e1

Observation 39c37b99-6a8e-48bf-8e40-05f3efd272b5 · outbound

This paper cites Nltk: the natural language toolkit.

HIVMedQA: Benchmarking large language models for HIV medical decision support Nltk: the natural language toolkit

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.457002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.637022Z digest=sha256:cb988f26a2c748fc7c26d7783c69b1f2d345228fc4b9ee2b04d6d4756bce9b2d

Observation d7074ab6-9d76-42e6-ad9c-628e233c462f · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.433360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.645060Z digest=sha256:a2fb1e86ad4040a34d9f06587d24ad202ac9aeba082b7962fbbe251f8206b751

Observation 611bcb64-c907-412e-b97e-675a73de8fda · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.412058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.652634Z digest=sha256:1382065a80a74e5262e08b58b328e6b7d386575a45b0e52f28f7dab2e3d8a0a9

Observation 21637a13-bd05-44d4-b07f-63821aae5d57 · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.385838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.663408Z digest=sha256:c9f1604e7f674db4b85f3c1f0e9e0d05696e1af6d3b4db789a515dca31a7ecc6

Observation 21272b7e-acc2-410c-8499-524d04da7d5d · outbound

This paper cites an unresolved cited work.

HIVMedQA: Benchmarking large language models for HIV medical decision support Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:42:19.365142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.670150Z digest=sha256:32747da79956a58b21ff6b619b5b67a8b8d2fac5e7517914ea17925e03262fb9

Observation 25b0ab7b-26fb-4b0f-acf9-01c3e1c304f1 · outbound

This paper cites - 1-2: The student’s answer shows partial understanding but contains notable misinterpretations.

HIVMedQA: Benchmarking large language models for HIV medical decision support - 1-2: The student’s answer shows partial understanding but contains notable misinterpretations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.332237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.676520Z digest=sha256:f80536ba0c87ed8808228a8af1e45412be5d467121f50244e67775aead59b59f

Observation 3cb5bd30-414a-4147-af87-b03d56049fd2 · outbound

This paper cites - Score low if the reasoning lacks clarity or is inconsistent with medical principles.

HIVMedQA: Benchmarking large language models for HIV medical decision support - Score low if the reasoning lacks clarity or is inconsistent with medical principles

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.315183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.681719Z digest=sha256:9fb3646667228f935956c5c77ee235a33ee4b46dd9cd2c933edde982c8d08c4c

Observation cceefff4-bf91-4f9d-9ca7-aebe4c019f74 · outbound

This paper cites - A lower score should reflect the severity and frequency of factual errors.

HIVMedQA: Benchmarking large language models for HIV medical decision support - A lower score should reflect the severity and frequency of factual errors

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.296570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.686903Z digest=sha256:65e2947d6a3609580c8ff1ca23b21d0cc44d901027d52e522e0b282af46503ed

Observation 0f45edbf-6d13-44e2-907e-40830b13a740 · outbound

This paper cites - A perfect score requires complete neutrality and sensitivity.

HIVMedQA: Benchmarking large language models for HIV medical decision support - A perfect score requires complete neutrality and sensitivity

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.277039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.692350Z digest=sha256:ee1d6e27507a6383d5b126b6e82974becdb06e9978b47cf95bf2694d209868c6

Observation 7b71bcc7-b87a-49af-b57e-7544a5405a40 · outbound

This paper cites - Perfect scores require clear evidence of safety-oriented thinking.

HIVMedQA: Benchmarking large language models for HIV medical decision support - Perfect scores require clear evidence of safety-oriented thinking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:42:19.256427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:42:18.697628Z digest=sha256:dcdbac9406422b8fbb9141ec1fd57c4d8aa50821575eea02083839fe3552e367

Pith citing papers

Observation 19b33552-b332-4037-a819-f9cd47f40acb · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment HIVMedQA: Benchmarking large language models for HIV medical decision support

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.662324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:6de488d6f249c7b4eb8bcc100d373fe4b3c8d13f914061c26aea63f5b684abee