Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Match the Conclusions of Systematic Reviews?

As of 19 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2505.22787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22787 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:05:45.754584Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:57:11.817415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.432897Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abfb9d70-afe7-4cb1-855e-741140657f6c · outbound

This paper cites Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases.

Can Large Language Models Match the Conclusions of Systematic Reviews? Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.999188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.627088Z digest=sha256:65233ff7cc5c94d1d6671d8a89d1ca5300fa86c13c2fbac1410138a4b686ed95

Observation 8d6b0df0-f2ba-4cb7-817e-201fb4cac29a · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:51.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.702126Z digest=sha256:6ff3dd9312d295cd5076f1541d1c68aa73cb827f317507799762a36da5c59afc

Observation 8823c06f-d137-4b0f-991b-818d8f54718c · outbound

This paper cites The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:41.788855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:41.788855Z digest=sha256:9c9465428d0183cf992f6eaed1f44d4a20f44c23eb0707118ef2feb7c9bffc82

Observation 22686ff8-3484-41d7-a983-d27f398c74c3 · outbound

This paper cites How to optimize the systematic review process using ai tools.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to optimize the systematic review process using ai tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.728714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.857108Z digest=sha256:6fb654a6cb14d1e32f74e65aeb7b10501ce0663168f60a3f48dadf616cbf1fe7

Observation 3ef3b132-53c3-4c9f-b155-040c13c904d1 · outbound

This paper cites Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses.

Can Large Language Models Match the Conclusions of Systematic Reviews? Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.608818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.917310Z digest=sha256:59f4f10149736b5ef6eb5c262665afc3cde25b70b16416364d7e6a95212e8260

Observation 41609f5e-f087-4fa1-a9c7-11f8ba592707 · outbound

This paper cites Deep research system card, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deep research system card, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.433866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.001265Z digest=sha256:9a3e649c46baddc2ff10b6fbf2531f9352379632e26d1f69819fc32768947953

Observation 94eaa987-14a8-4cf7-9b2b-938abb39df49 · outbound

This paper cites Gemini deep research – your personal research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gemini deep research – your personal research assistant, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.336859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.069304Z digest=sha256:384cf26a6a7e7d4feb111e1594afc51c3afea963c51d82f14a00b0ab5d138060

Observation 0f59b63d-401a-4432-b88c-b2d27d28c6cc · outbound

This paper cites Elicit: The ai research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Elicit: The ai research assistant, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.172471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.163546Z digest=sha256:077778c04d5722143249839ce56ffec421dbd8b74de80cbe406c82d1c5dc0192

Observation 00702845-aabc-4db0-8d57-5d07882d3dff · outbound

This paper cites Open evidence: Ai-powered medical information platform, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open evidence: Ai-powered medical information platform, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.025871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.237160Z digest=sha256:dcc5b1d3e27c1180d0d2074f13aa312daadd9c5f7cdd0d7b92f7763d7b383cd5

Observation 63bb4f11-83cd-4faa-8b6c-af94bbb4ac6d · outbound

This paper cites Food and Drug Administration.

Can Large Language Models Match the Conclusions of Systematic Reviews? Food and Drug Administration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.851003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.333145Z digest=sha256:279f647a5a494e290e3db932c01bfc7bb32d62b4d0d792709c1523a0a6e9b08f

Observation fd923c07-34dd-4cab-ae8e-3e618ca7096e · outbound

This paper cites Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report.

Can Large Language Models Match the Conclusions of Systematic Reviews? Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.369820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.369820Z digest=sha256:66f6bbbbe0991d55fd5acb0cdb36858303a0b05bb0b41ef490fc032ffb6fe49f

Observation 35007af2-d8e3-4837-b6b3-21daba273a0b · outbound

This paper cites Can large language models reason about medical questions? Patterns , 5(3), 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can large language models reason about medical questions? Patterns , 5(3), 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.753052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.424003Z digest=sha256:0eccdf0eb401376431915ca00240b2863c08f2a35b231b39c2cd148489532e58

Observation 182459d2-29d9-417e-9a0e-fbbf805386af · outbound

This paper cites Medalign: A clinician-generated dataset for instruction following with electronic medical records.

Can Large Language Models Match the Conclusions of Systematic Reviews? Medalign: A clinician-generated dataset for instruction following with electronic medical records

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.618921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.533830Z digest=sha256:7d0c89221ea0058e0a4e2095a6cdd98c2f26940b3b4ea9763797a9335b275b19

Observation 85ac5650-3983-4a41-90cf-dd91d884f93c · outbound

This paper cites Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.445363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.587478Z digest=sha256:fb3a0098e1b97e18d67c3c72328ea341f72ddfcbf7ac50fc6e1a160a6bcb7691

Observation 2e81b9be-4c87-4a88-86df-6691597059e5 · outbound

This paper cites Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.285844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.700952Z digest=sha256:41ca33dbf8bde83e2ed6f5949f241aead5472a0d45556f31ff664489987dff0b

Observation 9f2c0cdb-5233-4d15-9719-d9efeaca81e6 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:50.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.778063Z digest=sha256:769a42db6c3be00644a1d20a8e9c0b2c77311a3e19785c569ea14acd25aa6b53

Observation 38605347-300d-4d31-a656-f16fa7949a8f · outbound

This paper cites Assessing the risk of bias in randomized clinical trials with large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessing the risk of bias in randomized clinical trials with large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.934019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.850157Z digest=sha256:10300c79398dfe8dbf0c3428c3cd3c42fe8e1e20e537feb36772d048162bc0c5

Observation 98e86c26-42ec-4b84-b016-a820ada89bb2 · outbound

This paper cites BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature.

Can Large Language Models Match the Conclusions of Systematic Reviews? BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.893682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.893682Z digest=sha256:c0a79b94b132e446c2e5e3bf3f0aafa5b92a69d592b00342c0136df25cda93a3

Observation d67b7e0e-e96d-483a-86f3-58ec3d060384 · outbound

This paper cites o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \.

Can Large Language Models Match the Conclusions of Systematic Reviews? o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.738498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.975698Z digest=sha256:0f156987d9e1a30ba8a1f41f861a6f75bbc2c54b6349818c0d1485a31a639f48

Observation 25f5ae01-0e52-451e-8498-0e05348f9d19 · outbound

This paper cites Generative artificial intelligence use in evidence synthesis: A systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Generative artificial intelligence use in evidence synthesis: A systematic review

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.615255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.044004Z digest=sha256:133f9bbdaf76231eb79cd45b3dd4a2b01ebb29adcd7f02d2c47b9b09d696475d

Observation ba71cfcd-0fdc-43bd-b69e-7113e17e0ee8 · outbound

This paper cites M ed REQAL : Examining medical knowledge recall of large language models via question answering.

Can Large Language Models Match the Conclusions of Systematic Reviews? M ed REQAL : Examining medical knowledge recall of large language models via question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.417678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.127980Z digest=sha256:191dc60563e1cd20c8f0cb35eeafc60c38da43e10efe950e16f1e4a3f6a82055

Observation 4ee122c0-75c4-4ee7-9002-b1a475e4b530 · outbound

This paper cites H ealth FC : Verifying health claims with evidence-based medical fact-checking.

Can Large Language Models Match the Conclusions of Systematic Reviews? H ealth FC : Verifying health claims with evidence-based medical fact-checking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.287045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.193185Z digest=sha256:b6bdfef0b9e4c12b14390740abb02d19be4936afa32cd77b3e4aa305a60be0b6

Observation 116dc038-7c89-44ff-94f7-91293199787a · outbound

This paper cites What evidence do language models find convincing?, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? What evidence do language models find convincing?, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.071576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.238814Z digest=sha256:ae0353bc9423bf4b8f5ef4d153e8c6db0658455904343477c9ef934e84717e82

Observation 1d2fea88-ace9-4c10-86be-61cdf76911e7 · outbound

This paper cites Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.888788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.332339Z digest=sha256:dcb4886813fe977ebc65dae7658c5d5a23ce722b599f730ac32d60fcd8c90450

Observation 2fe2347a-368f-481e-9843-adec69ddf050 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.713111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.408496Z digest=sha256:be1339fb88d295f86c1afd54d568750af60d3709d1e5b2bf1ffe3270a184eba5

Observation 153c3248-884e-4b22-b754-fe928d4bb7b1 · outbound

This paper cites Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.535178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.481306Z digest=sha256:a5f065c92a1c2b151c94122541c7d646a71ac1f44ceb88c386e327932a5c7a51

Observation c70947c1-3b0c-4aa1-b33e-b3a69558d131 · outbound

This paper cites How to write a cochrane systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to write a cochrane systematic review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.332514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.544854Z digest=sha256:64826a4a8cbe8d2e6d2c72c24261e625e2d58bcfc7ac4cfd66cf52031422ff21

Observation b0a7dd8d-16d1-4069-84d7-ff2ff483006e · outbound

This paper cites Quality of cochrane reviews.

Can Large Language Models Match the Conclusions of Systematic Reviews? Quality of cochrane reviews

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.191484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.584060Z digest=sha256:357d7386e4eadbd870072676e0da141e625d2a8832ecf5cc3c6108842262e7e4

Observation 0ee4b7c3-0587-4990-a353-0b1af38a86f4 · outbound

This paper cites What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011.

Can Large Language Models Match the Conclusions of Systematic Reviews? What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.996527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.680202Z digest=sha256:c9f59d9d3b1cd68b51f081c600eaf933bfa9a3d9f4d58bc09ffdb797b59dfb11

Observation f5763b0a-a6c1-4835-8b0a-348f3ea98e09 · outbound

This paper cites Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.846344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.833178Z digest=sha256:eb8cf25a8e203e7643c6e65971ce6da0879eebf78fa43847279b3fe96cbb9564

Observation 57ebdff6-cfed-4775-83c3-ef9ba273a830 · outbound

This paper cites Bethesda (MD): National Center for Biotechnology Information (US), 2010-.

Can Large Language Models Match the Conclusions of Systematic Reviews? Bethesda (MD): National Center for Biotechnology Information (US), 2010-

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.705602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.916594Z digest=sha256:62fa2872dae94bcd5aa1f2a64754b7cd5992f2832a63eb5ff41d09d4e8354358

Observation ddc25f37-4587-4ab6-9032-9003f348532f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:47.583430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.012378Z digest=sha256:bec721a6aad37b14646c3e4fb940d310639933d11cb7b69343a299d17344e222

Observation 2eebce0a-e7a3-4f98-b7aa-e7752544f1f9 · outbound

This paper cites Assessment of the strength of recommendation and quality of evidence: Grade checklist.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessment of the strength of recommendation and quality of evidence: Grade checklist

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.424267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.097237Z digest=sha256:cd2587a19d4d8f363c2bc6a784657a394f067d9b858f48ee3504390704d9f513

Observation ff297756-cfd5-48f9-916d-c8f0ef492680 · outbound

This paper cites Openai o1 system card, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openai o1 system card, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.179454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.179454Z digest=sha256:fea9f60827241b06b024897bc32abddabfb638a165248753fcbe4d9ff48882d9

Observation 052921c6-0370-49bf-ad06-b2da74cec50a · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.270040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.270040Z digest=sha256:a2422a7c27da16f83b62f1862b9f61b9cbaf1ac1d9be5d20f91f2f22dd9f9496

Observation a2576a7f-030f-4e82-b9a4-b5681c7e2546 · outbound

This paper cites Open Thoughts.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open Thoughts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.248285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.358789Z digest=sha256:a6df58c088f0a5e3aac4e343652ae2fcba9788fff9a847da77ca4379d74818a7

Observation 74080c23-941b-4d06-9368-74c6f7e90f55 · outbound

This paper cites Gpt-4 technical report, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gpt-4 technical report, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.428543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.428543Z digest=sha256:8ba7c56b75146b8a6cfdf5d7647d4133caa10f3d3fea12dc0d4b5d05073eedfb

Observation 305a4698-4bc5-4ff5-8564-cad8cac40ee3 · outbound

This paper cites Qwen3, April 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen3, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.486772Z digest=sha256:4786987ed71c99f32f7e2e4ac18ab0a43174abe21f435b742cec9189a3233871

Observation 4f7c6598-9e4c-44fa-b82c-930f310679ac · outbound

This paper cites The llama 4 herd, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 4 herd, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.057232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.547522Z digest=sha256:fc73330f1ffe76aff7b9eb553a593366a63991a775453d5a8dc28f1bb643e963

Observation f7222a5f-1b41-44f6-817a-c64f0cd26817 · outbound

This paper cites Huatuogpt-o1, towards medical complex reasoning with llms, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Huatuogpt-o1, towards medical complex reasoning with llms, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.623175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.623175Z digest=sha256:a31e7177207b2cc355f8fe855c83ba606cd9eb8b57eb488daa5bf745d19fd968

Observation 935206a9-ac73-4ecf-90e1-0bddbf237818 · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openbiollms: Advancing open-source large language models for healthcare and life sciences

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.888598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.724796Z digest=sha256:b94693e4abc8d3e711ff7256ddecb9c8b70783314c981d15b717397e638bbacc

Observation 01a60bc0-9a29-433e-9a9d-b6db76a4c5b4 · outbound

This paper cites Refinedocumentschain.

Can Large Language Models Match the Conclusions of Systematic Reviews? Refinedocumentschain

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.742395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.817332Z digest=sha256:28ebf3a4e5e9d4038932445bd35d9131dd4d09c8511ccb3a2a6081b4a3b52e2d

Observation 0b21fda6-023b-4ded-8cba-254e44ea342a · outbound

This paper cites An introduction to the bootstrap.

Can Large Language Models Match the Conclusions of Systematic Reviews? An introduction to the bootstrap

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.604185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.916032Z digest=sha256:6867248bec8106c7c05f629fad3eed839b17dcb178c29297d84fc6b0c11f6a0f

Observation d6ba0e78-aec5-4990-8dad-eee64931b492 · outbound

This paper cites Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.983710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.983710Z digest=sha256:21cfce4b8d41950398a39974b23bbf29a150ae722be61e0d7f41f9b4316e1822

Observation a9a25759-e103-4159-9809-9ffbe1bc9358 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long-context LLMs Struggle with Long In-context Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.055957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.055957Z digest=sha256:b18c4841627ae91f456666142a81e42655c91e1be4062d22adaa652d7b6c32ee

Observation c1bc1f40-dfa1-4665-8f19-8e8aece3d297 · outbound

This paper cites Large language models are overconfident and amplify human bias.

Can Large Language Models Match the Conclusions of Systematic Reviews? Large language models are overconfident and amplify human bias

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.149556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.149556Z digest=sha256:f3f17d220d8de0e33d6ac99b898cbfb36e4b281654dbef34f74a4d6d2272665a

Observation f740f987-927e-461a-aba9-43f980797571 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.189971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.189971Z digest=sha256:b0ef7987d1b9fcf324eee7896a6e2c43b7e37cd63f74dc738cc32b596acb3bd6

Observation 607a5fc8-c3fc-456a-bcae-c39cd7a68ee4 · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Can Large Language Models Match the Conclusions of Systematic Reviews? Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.303527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.303527Z digest=sha256:1c9c433116af98f3394d1836094e8bd988dc4c76a86392208e114bc5a78c2bf8

Observation c88dedd6-9da0-4f3b-bbde-83e6c12e97fa · outbound

This paper cites Fine-tuning is fine, if calibrated.

Can Large Language Models Match the Conclusions of Systematic Reviews? Fine-tuning is fine, if calibrated

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.443913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T13:05:45.381013Z digest=sha256:b4c1ee02367fda153a8718fab1aa2607c179ed4364c4a545d754a57c57f3efa2

Observation 8ec6626c-2a83-4b3f-978a-1cc08e17c7ec · outbound

This paper cites Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data.

Can Large Language Models Match the Conclusions of Systematic Reviews? Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.447639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.447639Z digest=sha256:b45ef2524f190eec4cd55a03734b29ffacee126bbe08ebb64e8e26cfce42b0bc

Observation fd2586f8-4202-4c76-92c1-6f458ae2bf17 · outbound

This paper cites FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?.

Can Large Language Models Match the Conclusions of Systematic Reviews? FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.491178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.491178Z digest=sha256:dcb76ebea6279f2dfdc80bfdb04f6bad1d56e0d36d5c986a1516b342bebc22a1

Observation d4fce9a9-c421-4326-add9-a0d887425872 · outbound

This paper cites Deepseek-v3 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-v3 technical report, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.578828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.578828Z digest=sha256:2cb9e2fcdf9b8851765ad10d982bdba5298ed3ad0887170a508bc0098702b09f

Observation 0fb222eb-0441-4a2d-a7ac-c47d6547eed8 · outbound

This paper cites The llama 3 herd of models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 3 herd of models, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.620661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.620661Z digest=sha256:9cf9ea6f8415344e1cd95d5377842eab17ee62f83dd99fa947e0c945239b752a

Observation adbe0757-7821-4c02-bd0d-a20069d1bb3f · outbound

This paper cites Qwen2.5 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen2.5 technical report, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.680263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.680263Z digest=sha256:78f4ffc0a4ddcd9ea00798bd1bc55e442dad58b9e7fa292e37cd4d2644f612b0

Observation 5c7b44d0-a5e8-48b7-b2d3-b21b530bc83d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.754584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.754584Z digest=sha256:d17fd071776a702389751ef7a9340016aad5d6bb0d2b72727a328cb50eacd4e0

Pith citing papers

Observation b9ba4ec6-9052-4ae7-a9b4-e667da51b05f · inbound

Treatment, evidence, imitation, and chat cites this paper.

Treatment, evidence, imitation, and chat Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.764062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T08:21:05.698812Z digest=sha256:776ffed72937fcec172a50c40df7306bff56e13b64ea3cf97e3abc98b30bd010

Observation 38004b69-c1a4-422f-b45d-332e23d7fb2a · inbound

Evaluating Large Language Models for Evidence-Based Clinical Question Answering cites this paper.

Evaluating Large Language Models for Evidence-Based Clinical Question Answering Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:57:11.817415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:57:11.817415Z digest=sha256:7b22f1a8bcc3d0e9fef37a4a1f215a5159199d3289cbecd9927a42e08c9c51cc

Observation 0aeb246b-c148-408e-aa59-a286729bfe92 · inbound

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison cites this paper.

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.026099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T06:03:59.798126Z digest=sha256:92346ad4c50b720921aef159ab1b05b43ed1e6f9936fadcdc86e8ec02aebb043

Observation 94e14c8b-1f7f-4c93-9106-9f4d9a7c0e67 · inbound

Can AI Agents Synthesize Scientific Conclusions? cites this paper.

Can AI Agents Synthesize Scientific Conclusions? Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.435176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T13:05:07.718882Z digest=sha256:6b0e5ee8ff71d357a5da3ea7c012f0aa74698efde0e8ff2682503ad36d0058a1