Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Match the Conclusions of Systematic Reviews?

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2505.22787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22787 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:05:45.754584Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:03:59.798126Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.432897Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abfb9d70-afe7-4cb1-855e-741140657f6c · outbound

This paper cites Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases.

Can Large Language Models Match the Conclusions of Systematic Reviews? Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature databases

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.999188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.627088Z digest=sha256:f985d39338ccef0f428448fedfeb4e0e3125c177de71e09ce7dfa71cab4b9d5f

Observation 8d6b0df0-f2ba-4cb7-817e-201fb4cac29a · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:51.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.702126Z digest=sha256:4072ba7382e90e6757738464fabc7cbb3249168c7eedab3cddba03689e6fed62

Observation 8823c06f-d137-4b0f-991b-818d8f54718c · outbound

This paper cites The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:41.788855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:41.788855Z digest=sha256:823c7dc7d15af5e0dc230ab51327f61d12fa3a137c7cd31d3cfb49c443131f83

Observation 22686ff8-3484-41d7-a983-d27f398c74c3 · outbound

This paper cites How to optimize the systematic review process using ai tools.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to optimize the systematic review process using ai tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.728714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.857108Z digest=sha256:60181af3916c2e4a83175633377605b37d70bcd4d4aad7ccd8ad6ad018c74207

Observation 3ef3b132-53c3-4c9f-b155-040c13c904d1 · outbound

This paper cites Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses.

Can Large Language Models Match the Conclusions of Systematic Reviews? Future of evidence synthesis: Automated, living, and interactive systematic reviews and meta-analyses

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.608818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:41.917310Z digest=sha256:8c5b5698bbb7803ef3cab8113f2cc42ec8a28fecbffbc3a54696c563c9da6174

Observation 41609f5e-f087-4fa1-a9c7-11f8ba592707 · outbound

This paper cites Deep research system card, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deep research system card, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.433866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.001265Z digest=sha256:b7236972c01a71f286d288bae6bd8b8507df1e05aecbef72dfdf0f1a74de2f18

Observation 94eaa987-14a8-4cf7-9b2b-938abb39df49 · outbound

This paper cites Gemini deep research – your personal research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gemini deep research – your personal research assistant, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.336859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.069304Z digest=sha256:87f98e1ef28295b8ff6f9513afdbeb8a4b1e847e761a932689e7a3ec32a2fa73

Observation 0f59b63d-401a-4432-b88c-b2d27d28c6cc · outbound

This paper cites Elicit: The ai research assistant, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Elicit: The ai research assistant, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.172471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.163546Z digest=sha256:a285b5654c3d9937b87b64f12e093c09049599c9bcb5ad805579ed325c6beb1d

Observation 00702845-aabc-4db0-8d57-5d07882d3dff · outbound

This paper cites Open evidence: Ai-powered medical information platform, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open evidence: Ai-powered medical information platform, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:51.025871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.237160Z digest=sha256:20648b388e705a52b06be1be6f07663cd3235908984e42a789b5c39cf170c47f

Observation 63bb4f11-83cd-4faa-8b6c-af94bbb4ac6d · outbound

This paper cites Food and Drug Administration.

Can Large Language Models Match the Conclusions of Systematic Reviews? Food and Drug Administration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.851003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.333145Z digest=sha256:2c670ade827c7eceeae923b044066cae5bc0616c3debdc0ed104ffd669b20f83

Observation fd923c07-34dd-4cab-ae8e-3e618ca7096e · outbound

This paper cites Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report.

Can Large Language Models Match the Conclusions of Systematic Reviews? Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.369820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.369820Z digest=sha256:3ba6623edc7ce9b47b159c49fa5535544ad904d0cb646474f9123d8db0476d7c

Observation 35007af2-d8e3-4837-b6b3-21daba273a0b · outbound

This paper cites Can large language models reason about medical questions? Patterns , 5(3), 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can large language models reason about medical questions? Patterns , 5(3), 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.753052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.424003Z digest=sha256:2777ef586b43c43bb3c2599b5814dd07b4fc6b4aa03aeb9ef734f8e529bf4d1e

Observation 182459d2-29d9-417e-9a0e-fbbf805386af · outbound

This paper cites Medalign: A clinician-generated dataset for instruction following with electronic medical records.

Can Large Language Models Match the Conclusions of Systematic Reviews? Medalign: A clinician-generated dataset for instruction following with electronic medical records

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.618921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.533830Z digest=sha256:f50ff7ba5256c98db8bb1492afe2e290d4b04b008477bd5f1789f6a7fbdd7674

Observation 85ac5650-3983-4a41-90cf-dd91d884f93c · outbound

This paper cites Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Artificial intelligence to automate network meta-analyses: Four case studies to evaluate the potential application of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.445363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.587478Z digest=sha256:7d2892bf83e4f375d10b3f56f09299f8dd8178b9737fe7f4101bcfac58786843

Observation 2e81b9be-4c87-4a88-86df-6691597059e5 · outbound

This paper cites Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Applications of the natural language processing tool chatgpt in clinical practice: Comparative study and augmented systematic review

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:50.285844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.700952Z digest=sha256:face9095eaec2a7fb752c264b8ac7ecab1d3385fbf1a7a65fb5e89a8bf4c672d

Observation 9f2c0cdb-5233-4d15-9719-d9efeaca81e6 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:50.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.778063Z digest=sha256:589bb48da51009ea60d3a30696bb2751bf83f9d67f730eb81e659362e2666b8e

Observation 38605347-300d-4d31-a656-f16fa7949a8f · outbound

This paper cites Assessing the risk of bias in randomized clinical trials with large language models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessing the risk of bias in randomized clinical trials with large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.934019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.850157Z digest=sha256:6fb1facbbc0b5c438fc633475d6ee7318aadca9d1565caf69d0ab69375b096e3

Observation 98e86c26-42ec-4b84-b016-a820ada89bb2 · outbound

This paper cites BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature.

Can Large Language Models Match the Conclusions of Systematic Reviews? BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:42.893682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:42.893682Z digest=sha256:1cf081246bc032134e16c2826ef9c5cf6ec713ff284751d24ac70b5a83902e90

Observation d67b7e0e-e96d-483a-86f3-58ec3d060384 · outbound

This paper cites o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \.

Can Large Language Models Match the Conclusions of Systematic Reviews? o ws, Maria-Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel B \

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.738498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:42.975698Z digest=sha256:a0359a87b21563f36ef6c64edc334a60e42c226797df1ab8a654050ee17bde5a

Observation 25f5ae01-0e52-451e-8498-0e05348f9d19 · outbound

This paper cites Generative artificial intelligence use in evidence synthesis: A systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? Generative artificial intelligence use in evidence synthesis: A systematic review

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.615255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.044004Z digest=sha256:43cbe3461d381f2d4383e5b087ffa1954f501733f4630f5b93144dca55a90f3d

Observation ba71cfcd-0fdc-43bd-b69e-7113e17e0ee8 · outbound

This paper cites M ed REQAL : Examining medical knowledge recall of large language models via question answering.

Can Large Language Models Match the Conclusions of Systematic Reviews? M ed REQAL : Examining medical knowledge recall of large language models via question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.417678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.127980Z digest=sha256:1c72ac19aa07eb58cc2ec6224e5104db4cb0358a6bdb2350c0c97bed123c6e4a

Observation 4ee122c0-75c4-4ee7-9002-b1a475e4b530 · outbound

This paper cites H ealth FC : Verifying health claims with evidence-based medical fact-checking.

Can Large Language Models Match the Conclusions of Systematic Reviews? H ealth FC : Verifying health claims with evidence-based medical fact-checking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.287045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.193185Z digest=sha256:85e79569177d8a9f51cb3903a513ba3c33134e4ec75bda5bb36c16d8f033174c

Observation 116dc038-7c89-44ff-94f7-91293199787a · outbound

This paper cites What evidence do language models find convincing?, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? What evidence do language models find convincing?, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:49.071576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.238814Z digest=sha256:43a700998cf5a94bebe519659bb345d13ffcc966df3f480c0bd83f55d48c5e77

Observation 1d2fea88-ace9-4c10-86be-61cdf76911e7 · outbound

This paper cites Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Clasheval: Quantifying the tug-of-war between an llm's internal prior and external evidence, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.888788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.332339Z digest=sha256:0ded3d2a489a8a6ab9ef0ff945f967396078443df7690926d6c55334bfb1d8be

Observation 2fe2347a-368f-481e-9843-adec69ddf050 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.713111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.408496Z digest=sha256:a0507396c5b697297bb65771b941b1fb339dcf64e68f18523919b6b540b42c0f

Observation 153c3248-884e-4b22-b754-fe928d4bb7b1 · outbound

This paper cites Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Untangle the knot: Interweaving conflicting knowledge and reasoning skills in large language models, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.535178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.481306Z digest=sha256:e1bcd64dcd7ab2f7ca88357acbd806b7362d99eeba4289353e6ec94f7c6c7c34

Observation c70947c1-3b0c-4aa1-b33e-b3a69558d131 · outbound

This paper cites How to write a cochrane systematic review.

Can Large Language Models Match the Conclusions of Systematic Reviews? How to write a cochrane systematic review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.332514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.544854Z digest=sha256:9a3ba304452cddc3b7b6fc40c2e236b31ad1fb9551d19d62afb57c605e3d893e

Observation b0a7dd8d-16d1-4069-84d7-ff2ff483006e · outbound

This paper cites Quality of cochrane reviews.

Can Large Language Models Match the Conclusions of Systematic Reviews? Quality of cochrane reviews

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:48.191484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.584060Z digest=sha256:c7a7158cdfa64f4e36b01efb792f29cd5a02e3d391c4c0960daf3533923b6fd9

Observation 0ee4b7c3-0587-4990-a353-0b1af38a86f4 · outbound

This paper cites What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011.

Can Large Language Models Match the Conclusions of Systematic Reviews? What is a cochrane review? Epidemiol Psychiatr Sci , 20(3):231--233, Sep 2011

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.996527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.680202Z digest=sha256:ee1887c81ebc548641db8ae2306699274a0947c835dc479dec73df5baeecb56e

Observation f5763b0a-a6c1-4835-8b0a-348f3ea98e09 · outbound

This paper cites Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.846344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.833178Z digest=sha256:b8833fbd42c82b6301c226a726fa94d0a530c9429c70651339e04e7208dfdb43

Observation 57ebdff6-cfed-4775-83c3-ef9ba273a830 · outbound

This paper cites Bethesda (MD): National Center for Biotechnology Information (US), 2010-.

Can Large Language Models Match the Conclusions of Systematic Reviews? Bethesda (MD): National Center for Biotechnology Information (US), 2010-

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.705602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:43.916594Z digest=sha256:3b74e69553feb3c930a47b215e1d15a69ea7202bc5d1397b6cc1f585ec5c5972

Observation ddc25f37-4587-4ab6-9032-9003f348532f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Match the Conclusions of Systematic Reviews? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:05:47.583430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.012378Z digest=sha256:139eaf5e977fc5e4a1171edec37d6244de1f6a121ad955378b00baa58db89ae2

Observation 2eebce0a-e7a3-4f98-b7aa-e7752544f1f9 · outbound

This paper cites Assessment of the strength of recommendation and quality of evidence: Grade checklist.

Can Large Language Models Match the Conclusions of Systematic Reviews? Assessment of the strength of recommendation and quality of evidence: Grade checklist

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.424267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.097237Z digest=sha256:165059980275554fdb4237cddf66df243da329edc6328395c0321b6036688be7

Observation ff297756-cfd5-48f9-916d-c8f0ef492680 · outbound

This paper cites Openai o1 system card, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openai o1 system card, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.179454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.179454Z digest=sha256:8c6cf959863fbec79e7ff50b9dda864b2e1de8e6195ee014fc197d0de7189056

Observation 052921c6-0370-49bf-ad06-b2da74cec50a · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.270040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.270040Z digest=sha256:d1dc271af2523210e3b2435bdec2ef5f8376ccc7c8543ac36b63a24445cc361e

Observation a2576a7f-030f-4e82-b9a4-b5681c7e2546 · outbound

This paper cites Open Thoughts.

Can Large Language Models Match the Conclusions of Systematic Reviews? Open Thoughts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.248285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.358789Z digest=sha256:7b734c86e7238cded5a7cfd5c0cb0f67a527b2724ae35d132905640612b77750

Observation 74080c23-941b-4d06-9368-74c6f7e90f55 · outbound

This paper cites Gpt-4 technical report, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Gpt-4 technical report, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.428543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.428543Z digest=sha256:0b1c57aacdf8c6e0d3063b99a745bae658823034911eebcc80770637bdbd1ab7

Observation 305a4698-4bc5-4ff5-8564-cad8cac40ee3 · outbound

This paper cites Qwen3, April 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen3, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.486772Z digest=sha256:4e76528821fbdff6720fa087379052c49ed1f198c4f0c41b263d07c3cbbf7edd

Observation 4f7c6598-9e4c-44fa-b82c-930f310679ac · outbound

This paper cites The llama 4 herd, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 4 herd, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:47.057232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.547522Z digest=sha256:10c21780a2a4b9d0bba04e1fd57cd25ba92457cbbb3eac206907ecbba9bb32bd

Observation f7222a5f-1b41-44f6-817a-c64f0cd26817 · outbound

This paper cites Huatuogpt-o1, towards medical complex reasoning with llms, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? Huatuogpt-o1, towards medical complex reasoning with llms, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.623175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.623175Z digest=sha256:f32800047587b5cc11d1fdcd206d10d7f82ad434b37743af446b38de39bef19e

Observation 935206a9-ac73-4ecf-90e1-0bddbf237818 · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences.

Can Large Language Models Match the Conclusions of Systematic Reviews? Openbiollms: Advancing open-source large language models for healthcare and life sciences

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.888598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.724796Z digest=sha256:ed01c6e58bb5ec22e0a5bc7df93d069566561aa7bb6c0b51359d581244db3d61

Observation 01a60bc0-9a29-433e-9a9d-b6db76a4c5b4 · outbound

This paper cites Refinedocumentschain.

Can Large Language Models Match the Conclusions of Systematic Reviews? Refinedocumentschain

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.742395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.817332Z digest=sha256:a69e18e2bd1cba54ec9105004622536a19b6984a03d99e3a1e59f549c8fd0fa2

Observation 0b21fda6-023b-4ded-8cba-254e44ea342a · outbound

This paper cites An introduction to the bootstrap.

Can Large Language Models Match the Conclusions of Systematic Reviews? An introduction to the bootstrap

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.604185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:44.916032Z digest=sha256:8d9dfca17e90d5dbc0fc845ef0134e44c4f99817a98caadc83144128bd16316a

Observation d6ba0e78-aec5-4990-8dad-eee64931b492 · outbound

This paper cites Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:44.983710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:44.983710Z digest=sha256:8fa7f92bbca34ef1187cde3ae40caaf924be8ab40d2e70d5cadc27defc13ec5d

Observation a9a25759-e103-4159-9809-9ffbe1bc9358 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Can Large Language Models Match the Conclusions of Systematic Reviews? Long-context LLMs Struggle with Long In-context Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.055957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.055957Z digest=sha256:25c2be9ac751809bbd4dd31901cba7b1857e131ab34de19b32a70ea0e9873cec

Observation c1bc1f40-dfa1-4665-8f19-8e8aece3d297 · outbound

This paper cites Large language models are overconfident and amplify human bias.

Can Large Language Models Match the Conclusions of Systematic Reviews? Large language models are overconfident and amplify human bias

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.149556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.149556Z digest=sha256:c607580f9cd313f6c6bd1f70631d10dd99686648d66da99fa64131627577c243

Observation f740f987-927e-461a-aba9-43f980797571 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Can Large Language Models Match the Conclusions of Systematic Reviews? Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.189971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.189971Z digest=sha256:ecbffc93fb0559f71d1342fd9cd18eed13e1d69baaede06ad188ea6fbbd9f585

Observation 607a5fc8-c3fc-456a-bcae-c39cd7a68ee4 · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Can Large Language Models Match the Conclusions of Systematic Reviews? Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.303527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.303527Z digest=sha256:9ed2f24c7dea9667e6e4925551766d58f4ecc212eb448495c4bd6c38bf370b8d

Observation c88dedd6-9da0-4f3b-bbde-83e6c12e97fa · outbound

This paper cites Fine-tuning is fine, if calibrated.

Can Large Language Models Match the Conclusions of Systematic Reviews? Fine-tuning is fine, if calibrated

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:05:46.443913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:05:45.381013Z digest=sha256:8d62fb4e788fac99b19613b1e93a178949542cfb767daf231b820c20182e8cab

Observation 8ec6626c-2a83-4b3f-978a-1cc08e17c7ec · outbound

This paper cites Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data.

Can Large Language Models Match the Conclusions of Systematic Reviews? Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.447639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.447639Z digest=sha256:de0ed0d1276fa462c09e3ce63b8de2431ba3d6e081ca9884f313e4a6b1904a7a

Observation fd2586f8-4202-4c76-92c1-6f458ae2bf17 · outbound

This paper cites FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?.

Can Large Language Models Match the Conclusions of Systematic Reviews? FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.491178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.491178Z digest=sha256:65617130fb969661dc19bcaf543c6d05180a4c8a5e3510eae5073e7bb0ce0caf

Observation d4fce9a9-c421-4326-add9-a0d887425872 · outbound

This paper cites Deepseek-v3 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Deepseek-v3 technical report, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.578828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.578828Z digest=sha256:1203dc1afa7671e46e63a1e992bfde9e675b9fde934ea0fe7fc53fa39582661a

Observation 0fb222eb-0441-4a2d-a7ac-c47d6547eed8 · outbound

This paper cites The llama 3 herd of models, 2024.

Can Large Language Models Match the Conclusions of Systematic Reviews? The llama 3 herd of models, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.620661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.620661Z digest=sha256:b6445a10c6ec44e403c920ad70b9d0dfc228b6cba3434bbb0e0ed2884d2474d6

Observation adbe0757-7821-4c02-bd0d-a20069d1bb3f · outbound

This paper cites Qwen2.5 technical report, 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwen2.5 technical report, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.680263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.680263Z digest=sha256:7eb34846e07dd82adacd06b83a5cc177fc8ed663490147bd4ba9401358628ae0

Observation 5c7b44d0-a5e8-48b7-b2d3-b21b530bc83d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Can Large Language Models Match the Conclusions of Systematic Reviews? Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:45.754584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:05:45.754584Z digest=sha256:1fbb52b4a218cde8967a3fe05905a6442155127558cba1c4e9f28aef7ebfb6c4

Pith citing papers

Observation b9ba4ec6-9052-4ae7-a9b4-e667da51b05f · inbound

Treatment, evidence, imitation, and chat cites this paper.

Treatment, evidence, imitation, and chat Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.764062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T08:21:05.698812Z digest=sha256:a8a9203a5cebae5724af76d2afd90c894f803d7c9de23777d4b4479866e084a3

Observation 0aeb246b-c148-408e-aa59-a286729bfe92 · inbound

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison cites this paper.

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.026099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T06:03:59.798126Z digest=sha256:a9a71b9919d9fb9fdf64910d094c8240ac43795e9fa2d4261b16fca0f67f1ffe

Observation 94e14c8b-1f7f-4c93-9106-9f4d9a7c0e67 · inbound

Can AI Agents Synthesize Scientific Conclusions? cites this paper.

Can AI Agents Synthesize Scientific Conclusions? Can Large Language Models Match the Conclusions of Systematic Reviews?

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.435176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:05:07.718882Z digest=sha256:b23a5d9546800b6bdd30e4a67a19f99036a7257a8a4c8104ed72a9d94af1be07