Pith. sign in

Paper Citation Record · LEDGER

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

As of 9 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2507.15152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15152 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:46:08.376749Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:53.040834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:13:54.647415Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact16
  • verified fuzzy38
  • unresolved37
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59abdfec-2b79-400f-abb4-6896d381a4af · outbound

This paper cites Research Synthesis and Meta-Analysis: A Step-by-Step Approach.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Research Synthesis and Meta-Analysis: A Step-by-Step Approach

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.653529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.653529Z digest=sha256:639b4dcbf799294d50da285799657c6fd783cd00b73c2bad74cb37dd41fe0af7

Observation 1f67cc0f-6e6e-4c9c-8dd2-dd9074bee257 · outbound

This paper cites Analysing data and undertaking meta-analyses, chapter 10, pages 241–284.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysing data and undertaking meta-analyses, chapter 10, pages 241–284

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.664179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.664179Z digest=sha256:3b37a9b7c0a47e596075aa9a07465d06433149eeb87e78bbb4fc30dc17165f8a

Observation d0762771-aa93-47e0-aeb8-895ffc8ebf2c · outbound

This paper cites Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.671805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.671805Z digest=sha256:eb6c02146846e539872225b1dadcc9e123b2a68f49ec7ce6819b22ee34a33759

Observation efdddffb-8142-469b-909e-5f43339c72f5 · outbound

This paper cites Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.038582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.680841Z digest=sha256:9618006dbc8b952681ff0dded4c02b1bbcf98e6b44e31c1b7d4c87293e360139

Observation 70ea14fd-5453-4d30-8447-05c6b28e9f97 · outbound

This paper cites Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.020674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.685632Z digest=sha256:4c8d4b6e0bae4bc506483a5e63ba49d6a57f9e5831423b67e31a1432c9186914

Observation 975c4024-e2d0-4308-ada1-f84d80119446 · outbound

This paper cites Toward systematic review automation: a practical guide to using machine learning tools in research synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward systematic review automation: a practical guide to using machine learning tools in research synthesis

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:07.690733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.690733Z digest=sha256:8ebc76cf468c49eaf27c3d47d8ad007cc0274ca4c47cb06784d12e77480c14bb

Observation b0629e68-bf64-4206-99b4-91252af8cde9 · outbound

This paper cites Exact: automatic extraction of clinical trial characteristics from journal publications.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exact: automatic extraction of clinical trial characteristics from journal publications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.697875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.697875Z digest=sha256:e7ffb7edaa1c88e58960acaea4e6b704994f4dbf134abc99923db56fd95ee98e

Observation fdd4bafa-4184-4928-a2ff-b8fe42880703 · outbound

This paper cites Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz

Reference 8

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.968448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.746197Z digest=sha256:f8360c7b1c04ea8f75496781310e57939dc440b566433d2d3217181a8319258f

Observation ca059850-7817-4bff-a971-e79644f87bd6 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 9

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.459526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.763071Z digest=sha256:db8d620e09663eca974785ff99d646b961fbd40c3622942717953127e48c09ba

Observation 1c86dbba-aac8-4c3c-ac6b-60cf0a5c8cc7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.784022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.784022Z digest=sha256:e040683e8aa09736cfcd3e91dafc843fde6e4b15add76257205865c8d18a383f

Observation bc08e3c2-ce9c-41ab-ae05-793e5c754ed0 · outbound

This paper cites Automating meta-analyses of randomized clinical trials: a first look.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automating meta-analyses of randomized clinical trials: a first look

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.801622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.801622Z digest=sha256:2ca9e6548649e01ec81d8d6f2d9e872910669c510d41b72da676da64dd0f8320

Observation c032b9f2-8819-4519-aed0-7709bda04041 · outbound

This paper cites Katz-Rogozhnikov, Kush R.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Katz-Rogozhnikov, Kush R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.812345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.812345Z digest=sha256:38a3483f0949aa91bc3d44a4eb7008a3dcb2e9a841d02e6e09930bdee7ecfa36

Observation ec473968-5324-44d2-a235-b93418bdb282 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.817822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.817822Z digest=sha256:05fcc97b27792fff66284872de4bccb7a88b6c7790c499bb476c95fcbea1f1e5

Observation 25b1e172-1af5-4a79-be85-76ea657b1997 · outbound

This paper cites Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved].

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved]

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.928184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.824462Z digest=sha256:e12bf496d8f1cef74d188638de46d43b08e759e8bd81f3b9b182e4c2e9036b66

Observation 094ae696-9487-4bcc-8bfd-ea6e8255c58e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.829830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.829830Z digest=sha256:0911d941d1a06b0251c4cfa3023b686e99e9a5fdb0a76f161152c54962410fee

Observation d8fbea24-066b-455d-b1ca-bda23239e5b1 · outbound

This paper cites The data is in: Deciding when to automate screening in your slr, November 2023.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The data is in: Deciding when to automate screening in your slr, November 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.834832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.834832Z digest=sha256:8252a434e26f9d1645eb08e1e2de831ca37687d27de1190967117a553dc9d07b

Observation f63edf05-4f58-41af-89ae-082feb262b6e · outbound

This paper cites Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.839801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.839801Z digest=sha256:a73a5f64e30813978aa1bde1270db9206a87ad65ca174883017347affc4fbf0b

Observation 28e34ed0-8485-4fcb-8124-2dad653175dd · outbound

This paper cites MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.854237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.854237Z digest=sha256:c47b61d707ec0bb23bce5fe95673af1e9e10a50dccb0545c5c7b20768edbc1ca

Observation e7c78217-16e8-4a6f-8fd3-64c498887a46 · outbound

This paper cites Chatgpt: Large language model (mar 14 version).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Chatgpt: Large language model (mar 14 version)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.864218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.864218Z digest=sha256:401fd591fbfca4c2927711778fe048a3f908409bdf9cc232c303760484c8fe80

Observation cac6f8e3-00ef-4460-b6f3-4ea7f1d4f000 · outbound

This paper cites Claude 2 model announcement.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Claude 2 model announcement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.870017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.870017Z digest=sha256:d2400a2a74d14d7a1efdfe113c16bb7ef4de82257578bd530ff9025ee2b4dcae

Observation 583f778e-b019-49a5-aa43-88fadb7779a7 · outbound

This paper cites Zero-shot infor- mation extraction for clinical meta-analysis using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zero-shot infor- mation extraction for clinical meta-analysis using large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.875116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.875116Z digest=sha256:c3d9f16870b383e0f2adee3e8bb4136fd618db8f91ef161319fd2706a564418e

Observation 75556778-c65e-4b3b-8f2d-8b86090c2e79 · outbound

This paper cites Performance of two large language models for data extraction in evidence synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Performance of two large language models for data extraction in evidence synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.880456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.880456Z digest=sha256:a6f0a06b45c0fb52087f50c5d82b176d8901f115e37e81deac613e081578480f

Observation 6705cf30-2390-4d27-9f7e-b02c0a6a2512 · outbound

This paper cites Automatically extracting numerical results from randomized controlled trials with large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automatically extracting numerical results from randomized controlled trials with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.885286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.885286Z digest=sha256:719fb4e2d9d2b5f704df4dd14cc09e9b3b56b1f9a260404897763e099220cc8a

Observation fe259602-0acf-4aa5-9e96-25afc38508d9 · outbound

This paper cites Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.891121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.891121Z digest=sha256:1bd7aa5de4a51dd671fd466a5d1925bed89db0700f50150a46f30ce69ec51b32

Observation 94bbea6d-b5f3-4320-89d7-15055bdf1e24 · outbound

This paper cites Lee, Shigeki Yamada, and Tomohiro Mizuno.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Lee, Shigeki Yamada, and Tomohiro Mizuno

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.896068Z digest=sha256:d6cd466b4b2f98154fb3f55e687200fb031ecc23509648373a3c9cfd689a3017

Observation 1cfbaa80-85fd-4a5a-97d0-24a0a022a479 · outbound

This paper cites Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis

Reference 26

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.370112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.901443Z digest=sha256:ab5e7ac15ee44c7691a496566f2719bb7ae56961d412aefe572f3fc097c0764a

Observation 96dc95e7-b78e-47d4-bdca-e36962197bcf · outbound

This paper cites Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris

Reference 27

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.829269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.906139Z digest=sha256:6d7c4453a358de4b186f7606e54b1032b24fd60d3a487ac1f8685a38387f74c6

Observation 64105b66-d969-41af-a918-8da6e2c4b62b · outbound

This paper cites Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.285708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.913293Z digest=sha256:3c2f8ff491fde5132137226bef4266d0f416e16cff1785c4626debbad4e5cfa8

Observation 8b698dba-6c1e-4100-a69b-c5908f7b0a20 · outbound

This paper cites Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim

Reference 29

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.799935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.918186Z digest=sha256:972682a211c4f99e39dc40874a0557130fcea25f7398384de95e5a7575a87c04

Observation 5c6db0f3-9237-4145-a5d2-9b5c829d2f1e · outbound

This paper cites V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.774773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.924380Z digest=sha256:238d08371efe0fd6d16d4bd8b7d9413adbd37e2f9f64bd8ffe7d13304b2d77ee

Observation 34192b22-a459-443b-8287-0cf5870d8db2 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gpt-4o mini: Advancing cost-efficient intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.934074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.934074Z digest=sha256:7f90e367a9ba3813553d7221336bfc3b292741627e791e97af955ae5ba51be24

Observation c32e8170-e0d4-4921-9966-1a46b81557ab · outbound

This paper cites Gemini 2.0 flash.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gemini 2.0 flash

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.939758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.939758Z digest=sha256:0c4955e84ea5972c56d24acc0348fa7969068b82c41d8a3ab776f1596a4ccf8d

Observation 93682bef-d44e-4c7c-8a57-315a4becb015 · outbound

This paper cites Grok-3 language model.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Grok-3 language model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.945503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.945503Z digest=sha256:f464cfc00e186c919928197ff918c72be49b4c6d3af40fa71830e5481a56234a

Observation 8b9cbaff-4f41-4132-857c-c60de2b330d0 · outbound

This paper cites The impact of temperature on extracting information from clinical trial publications using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The impact of temperature on extracting information from clinical trial publications using large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.951933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.951933Z digest=sha256:a492467f1a7ac709f2ab35667fc444302c0656f67d5a9b7533aa1019b4cedc95

Observation 4bbf6358-214b-43d1-afdb-4f0aea3b2d51 · outbound

This paper cites AI-Assisted Data Extraction for Systematic Reviews in Education.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction AI-Assisted Data Extraction for Systematic Reviews in Education

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.958889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.958889Z digest=sha256:a84bd3a7aa91a122341dd0e204c4988e783da0c9cc3a6a7797697efa10c625f3

Observation 178abf96-a100-40cc-b4ea-fefa51c92b6a · outbound

This paper cites Use gemini 2.0 to speed up data processing.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Use gemini 2.0 to speed up data processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.965054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.965054Z digest=sha256:162ab3cb95073d5780701bf1abbab585ba605a427a631f125f97749d9012310d

Observation f654e4b5-931a-4750-a1f4-2c9e8ea67a82 · outbound

This paper cites Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.748746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.971513Z digest=sha256:ef185affb96498f31ad6bef249b85f3700cf3402986914596f36c755754aea9b

Observation 759dce94-632b-4544-a756-d882cdaf18b0 · outbound

This paper cites Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.347631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.976676Z digest=sha256:3d245fdb3656b56994747694e1fdf3258ef6a8ba0f41b40c811e4beec897345f

Observation fc9056f3-42b0-4a30-89cd-34b59992be83 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Reflexion: language agents with verbal reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.318292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.989820Z digest=sha256:6b2855d080c15f0f93351489a24ba6af96bc3f58c259bf7204427fe20c5ffe12

Observation 2acec01b-d982-4110-aa42-d3a6d238ea8e · outbound

This paper cites Towards mitigating LLM halluci- nation via self reflection.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Towards mitigating LLM halluci- nation via self reflection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.994474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.994474Z digest=sha256:8e7b7a21c8d07986dce489ce7f933cbd5a0f2891f7fe7705eb0dd3f80f1357f8

Observation 6c7a93d2-95b1-41d6-8356-4ae8b04c0c95 · outbound

This paper cites When hindsight is not 20/20: Testing limits on reflective thinking in large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction When hindsight is not 20/20: Testing limits on reflective thinking in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.295126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.000826Z digest=sha256:f25870cbaf22c9e264e8bec02083ef73fe64d79cd65c74433510aaa147f35466

Observation ccc81064-a620-4159-a80e-ba3512ca49d4 · outbound

This paper cites Dietterich.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Dietterich

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.267940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.011272Z digest=sha256:b1725bfe6119d7ae97b6a853c005bb5b0fe5854ca14b5d3d84882d0714c50e42

Observation 23f4b757-ac28-47ef-8306-fc3ec919bcb0 · outbound

This paper cites A survey on ensemble learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction A survey on ensemble learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.018396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.018396Z digest=sha256:27224e15a2498114d9e6e67d7b3dafd0dbbea0391d86774b98fbb8d5fb4efbc8

Observation e84df5e4-15d8-4caa-aca3-aba9d2a57a96 · outbound

This paper cites Ensemble pretrained language models to extract biomedical knowledge from litera- ture.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Ensemble pretrained language models to extract biomedical knowledge from litera- ture

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.645522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.023179Z digest=sha256:bf034a02e2861e6a224a84877bcf747673dd624046aef4a8ed131736333fee97

Observation 8d7ef2bc-7931-463f-922d-5202489506dd · outbound

This paper cites Zhang and A.L.P.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zhang and A.L.P

Reference 45

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.618170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.030348Z digest=sha256:4d4b003985fa3f181b43d5fcd2ee8bf98763459ae76340d15610691375533923

Observation 31adf290-91b9-4d02-862c-b19873f71cee · outbound

This paper cites Comprehensive testing of large language models for extraction of structured data in pathology.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Comprehensive testing of large language models for extraction of structured data in pathology

Reference 46

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.587520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.036361Z digest=sha256:1b7941573c3d5ab5eb14f29ee60bcfb7d6838780baf76e84185285ed9979d395

Observation d28dc553-4477-4edf-b7d3-e393a2f59d3f · outbound

This paper cites Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.546789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.041977Z digest=sha256:5fc771d8065f6fe5ed13e075aba5eafee1c7bf17ff34d63a58b577ec568e2d13

Observation 18e1f353-74a7-4807-89c8-99ee1df24e4f · outbound

This paper cites Match, compare, or select? an investigation of large language models for entity matching.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Match, compare, or select? an investigation of large language models for entity matching

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.243358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.048626Z digest=sha256:9b54da5f89df15a3d243dcaa25bbf7e52c7756bd7bc69d425266b2ad9fca67ad

Observation cc3c1a78-35ff-4e69-a324-9c56813ca94d · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 49

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.516502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.053892Z digest=sha256:c4e88228d0350a1dad8e7b83efe95e05e4b5a4294d11ea771d52caf125993ea3

Observation 84af77db-91b6-4c1a-b14e-2095284f718e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.061584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.061584Z digest=sha256:7c57be9ab31044933123a048228ab89a77c7fa408b708cf7b8075bfa0d17749d

Observation 3f397eca-7685-4cf9-8b04-69ba6bc821d3 · outbound

This paper cites Guyatt, Andrew D.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Guyatt, Andrew D

Reference 51

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.492238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.068084Z digest=sha256:02fe3022bf0dde17f2e01e6393d2186b943ae248d7111f0b12c7e8ad8642d98f

Observation 8614a940-0e89-4af5-add4-f99fb77eb40c · outbound

This paper cites Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.075310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.075310Z digest=sha256:b06c470b8d427b8fc864e53a486a4cbb7bb4608becb279db5a5f0f15e0b7f59d

Observation 449df189-eab1-4838-8f8f-8e5dc1abdc35 · outbound

This paper cites Agentic reasoning: Reasoning llms with tools for the deep research,.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic reasoning: Reasoning llms with tools for the deep research,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.198891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.083557Z digest=sha256:a2533739c72d62408cf326c82215cff0b3b0d1892019334d5ca43f3d0d874dc3

Observation 2aa669b4-d09e-46ea-aa25-0ac7f3da646a · outbound

This paper cites TART: An open-source tool- augmented framework for explainable table-based reasoning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction TART: An open-source tool- augmented framework for explainable table-based reasoning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.171211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.099934Z digest=sha256:8d770aad7dafbd5d5a928dfbe80e4ddbff5e80feeb23449764e3924208290452

Observation b81958ca-a660-4689-bb8e-db65fdbe01a3 · outbound

This paper cites Medical hallucination in foundation models and their impact on healthcare.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Medical hallucination in foundation models and their impact on healthcare

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.105709Z digest=sha256:183c7abfc8016c32a4372caa57f5fa631fac009d1fbcd553e3a9ec4a190b4ebc

Observation cabcab11-7bd8-46b3-ae71-b44625490a07 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.111858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.111858Z digest=sha256:22c6d2b42050e9b94a6bc5341ddd6fb330f5847ded1066f968aecabd641340bc

Observation 2b1898d8-f3f5-489c-9aee-2d67e449992e · outbound

This paper cites Potential roles of large language models in the production of systematic reviews and meta-analyses.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Potential roles of large language models in the production of systematic reviews and meta-analyses

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:08.118777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.118777Z digest=sha256:22e85f990ff39faae0866c135c6f85988d5f66aca6759b4755da316410f26852

Observation 93db9ab6-c6b9-4861-98f0-f20db0315ff6 · outbound

This paper cites justification.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction justification

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.153899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.124679Z digest=sha256:eeac9bc3356875bb159fcde922dc0befaf78dda1b0647a0c049d9e98ff1b81ed

Observation e1d54069-2ad1-44a4-9a0c-74c7a7d12d35 · outbound

This paper cites other_time_points.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction other_time_points

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.136844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.130975Z digest=sha256:e6b647619b080a93d9a0c267644f0541d5fab5b486ead91d95d168047ffe4496

Observation 69682989-e509-4562-9ede-0699148ce6d3 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.117550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.136861Z digest=sha256:b14f5c595539b8437ab938d98f044c541327b2bdbfe51e21e023bba65a9a0a6d

Observation 5af7e017-6330-4c4b-9a15-b03cafc5d496 · outbound

This paper cites null"`, NOT `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"`, NOT `

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.100782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.142821Z digest=sha256:74741f1347339824d2907ecfcd86dcc02f1f89475198c2acdace173eb9376fe9

Observation 502eff7e-a05c-4706-a710-8504908387b9 · outbound

This paper cites more common in the intervention group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction more common in the intervention group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.078892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.149296Z digest=sha256:507b75b574a50c3db78a034f357c7f6a5569cd0316379a53546318a15a2086d4

Observation 03ae9405-e1de-40c5-a58f-b5f4d3b0af26 · outbound

This paper cites data_conflicts.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction data_conflicts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.058081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.156022Z digest=sha256:f60f072c26335d3abe5f981437997082938232ff2ada6c27f75d80d856d758fe

Observation cbdb9462-4a71-409a-9055-a8c6b0f3411e · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:46:10.036874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.167425Z digest=sha256:9448700cb62ecc8b56e6500dca6f2ca9559576cc04a206843864439665177564

Observation be175a7f-d594-4d4a-b249-ba68e6a3e99c · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.018641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.175488Z digest=sha256:29b92c16fc44f82a382bc5289eadddd5cf7a5081691f2c49ab073362807052ef

Observation b871bc30-9e35-4490-80b3-93fad8c85ad0 · outbound

This paper cites - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting)

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.000536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.180143Z digest=sha256:edc639ae6683b240fdb08488d70fcb60064bad7cf699bc1098656c8ad6b439dc

Observation 011d302c-57c1-4b6c-9e08-c0b603624a95 · outbound

This paper cites - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg)

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.981621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.185009Z digest=sha256:cb3b83a62f028eb491142de8f04e16b1365cbfe3851ea7cc4a4fc79c69d22a4e

Observation 7e25a773-219d-4351-b77e-f33ec5ef81cf · outbound

This paper cites - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.956642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.190287Z digest=sha256:f1cecdc023731a134edb6c3045dbe63c49804cf9358f427544086389108dffc4

Observation eeed05e4-b1e0-4c03-8474-33f38239b288 · outbound

This paper cites source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.929667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.194780Z digest=sha256:8a1401b825eaa3a4fbff589f4af6578cbf3187529aa20afec38ab06971a178c2

Observation b82d1c6d-022f-4f65-b6be-1fa851dfde95 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.910425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.200814Z digest=sha256:6386b1a8f181aeb8053edc0605608d687cb6a75121e25e7da5b0b8f7d0028ad2

Observation 3c9eca83-c241-4135-8854-e64f65e33fec · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.890938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.207241Z digest=sha256:afa86da714bf8d70d182d620821772bd94a5dc061544aacdba77d426955cef5a

Observation 8fdb2b6c-3a11-49f8-9877-ccc8cae1ee5a · outbound

This paper cites null"` or `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"` or `

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.871709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.212807Z digest=sha256:ebca25db33ff03617d64d629438f0c2804707f39217d4ac9027c48b6f3dc1294

Observation 92559793-e3f6-4e01-bbf9-3f5fdee6ad67 · outbound

This paper cites revised_value.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction revised_value

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.855205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.217778Z digest=sha256:493e68421b2ea0086eaa7223e1f7564e71dd9158c84937ec54f86057e413399d

Observation 9fc1478b-13be-463e-83ff-8a368299de13 · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.835083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.222572Z digest=sha256:23a25e07bd5154495f0029bf4bf77921bd535dcedb71a372d274134700053c26

Observation 90943ae5-6f81-415a-8a34-8356cbc00d3f · outbound

This paper cites confidence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction confidence

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.813283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.234818Z digest=sha256:3296847fb330136e43b680549fce19841f8d356e55fb5d647e1e99560668dded

Observation aa8d47f0-2708-41cd-a7a1-9c15b0df1e89 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.793979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.240837Z digest=sha256:24cb0b5c0428e25134d9130fa72878b75fa22eafd1a3f506f53ff24091b82b9a

Observation d2d90d35-360b-48b9-b65e-437e63df213b · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.770068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.246699Z digest=sha256:d77da9fbeac0f54699f5bb7ca87ec4ee3603c44ac6b15a20af055d53310e36c4

Observation 41fcd1e8-5048-4941-909e-7a53cd38dea4 · outbound

This paper cites Just return the final merged JSON object.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Just return the final merged JSON object

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.753155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.252902Z digest=sha256:7cf0d5450e4dc7d0a64aa68d4475352aea86dba425090919faab2e7a881ea631

Observation 9385cae5-652e-45cb-8c22-60b0402bf079 · outbound

This paper cites LGL_group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction LGL_group

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.713347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.270446Z digest=sha256:7da67ecb4b622759c27773ec4b623b34454ceeea1a256297c7cdc172c3a51f43

Observation bd83a4e9-a0df-4780-ada0-9ec931fa8fcd · outbound

This paper cites **Note:** EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** EXT fields may be nested

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.694694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.277351Z digest=sha256:8be6d66c39231fb20f782fa418c988b328851172a09156e373afbbcabb01c2a1

Observation 6c0500bc-867b-4458-806d-71ba0e2a6d13 · outbound

This paper cites kg/m²"` and `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction kg/m²"` and `

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.679005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.284067Z digest=sha256:7ad8eb081bf717d9690d7d66674fd6d88eb82e77bfdcd4089ca8faa3b01ccdd9

Observation 7176815e-0209-41f6-8c8e-5775018994e1 · outbound

This paper cites low glycemic load diet.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction low glycemic load diet

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.664051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.291370Z digest=sha256:439b19a5e783019c448c0cbc8ad54658b44ccb8007b1da983158a1c0ed920b21

Observation 17f12667-e5e3-4b97-af9c-4857f228a268 · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.649218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.296906Z digest=sha256:79d30e749b69b5214de0c91a43e674684b7204c5673ba72b524f30aab6d19521

Observation 1be37d85-a7de-4cf8-8428-1f2016d914b7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.635116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.305674Z digest=sha256:9d0d8d41999bf7bb41c9857f214378da58d64777a983677b5b8bf07e0a19fefb

Observation 2edfcde8-35c9-4f4d-9765-3301503073e9 · outbound

This paper cites randomised controlled trial.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction randomised controlled trial

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.620028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.316558Z digest=sha256:7df5a9c5fc744993dfc4c35d96bfb19660a011774bb46ffddd3a96bd2b1119be

Observation fc808152-11b7-4421-add0-9536fa7c8b8c · outbound

This paper cites You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.605521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.322330Z digest=sha256:6847004724b587ac60dffdcfbf5c7215c158afc35b2cef7aa8af9e6028ae2349

Observation cdf87145-1676-4d00-b375-3100fed7d3fa · outbound

This paper cites not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction not reported

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.589766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.328561Z digest=sha256:99182c056e7e7fdc0834192830388e8e4d7318f5bef6ff0f6a77b199f92685f8

Observation 685b9bca-5db9-43cd-9dad-2c4663eac3aa · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.575016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.335133Z digest=sha256:e34b7b213df91408fc8d2063ddc40d22e26d82c0b5335930efcc936bf1801549

Observation 8c836ecb-5ab3-4107-a7c8-4d359f7eae8a · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.558785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.340397Z digest=sha256:290252cc3ee97f558eb20b73b3ed1417092d22bacf92500ced54159154330010

Observation 606b3b47-a88b-4036-b68c-99e978ad0083 · outbound

This paper cites This is the preferred method.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction This is the preferred method

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.731507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.345245Z digest=sha256:f1fafad6d92ae0f189a1520fa2e558740daa73c06f4f61aa002da043fbeca247

Observation eb4f5bfe-8964-4cb9-82f6-44340197761e · outbound

This paper cites study_characteristics.PC.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction study_characteristics.PC

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.543680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.350379Z digest=sha256:71a3737b3fd2eca66d93a1ad9fc9b0ff6083dfde71d876aeff7897de798bb389

Observation 6f81d274-411b-4662-8a29-b2a4b800eb0c · outbound

This paper cites **Note:** GT and EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** GT and EXT fields may be nested

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.528298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.357558Z digest=sha256:2b43bb42f55abbb52d4dcdaefea4f94fc91fe3d075d1a5a643c28ef243e5f4b4

Observation d9c6a707-ca25-4048-ab83-9962fac63c46 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.513048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.363013Z digest=sha256:4f2427da4c1a22d9e5450cfc8fa72dd0a43cebff75f80e9db0b31910974e8cd2

Observation f85a4249-4b59-4dd7-8a96-a6d65085bd5f · outbound

This paper cites Hallucinated.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Hallucinated

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.495300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.370459Z digest=sha256:839bd4d8d184199e2c0ec4cb8bba4fdcb33b2f57f1cbaed56a9e64af17ff254f

Observation 630a5e2b-c286-48be-b87e-e5bf735924fd · outbound

This paper cites Not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Not reported

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.480859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.376749Z digest=sha256:3c4e8274ff566f5578115b9a5e5a39efcd1c83ebd74898ddfff5780d55287450

Observation 6eb66810-3256-4c2e-ac1c-081f1c2e57ba · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 2010

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.993029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.723776Z digest=sha256:637de5c843d22c69f8b402c4f8a79bbf82ebbd5ec5e13a8d6866739432d0ae97

Observation f2e4c449-39dc-4619-bcfe-c1fbc710d46c · outbound

This paper cites doi:10.2196/33124.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.2196/33124

Reference 2021

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.908021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:07.844989Z digest=sha256:4b5479111232b5aaca6d8e1882d260111d474b0de9d394f34939e4c759c8e2b1

Observation d98a3bf5-63c6-4afe-962e-99dc9b911760 · outbound

This paper cites doi:10.18653/v1/2024.findings-naacl.237.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.18653/v1/2024.findings-naacl.237

Reference 2024

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:46:08.006169Z digest=sha256:d79198985f7cc34d701eecb8af2b77006e01c092ce333429f767329c6e3a6156

Observation eb66b4e2-c9ff-40f7-bcf8-a057212be15a · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.090594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.090594Z digest=sha256:6546dd8cf4299f17bbeaffeddca36e27126fcc5d56a0ec2fc2578e88cff4e923

Pith citing papers

Observation 10d2e2ae-f6ec-48f1-a8fa-6aa090174121 · inbound

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis cites this paper.

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:54.709479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:13:53.040834Z digest=sha256:d70af84af1692675be7683431c9905a0421dbfd7a4db558385561166841796ab