Pith. sign in

Paper Citation Record · LEDGER

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

As of 20 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2507.15152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15152 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:46:08.376749Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:53.040834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:13:54.647415Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact16
  • verified fuzzy38
  • unresolved37
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59abdfec-2b79-400f-abb4-6896d381a4af · outbound

This paper cites Research Synthesis and Meta-Analysis: A Step-by-Step Approach.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Research Synthesis and Meta-Analysis: A Step-by-Step Approach

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.653529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.653529Z digest=sha256:71161ebf98e246ca67d9c5edcce959eb63984c32eb1329a5f366e265772ff292

Observation 1f67cc0f-6e6e-4c9c-8dd2-dd9074bee257 · outbound

This paper cites Analysing data and undertaking meta-analyses, chapter 10, pages 241–284.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysing data and undertaking meta-analyses, chapter 10, pages 241–284

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.664179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.664179Z digest=sha256:c2ad83506306ca96229f765c34bd57bdbbe4c46cea8bef0e22fd45d409b4a865

Observation d0762771-aa93-47e0-aeb8-895ffc8ebf2c · outbound

This paper cites Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the prospero registry.BMJ Open, 7 (2), 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.671805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.671805Z digest=sha256:601d7afbf6d8ab2d61c31038ddebd21375dc53bb330a186e726899ea3ff97b8d

Observation efdddffb-8142-469b-909e-5f43339c72f5 · outbound

This paper cites Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Higgins, James Thomas, Jacqueline Chandler, Miranda Cumpston, Tianjing Li, Matthew J

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.038582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.680841Z digest=sha256:982711be9b348503392824c285a8827e51f5a9fb1676bf447e1d1ca3a562f6ff

Observation 70ea14fd-5453-4d30-8447-05c6b28e9f97 · outbound

This paper cites Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Validity of data extraction in evidence synthesis practice of adverse events: repro- ducibility study

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T15:46:09.020674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.685632Z digest=sha256:c4faa7c01db6b7679887ecf67a751b89d7aea983d10b92b5c2ea7401bdf1c459

Observation 975c4024-e2d0-4308-ada1-f84d80119446 · outbound

This paper cites Toward systematic review automation: a practical guide to using machine learning tools in research synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward systematic review automation: a practical guide to using machine learning tools in research synthesis

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:07.690733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.690733Z digest=sha256:10a628cf927ee4692052043a4197ea97a6a9940b1c6c4450ebf72f8c701271e2

Observation b0629e68-bf64-4206-99b4-91252af8cde9 · outbound

This paper cites Exact: automatic extraction of clinical trial characteristics from journal publications.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exact: automatic extraction of clinical trial characteristics from journal publications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.697875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.697875Z digest=sha256:e77c553aa87b07d304f4eaf43ce8a13ac6d6086c6766da8a0d0f662920c35e9c

Observation fdd4bafa-4184-4928-a2ff-b8fe42880703 · outbound

This paper cites Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Summerscales, Shlomo Argamon, Shangda Bai, Jordan Hupert, and Alan Schwartz

Reference 8

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.968448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.746197Z digest=sha256:2e66ee4a7b053c162555e7a1a3b494fd3f7425bead039419d21aefe469fb60fc

Observation ca059850-7817-4bff-a971-e79644f87bd6 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 9

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.459526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.763071Z digest=sha256:39e1bba1c23bd4e5516437947f130532fff9e39b94f111b71b4d484d15811fda

Observation 1c86dbba-aac8-4c3c-ac6b-60cf0a5c8cc7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.784022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.784022Z digest=sha256:583d58833fa9ec9478f0c47c9ecb98ac87174c03bb27e69cad9f0a7fdd979797

Observation bc08e3c2-ce9c-41ab-ae05-793e5c754ed0 · outbound

This paper cites Automating meta-analyses of randomized clinical trials: a first look.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automating meta-analyses of randomized clinical trials: a first look

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.801622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.801622Z digest=sha256:d325daf9e299c4f18031202a644ae455e8f4485dda1dcf7b029d18b0963b0618

Observation c032b9f2-8819-4519-aed0-7709bda04041 · outbound

This paper cites Katz-Rogozhnikov, Kush R.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Katz-Rogozhnikov, Kush R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.812345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.812345Z digest=sha256:e47c68e032444537d191d0174a5527bf463998255467020002e5429df6a2ecd5

Observation ec473968-5324-44d2-a235-b93418bdb282 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.817822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.817822Z digest=sha256:9fec11871e18b4670081c1649c5b73d3ff749871ed93c0125af39f2f7e43093a

Observation 25b1e172-1af5-4a79-be85-76ea657b1997 · outbound

This paper cites Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved].

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Data extraction methods for systematic review (semi)automation: Update of a living systematic review [version 2; peer review: 3 approved]

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.928184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.824462Z digest=sha256:d3f0b1d07de7a2276854775fb075af2c7bf2d8b9c7c231bf375e2c91ca56e839

Observation 094ae696-9487-4bcc-8bfd-ea6e8255c58e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.829830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.829830Z digest=sha256:147c157cd2d332fe0c616c39f4784e6c1a4da0e35c9d060906e10a7fb68901d5

Observation d8fbea24-066b-455d-b1ca-bda23239e5b1 · outbound

This paper cites The data is in: Deciding when to automate screening in your slr, November 2023.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The data is in: Deciding when to automate screening in your slr, November 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.834832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.834832Z digest=sha256:07935c88c9ae6f0173e5ccc033187e63d5131699b16979db659351c25b03ee8a

Observation f63edf05-4f58-41af-89ae-082feb262b6e · outbound

This paper cites Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Toward automated data extraction according to tabular data structure: Cross-sectional pilot survey of the comparative clinical literature

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.839801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.839801Z digest=sha256:4e641c7ce3f50db90b238bf3d35470e920472720598f9b8f8ff94af90f4e5900

Observation 28e34ed0-8485-4fcb-8124-2dad653175dd · outbound

This paper cites MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction MetaMate: Large Language Model to the Rescue of Automated Data Extraction for Educational Systematic Reviews and Meta-analyses, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.854237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.854237Z digest=sha256:e682a60b077e1b225ec4cef9bc24891995f3f6a7e9a6e13d3864f27a5420497e

Observation e7c78217-16e8-4a6f-8fd3-64c498887a46 · outbound

This paper cites Chatgpt: Large language model (mar 14 version).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Chatgpt: Large language model (mar 14 version)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.864218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.864218Z digest=sha256:b15d28376c06c910eca71f1a2e1ad05b8ec5c509ab16655a04949a6cbdcdd4cc

Observation cac6f8e3-00ef-4460-b6f3-4ea7f1d4f000 · outbound

This paper cites Claude 2 model announcement.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Claude 2 model announcement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.870017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.870017Z digest=sha256:5e3abff6f6e18b448a5b1d38c47140a97be487b0b46c313fe70491ed4beef6e1

Observation 583f778e-b019-49a5-aa43-88fadb7779a7 · outbound

This paper cites Zero-shot infor- mation extraction for clinical meta-analysis using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zero-shot infor- mation extraction for clinical meta-analysis using large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.875116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.875116Z digest=sha256:8f348c7390b64ed700f7503ed42051a9f49202b63cf36bad27ad099e966a5984

Observation 75556778-c65e-4b3b-8f2d-8b86090c2e79 · outbound

This paper cites Performance of two large language models for data extraction in evidence synthesis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Performance of two large language models for data extraction in evidence synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.880456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.880456Z digest=sha256:bcf19e6347f7025fbb62153b45ef0cf4f1d54bb3029363306118c813eb2bdf74

Observation 6705cf30-2390-4d27-9f7e-b02c0a6a2512 · outbound

This paper cites Automatically extracting numerical results from randomized controlled trials with large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Automatically extracting numerical results from randomized controlled trials with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.885286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.885286Z digest=sha256:e568e385786adcbe7b663e1f6be9674f213caecb73090b78ea8341b519b13511

Observation fe259602-0acf-4aa5-9e96-25afc38508d9 · outbound

This paper cites Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Exploring the use of a large language model for data extraction in systematic reviews: a rapid feasibility study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.891121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.891121Z digest=sha256:4ac7948e25ad371b8285b308bdef4a61723b3b1094c5f04e035f27d337ca9603

Observation 94bbea6d-b5f3-4320-89d7-15055bdf1e24 · outbound

This paper cites Lee, Shigeki Yamada, and Tomohiro Mizuno.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Lee, Shigeki Yamada, and Tomohiro Mizuno

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.896068Z digest=sha256:22b55d1dc165b4b816ebb2f14fa327f8367bdf1e12101eb639cd4881a11ac581

Observation 1cfbaa80-85fd-4a5a-97d0-24a0a022a479 · outbound

This paper cites Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effects of the Modified DASH Diet on Adults With Elevated Blood Pressure or Hypertension: A Systematic Review and Meta-Analysis

Reference 26

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.370112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.901443Z digest=sha256:222505b884b56cb7db1a2db1ea3450d7072bfdf36ca6ed76c9a4fdb8cdde7b27

Observation 96dc95e7-b78e-47d4-bdca-e36962197bcf · outbound

This paper cites Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Abdelrahim, Nivine Hanach, Refat AlKurd, Moien Khan, Lana Mahrous, Hadia Radwan, Farah Naja, Mohamed Madkour, Khaled Obaideen, Husam Khraiwesh, and MoezAlIslam Faris

Reference 27

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.829269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.906139Z digest=sha256:53f53f6baaef64ecd942d54ba2b3e5fed058430fdbeb370f7956197deec965da

Observation 64105b66-d969-41af-a918-8da6e2c4b62b · outbound

This paper cites Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Effect of dietary glycemic index on insulin resistance in adults without diabetes mellitus: a systematic review and meta-analysis

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:46:09.285708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.913293Z digest=sha256:85c42e59957f38cd30585b2994102fd8ceeef8518566c91882f8b2555107b0bb

Observation 8b698dba-6c1e-4100-a69b-c5908f7b0a20 · outbound

This paper cites Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Choi, Min-Sun Gu, Seo-Yeong Ko, Jae-Hee Kwon, Ja-Young Han, Jae Hyun Kim, and Myeong Gyu Kim

Reference 29

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.799935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.918186Z digest=sha256:6c1a58e77a99bb9f0a541f62dbb373725305ed13476f9395a4e073bc8afab88a

Observation 5c6db0f3-9237-4145-a5d2-9b5c829d2f1e · outbound

This paper cites V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction V olar locking plate vs cast immobilization for distal radius fractures: a systematic review and meta- analysis

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.774773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.924380Z digest=sha256:ac5650201b4bcaf15fa31f9f16ee90cfeae0625263735a22403bd73dc1639657

Observation 34192b22-a459-443b-8287-0cf5870d8db2 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gpt-4o mini: Advancing cost-efficient intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.934074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.934074Z digest=sha256:dff7ff0217ca22eb86dfc35c5f1cc2ebf14d6638a7690a5ff0e4ea817f868b5d

Observation c32e8170-e0d4-4921-9966-1a46b81557ab · outbound

This paper cites Gemini 2.0 flash.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Gemini 2.0 flash

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.939758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.939758Z digest=sha256:865a53cf0a8dd310309fc67add0e9b514ce20cb18b78e15bbe2fb9b6fc6389d6

Observation 93682bef-d44e-4c7c-8a57-315a4becb015 · outbound

This paper cites Grok-3 language model.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Grok-3 language model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.945503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.945503Z digest=sha256:e5d2950d7079a43759e184dc6fd925a1f133ada855b8694ee76737b8a5410aaf

Observation 8b9cbaff-4f41-4132-857c-c60de2b330d0 · outbound

This paper cites The impact of temperature on extracting information from clinical trial publications using large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction The impact of temperature on extracting information from clinical trial publications using large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.951933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.951933Z digest=sha256:90483c8823fc46f994fe958bd41810e367c9c67d12faa1ac44118eb07e09f05c

Observation 4bbf6358-214b-43d1-afdb-4f0aea3b2d51 · outbound

This paper cites AI-Assisted Data Extraction for Systematic Reviews in Education.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction AI-Assisted Data Extraction for Systematic Reviews in Education

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.958889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.958889Z digest=sha256:8fa9398fd385951cb4ab23118777b1667bc067e4ae426ad8b066e1ebbe18808b

Observation 178abf96-a100-40cc-b4ea-fefa51c92b6a · outbound

This paper cites Use gemini 2.0 to speed up data processing.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Use gemini 2.0 to speed up data processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.965054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.965054Z digest=sha256:ae4a00b5c02ce7f75b90529d476d13922f493c197155848e2fb5b996fa2e8114

Observation f654e4b5-931a-4750-a1f4-2c9e8ea67a82 · outbound

This paper cites Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Harnessing ai for integrative medicine: Exploring grok 3’s role in researching qigong, tai chi, yoga, and mindfulness for college students’ mental health

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.748746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.971513Z digest=sha256:712ccd7c7f5c394a57b9e63f9efe604fa36cb18ffe33f63c199997bae1d83c31

Observation 759dce94-632b-4544-a756-d882cdaf18b0 · outbound

This paper cites Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Microsoft adds elon musk’s grok-3 to azure, citing health- care and science use cases

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.347631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.976676Z digest=sha256:98c9b2d6859059639a9963afec25a73e90058f609791c897e90013792de390a6

Observation fc9056f3-42b0-4a30-89cd-34b59992be83 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Reflexion: language agents with verbal reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.318292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.989820Z digest=sha256:d0a81447639f54a4d18afe75ab2ef8d31b042ae68ec625655decff4a81770573

Observation 2acec01b-d982-4110-aa42-d3a6d238ea8e · outbound

This paper cites Towards mitigating LLM halluci- nation via self reflection.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Towards mitigating LLM halluci- nation via self reflection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:07.994474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:07.994474Z digest=sha256:123924b7f84e6ce3c332d3628b8ad949d72f3e4974490888ec1c61e40f88add4

Observation 6c7a93d2-95b1-41d6-8356-4ae8b04c0c95 · outbound

This paper cites When hindsight is not 20/20: Testing limits on reflective thinking in large language models.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction When hindsight is not 20/20: Testing limits on reflective thinking in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.295126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.000826Z digest=sha256:927b0e8fda58d1f06a5155bbf67983a789c433c6dea9b95f278cc48fdf4223e1

Observation ccc81064-a620-4159-a80e-ba3512ca49d4 · outbound

This paper cites Dietterich.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Dietterich

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.267940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.011272Z digest=sha256:be5952d94e028f85e63b653657fcbac5901aaef7684f6b18c2dd45713b8e69ac

Observation 23f4b757-ac28-47ef-8306-fc3ec919bcb0 · outbound

This paper cites A survey on ensemble learning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction A survey on ensemble learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.018396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.018396Z digest=sha256:fda5560b07f6cd6e5dc4a1e0fd898483a8e14627e94c752e537e80510ce8c56a

Observation e84df5e4-15d8-4caa-aca3-aba9d2a57a96 · outbound

This paper cites Ensemble pretrained language models to extract biomedical knowledge from litera- ture.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Ensemble pretrained language models to extract biomedical knowledge from litera- ture

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.645522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.023179Z digest=sha256:4418c3c83a6b4b1eacf1af0e3b68ca8509ba210dcad26e6f56546693b69aa667

Observation 8d7ef2bc-7931-463f-922d-5202489506dd · outbound

This paper cites Zhang and A.L.P.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Zhang and A.L.P

Reference 45

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.618170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.030348Z digest=sha256:5ebb76c71b7f3100277912e807023ec5b0aa9e18e6a00a0db2c792bbdbfb8840

Observation 31adf290-91b9-4d02-862c-b19873f71cee · outbound

This paper cites Comprehensive testing of large language models for extraction of structured data in pathology.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Comprehensive testing of large language models for extraction of structured data in pathology

Reference 46

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.587520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.036361Z digest=sha256:cf29f31f827017cd2dcaabdf733228188ad5ec584e9448f5dd183cd226a8e3a3

Observation d28dc553-4477-4edf-b7d3-e393a2f59d3f · outbound

This paper cites Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Innocence discovery lab - harnessing large language models to surface data buried in wrongful conviction case documents

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.546789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.041977Z digest=sha256:e592b90a89ac0af5f88661739eb28316dbd401639350bd5b642928373dd0b786

Observation 18e1f353-74a7-4807-89c8-99ee1df24e4f · outbound

This paper cites Match, compare, or select? an investigation of large language models for entity matching.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Match, compare, or select? an investigation of large language models for entity matching

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.243358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.048626Z digest=sha256:0b7ba47908c5b7086216edd0c5eb22bd59371f9ed3c9b48c4689afe39af7833e

Observation cc3c1a78-35ff-4e69-a324-9c56813ca94d · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 49

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.516502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.053892Z digest=sha256:1966f394db2f8e341bf115458f0a895369a98273102e29d6c3b3025282723f25

Observation 84af77db-91b6-4c1a-b14e-2095284f718e · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.061584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.061584Z digest=sha256:90a2bd9fc4bcaaeae636aebbeb5940a37e1eace5860a66656660b1ebd73a4bdf

Observation 3f397eca-7685-4cf9-8b04-69ba6bc821d3 · outbound

This paper cites Guyatt, Andrew D.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Guyatt, Andrew D

Reference 51

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.492238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.068084Z digest=sha256:c1557df52818f75ccd33a63dbf04792696aa167443c600447afc44c57a3a0ecf

Observation 8614a940-0e89-4af5-add4-f99fb77eb40c · outbound

This paper cites Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.075310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.075310Z digest=sha256:cd447261c68e3f916ce9850f37583023546babf8ece4dc346fcf61e999a04fb6

Observation 449df189-eab1-4838-8f8f-8e5dc1abdc35 · outbound

This paper cites Agentic reasoning: Reasoning llms with tools for the deep research,.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic reasoning: Reasoning llms with tools for the deep research,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.198891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.083557Z digest=sha256:44189993a6322b1092a4effca2f08bdcc48956d9ef90faacba278129ab3342aa

Observation 2aa669b4-d09e-46ea-aa25-0ac7f3da646a · outbound

This paper cites TART: An open-source tool- augmented framework for explainable table-based reasoning.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction TART: An open-source tool- augmented framework for explainable table-based reasoning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.171211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.099934Z digest=sha256:281e150029615337a85121afaa49d17c26f5b61be7e188650395408a1aa84524

Observation b81958ca-a660-4689-bb8e-db65fdbe01a3 · outbound

This paper cites Medical hallucination in foundation models and their impact on healthcare.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Medical hallucination in foundation models and their impact on healthcare

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.105709Z digest=sha256:5f8de8125bd3390a1b8fa375625e1cbff650fc7ba7dd29fa61e4ad6faf02179b

Observation cabcab11-7bd8-46b3-ae71-b44625490a07 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.111858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.111858Z digest=sha256:5851fa4d525cd0cfc9defbfcef044a2d9f42817a29a214f55f03f9792d24df61

Observation 2b1898d8-f3f5-489c-9aee-2d67e449992e · outbound

This paper cites Potential roles of large language models in the production of systematic reviews and meta-analyses.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Potential roles of large language models in the production of systematic reviews and meta-analyses

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:46:08.118777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.118777Z digest=sha256:4f947379ad4819e7cda0ad81881d43d874579802314e0936c34406f1d7f2cd03

Observation 93db9ab6-c6b9-4861-98f0-f20db0315ff6 · outbound

This paper cites justification.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction justification

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.153899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.124679Z digest=sha256:4305294d96d527f3c5b9631a1a380fdb21fb53a251cede2b9d19d1f64dddcac4

Observation e1d54069-2ad1-44a4-9a0c-74c7a7d12d35 · outbound

This paper cites other_time_points.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction other_time_points

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.136844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.130975Z digest=sha256:74fa7df47ee5b128fb66b333ab9e140ee42edf93d016332c9688c48dbf31cd5f

Observation 69682989-e509-4562-9ede-0699148ce6d3 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.117550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.136861Z digest=sha256:2b908fcc0a7d2d791c30f59eb261de5d9f3507e44d3997df06eb9907e62e87c9

Observation 5af7e017-6330-4c4b-9a15-b03cafc5d496 · outbound

This paper cites null"`, NOT `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"`, NOT `

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.100782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.142821Z digest=sha256:59bfff53610150f85561117e7fe4c8b9c1b3fb535bde5a1ae9002c619c9454c9

Observation 502eff7e-a05c-4706-a710-8504908387b9 · outbound

This paper cites more common in the intervention group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction more common in the intervention group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.078892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.149296Z digest=sha256:bf5f9b05bd067f954ce3d9e623eeb7ed3cc5748a38796d538d3218b63e84d834

Observation 03ae9405-e1de-40c5-a58f-b5f4d3b0af26 · outbound

This paper cites data_conflicts.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction data_conflicts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.058081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.156022Z digest=sha256:6b7865b5d058f9c247218e2236da976a007877107ce0ac744aba3fa0d660f7b0

Observation cbdb9462-4a71-409a-9055-a8c6b0f3411e · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:46:10.036874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.167425Z digest=sha256:407aa60252287c6d8e0fc4f472b76acdde22235af13136df5ce51ec1211a35b9

Observation be175a7f-d594-4d4a-b249-ba68e6a3e99c · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.018641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.175488Z digest=sha256:debea8315c7e01176ad3ced4a25232e956c7b3958cf26758703d2510eda7984c

Observation b871bc30-9e35-4490-80b3-93fad8c85ad0 · outbound

This paper cites - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify any *structural inconsistencies* (e.g., missing key study characteristics, incomplete sample size reporting)

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:10.000536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.180143Z digest=sha256:7a7978fd7f6256062bec6d93a6d2ba9e2260a5561210235a4ab734ed3878e9de

Observation 011d302c-57c1-4b6c-9e08-c0b603624a95 · outbound

This paper cites - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg).

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - *Unit consistency*: Verify all measurements use the correct units (e.g., blood pressure should be in mmHg)

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.981621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.185009Z digest=sha256:4829980cb5c4e0f0aedf9a03d3c26ad657c815f64dde71c15fa805f7c4510f7b

Observation 7e25a773-219d-4351-b77e-f33ec5ef81cf · outbound

This paper cites - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction - Identify discrepancies and data conflicts between different sections of the paper (e.g., abstract vs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.956642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.190287Z digest=sha256:b1cea6a00f3c390f6ae1b105ac1d3bf4f757db39696f9bb1a9ef76b9ea43d356

Observation eeed05e4-b1e0-4c03-8474-33f38239b288 · outbound

This paper cites source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction source" the most appropriate location in the paper for this data? If not, provide a more accurate source. - Confidence Justification: Is the assigned

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.929667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.194780Z digest=sha256:038dda34439f10d3e9fe3eedde5dcf07c98a2014a07afb6fed2641e4a07177d5

Observation b82d1c6d-022f-4f65-b6be-1fa851dfde95 · outbound

This paper cites needs_transformation.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction needs_transformation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.910425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.200814Z digest=sha256:7621d69f0671638a9d73ff908557ee8a00c36e785e74e2af77e68442d21b872e

Observation 3c9eca83-c241-4135-8854-e64f65e33fec · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.890938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.207241Z digest=sha256:8abf97aef22554b86235b9561fa91491e129871fb82e04d9c03864037a2c1e79

Observation 8fdb2b6c-3a11-49f8-9877-ccc8cae1ee5a · outbound

This paper cites null"` or `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null"` or `

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.871709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.212807Z digest=sha256:38d6a0d3aa78a9316b42e0e44d8351a9dccb1c99dab060027e61b56399fa90c1

Observation 92559793-e3f6-4e01-bbf9-3f5fdee6ad67 · outbound

This paper cites revised_value.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction revised_value

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.855205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.217778Z digest=sha256:b81770f43fbfe82a0256537bbbc26ab4cfafd33110641d4e0f365847dada3eab

Observation 9fc1478b-13be-463e-83ff-8a368299de13 · outbound

This paper cites pdf_status.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction pdf_status

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.835083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.222572Z digest=sha256:3d1402ffd20cdb41ea7f0bc14a9f1bdb5c22365cd5bc3156b93e7af75251027e

Observation 90943ae5-6f81-415a-8a34-8356cbc00d3f · outbound

This paper cites confidence.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction confidence

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.813283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.234818Z digest=sha256:f7c432d25980208a2cfaae7e40fb11605414d2a74cc28736e97c2dae61ee23d3

Observation aa8d47f0-2708-41cd-a7a1-9c15b0df1e89 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.793979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.240837Z digest=sha256:c8d5099e96df6532cb43f801546ede61faa8488f6038610cec5ad7c9e2a37aa0

Observation d2d90d35-360b-48b9-b65e-437e63df213b · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.770068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.246699Z digest=sha256:6ea9c5074069ccd5f6e54632f02ad48d2b44114a87ab1e2849642127d9f2a063

Observation 41fcd1e8-5048-4941-909e-7a53cd38dea4 · outbound

This paper cites Just return the final merged JSON object.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Just return the final merged JSON object

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.753155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.252902Z digest=sha256:4cd0d639ac9d0e689854da270ad5f50db508dbb326cd5cbc0973b38412c245ff

Observation 9385cae5-652e-45cb-8c22-60b0402bf079 · outbound

This paper cites LGL_group.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction LGL_group

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.713347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.270446Z digest=sha256:5a4db8dfaf8f8be7ac3f980a00547836216bdee808687c8bc8278d8bf8809b5c

Observation bd83a4e9-a0df-4780-ada0-9ec931fa8fcd · outbound

This paper cites **Note:** EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** EXT fields may be nested

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.694694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.277351Z digest=sha256:6fe022c40ab3677df3a31ea83dce26188dfeb6aa3928b900c8ed7f1c94c7ec25

Observation 6c0500bc-867b-4458-806d-71ba0e2a6d13 · outbound

This paper cites kg/m²"` and `.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction kg/m²"` and `

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.679005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.284067Z digest=sha256:a0b42fee9d98d7a36a53828f779d2e21d9d658e72dfae98e2561c007612dbf5c

Observation 7176815e-0209-41f6-8c8e-5775018994e1 · outbound

This paper cites low glycemic load diet.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction low glycemic load diet

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.664051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.291370Z digest=sha256:e06e59293b693ad9d18c0378e275bd35d85b99b3ef8de18db6fcadfc42fdd146

Observation 17f12667-e5e3-4b97-af9c-4857f228a268 · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.649218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.296906Z digest=sha256:ef8362a1c5f17eabe1c9a19a33b950845582eb1d74d0b2d29ef76870c0c22291

Observation 1be37d85-a7de-4cf8-8428-1f2016d914b7 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.635116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.305674Z digest=sha256:266bbc84845d9cc04e78ade48499b5c2186e478ef741882f6b4b98414e8811e1

Observation 2edfcde8-35c9-4f4d-9765-3301503073e9 · outbound

This paper cites randomised controlled trial.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction randomised controlled trial

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.620028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.316558Z digest=sha256:565edcdb27c1f495993ed28a4f1a927ca3ed2b88f5f6cb64b5188d8663a90cad

Observation fc808152-11b7-4421-add0-9536fa7c8b8c · outbound

This paper cites You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction You may refer to EXT field meta-information (e.g., `source`, `notes`, `confidence`) to aid in field matching, especially when EXT uses vague or ambiguous labels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.605521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.322330Z digest=sha256:9fe10017fcc5eee3a3c32ca8e439943af9e356afd5ec99f9d283690bbabc3e07

Observation cdf87145-1676-4d00-b375-3100fed7d3fa · outbound

This paper cites not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction not reported

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.589766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.328561Z digest=sha256:7bbf237054ea348d644f94f2e36bc5f9ceb570591460f8d7f3e205984832e4b6

Observation 685b9bca-5db9-43cd-9dad-2c4663eac3aa · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.575016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.335133Z digest=sha256:8e8521412d1571c748a288c26d638fe1b63bbc2217b7bdaae1363ec8e71a1d72

Observation 8c836ecb-5ab3-4107-a7c8-4d359f7eae8a · outbound

This paper cites null" or.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction null" or

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.558785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.340397Z digest=sha256:374a006230123716e0dd9359164bd140d1128c76e8f72c4ed58c4dcde96bba40

Observation 606b3b47-a88b-4036-b68c-99e978ad0083 · outbound

This paper cites This is the preferred method.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction This is the preferred method

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.731507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.345245Z digest=sha256:34dd6d81f2e2e07731c2f718e642cb4e21f3ca6f5f31257eb09d7bf0b4dc1117

Observation eb4f5bfe-8964-4cb9-82f6-44340197761e · outbound

This paper cites study_characteristics.PC.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction study_characteristics.PC

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.543680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.350379Z digest=sha256:eeb4d043adbcbb80109dd6e4ae3ca1959740d040265cf86526791d9d3cf78a0f

Observation 6f81d274-411b-4662-8a29-b2a4b800eb0c · outbound

This paper cites **Note:** GT and EXT fields may be nested.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction **Note:** GT and EXT fields may be nested

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.528298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.357558Z digest=sha256:164161b6b172e7f92d612446197b893d5bb9906d92c5c1d55334c530a907ba07

Observation d9c6a707-ca25-4048-ab83-9962fac63c46 · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:46:09.513048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.363013Z digest=sha256:0cec1aa58c66c588124045e6c518e8ce753f1a466bb32623680cdee9c5b7cf7e

Observation f85a4249-4b59-4dd7-8a96-a6d65085bd5f · outbound

This paper cites Hallucinated.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Hallucinated

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.495300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.370459Z digest=sha256:ae811dd979e56f0a72a607adf797b54db23a18dcd7986703bbc0cc2eb4fb36c7

Observation 630a5e2b-c286-48be-b87e-e5bf735924fd · outbound

This paper cites Not reported.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Not reported

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:46:09.480859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.376749Z digest=sha256:55c6676ba781389e31aee449d365c2c092374c0839061570d65c3dbdfeb71d42

Observation 6eb66810-3256-4c2e-ac1c-081f1c2e57ba · outbound

This paper cites an unresolved cited work.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Unresolved cited work

Reference 2010

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.993029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.723776Z digest=sha256:534d06b8ca5ba2e63d2cdbc783fa376155339a4a90313297d99df434b02813b1

Observation f2e4c449-39dc-4619-bcfe-c1fbc710d46c · outbound

This paper cites doi:10.2196/33124.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.2196/33124

Reference 2021

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:46:08.908021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:07.844989Z digest=sha256:d3d3284d301a772f15ab99e15d7c7d8261f410e36cb75efb56a960fd78fe30cb

Observation d98a3bf5-63c6-4afe-962e-99dc9b911760 · outbound

This paper cites doi:10.18653/v1/2024.findings-naacl.237.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction doi:10.18653/v1/2024.findings-naacl.237

Reference 2024

Resolution
verified exact
doi, observed 2026-08-06T15:46:08.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:46:08.006169Z digest=sha256:263e2e69a51a9df092f24e0dff0472aa3714e9389c9da10b8988a29ab6ede758

Observation eb66b4e2-c9ff-40f7-bcf8-a057212be15a · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:08.090594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:08.090594Z digest=sha256:d58a32a4eb2b14687611bedfa4fdaf9f07eb274720dab802923255965f051819

Pith citing papers

Observation 10d2e2ae-f6ec-48f1-a8fa-6aa090174121 · inbound

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis cites this paper.

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:54.709479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T17:13:53.040834Z digest=sha256:9f71a4797cd0162b5fbfc2b73ddf409a7eded12ce894edc9d01d92313288c0a0