Starling, a multi-agent LLM system, extracts ~6.3 million nuanced structured records from PubMed across six tasks with reported error rates of 0.6-7.7%, lower than several curated databases.
Large language models are few-shot clinical information extractors
9 Pith papers cite this work, alongside 297 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 4polarities
background 4representative citing papers
Hybrid neural-symbolic pipeline extracts (action, date) pairs from clinical notes at 0.99 Pair F1 by using BioBERT tagging plus deterministic time normalization, outperforming LLMs on a synthetic benchmark with OOV actions.
EvidenceNet releases disease-specific biomedical knowledge bases with 7,872 and 6,622 evidence records for HCC and CRC, plus graphs, extracted via LLM pipeline with reported high fidelity.
ACIE deploys an agentic RAG system for clinical information extraction that reaches 96.5% acceptance by nuclear-medicine physicians verifying extractions against source passages in a lymphoma registry study.
RECOVER is an LLM-powered RPM system for postoperative GI cancer care, built from 7 participatory design sessions and 5 patient interviews, then piloted with 4 staff and 5 patients to derive design strategies and responsible AI insights.
Using GPT-5.4 to clean labels in the CT-RATE chest CT dataset revealed 3.6% discordance with original labels, with radiologists supporting the LLM labels in 74-92% of reviewed cases.
A modular RAG pipeline with schema-constrained prompting, deterministic post-processing, and second-pass auditing reaches 80.36% F1 on observation extraction from nurse-patient transcripts using GPT-5.2.
SchemaRAG dynamically reduces large schemas via RAG for LLM information extraction, reporting up to 8.8% micro-F1 gain, 47% latency cut, and 48% token cost reduction on healthcare and e-commerce data.
wSSAS is a two-phase deterministic framework that uses hierarchical text organization and SNR-based feature prioritization to improve clustering integrity, categorization accuracy, and reproducibility when applying LLMs to large review datasets.
citing papers explorer
-
Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Starling, a multi-agent LLM system, extracts ~6.3 million nuanced structured records from PubMed across six tasks with reported error rates of 0.6-7.7%, lower than several curated databases.
-
Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline
Hybrid neural-symbolic pipeline extracts (action, date) pairs from clinical notes at 0.99 Pair F1 by using BioBERT tagging plus deterministic time normalization, outperforming LLMs on a synthetic benchmark with OOV actions.
-
Building evidence-based knowledge bases from full-text literature for disease-specific biomedical reasoning
EvidenceNet releases disease-specific biomedical knowledge bases with 7,872 and 6,622 evidence records for HCC and CRC, plus graphs, extracted via LLM pipeline with reported high fidelity.
-
Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why
ACIE deploys an agentic RAG system for clinical information extraction that reaches 96.5% acceptance by nuclear-medicine physicians verifying extractions against source passages in a lymphoma registry study.
-
RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care
RECOVER is an LLM-powered RPM system for postoperative GI cancer care, built from 7 participatory design sessions and 5 patient interviews, then piloted with 4 staff and 5 patients to derive design strategies and responsible AI insights.
-
Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset
Using GPT-5.4 to clean labels in the CT-RATE chest CT dataset revealed 3.6% discordance with original labels, with radiologists supporting the LLM labels in 74-92% of reviewed cases.
-
Retrieval-Augmented Large Language Models for Schema-Constrained Clinical Information Extraction
A modular RAG pipeline with schema-constrained prompting, deterministic post-processing, and second-pass auditing reaches 80.36% F1 on observation extraction from nurse-patient transcripts using GPT-5.2.
-
SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction
SchemaRAG dynamically reduces large schemas via RAG for LLM information extraction, reporting up to 8.8% micro-F1 gain, 47% latency cut, and 48% token cost reduction on healthcare and e-commerce data.
-
Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs
wSSAS is a two-phase deterministic framework that uses hierarchical text organization and SNR-based feature prioritization to improve clustering integrity, categorization accuracy, and reproducibility when applying LLMs to large review datasets.