OncoTraj releases a harmonized 813-patient dataset with audited splits for three tasks on osimertinib resistance, showing single-timepoint NGS features yield no model above chance while recovering a TP53 association.
hub
cc/paper_files/paper/2019/file/ ac52c626afc10d4075708ac4c778ddfc-Paper
35 Pith papers cite this work, alongside 2,791 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
A mapping of predictive distributions through the censoring mechanism yields proper right-censored versions of the CRPS, Brier score, energy score and other losses, with the marginalized form proven proper under conditional independent censoring.
A new upper bound is derived for the worst-case effect of selection bias on medical prediction model performance under partial observation of the selection process and target data.
CopFITi is the first marginalization-consistent copula for irregular multivariate time series, using normalizing flows for marginals and a Gaussian mixture copula for dependencies to reach new state-of-the-art joint density modeling.
Agent-based models of emergency departments generate synthetic EHR data to test whether machine learning models for length-of-stay prediction lose performance under mass casualty incident conditions.
Starling, a multi-agent LLM system, extracts ~6.3 million nuanced structured records from PubMed across six tasks with reported error rates of 0.6-7.7%, lower than several curated databases.
MedicalBench is a benchmark for implicit medical concept extraction and sentence-level evidence retrieval built from MIMIC-IV discharge summaries with human verification to test LLM reasoning on unstated medical ideas.
TiRex-2 is a recurrent xLSTM time series foundation model for multivariate forecasting with future covariates and constant-cost streaming that reports SOTA zero-shot results on GIFT-Eval and fev-bench.
HealthAgentBench is a new benchmark of 54 healthcare agent tasks where even the strongest frontier AI agent reaches only about 42% success rate on end-to-end clinical workflows.
A landmarking approach using latent class mixed models for dynamic prediction of time-to-event data that accounts for latent heterogeneity in longitudinal biomarker trajectories.
PORTER is a language-grounded EHR foundation model that uses text descriptions for events and a numeric pathway, matching fixed-vocabulary performance on 74 tasks while recovering 97.1% AUROC on unseen vocabularies and outperforming on MIMIC.
eCREAM-MedCorpus releases ~4M anonymized Italian ED clinical notes and a 6k-note 132-item CRF annotation set, with zero-shot Gemma/MedGemma CRF-filling baselines.
DEM distills XGBoost into a residual decision tree with a new fidelity metric for interpretable anomaly detection in WBAN data, reporting AUC 0.9964 and 0.9047 with 0.17ms inference.
LLMSurvival enables LLM-based survival analysis on tabular data by converting censored time-to-event tasks into pairwise comparisons, yielding small concordance gains over Cox and deep learning baselines on ICU mortality and fracture prediction.
Resampling clinical time series into uniform bins for offline RL reduces performance by up to 60% and causes retrospective evaluations to overestimate returns by 1.5-3x versus unprocessed data.
InvisibleInk achieves high-utility differentially private long-form LLM text generation at 4-8x the cost of non-private generation by isolating and clipping sensitive logits and sampling from a small superset of top-k private tokens without privacy cost.
WaveDetect reformulates machine-generated text detection as a time-frequency signal processing task by applying continuous wavelet transform to token probability sequences to reveal spectral fingerprints.
CHRONOS is a three-layer system for evolving data marketplaces that applies neural-ODE temporal decay, changepoint-aware Shapley valuation, and EXP3-IX private coordination to achieve 0.937 recall, 2.74 qps, 161 ms latency, and epsilon 4.25 at delta 10^-6.
A contrastive-learning ECG foundation model with multitask heads predicts post-MI outcomes better than training from scratch (AUC 0.794 vs 0.608).
Single-agent LLM frameworks outperform naive multi-agent systems in multimodal clinical risk prediction tasks and are better calibrated.
RCD balances relevance, coverage, and diversity in a knapsack-constrained selection framework, with experiments showing that selector choice and budget level determine optimal unitization strategies on clinical datasets.
Quantum Knowledge Graphs model context-dependent triplet validity and improve LLM medical reasoning accuracy by 1.4 to 6 percentage points over baselines.
A 22M-parameter hyperbolic model answers structured EHR questions with accuracy close to LLM-based systems (EHRXQA 89.5%, MIMIC-Instr 76.0%).
CXRMate-2 improves chest X-ray report generation via temporal embeddings and tractable RL, delivering metric gains and 45% acceptability in radiologist review with no significant preference difference on most findings.
citing papers explorer
-
OncoTraj: a public benchmark for longitudinal resistance prediction in EGFR-mutant non-small-cell lung cancer on osimertinib
OncoTraj releases a harmonized 813-patient dataset with audited splits for three tasks on osimertinib resistance, showing single-timepoint NGS features yield no model above chance while recovering a TP53 association.
-
Proper Scoring Rules for Right-Censored Survival Data
A mapping of predictive distributions through the censoring mechanism yields proper right-censored versions of the CRPS, Brier score, energy score and other losses, with the marginalized form proven proper under conditional independent censoring.
-
A Practical Upper Bound on Selection Bias Effects in Medical Prediction Models
A new upper bound is derived for the worst-case effect of selection bias on medical prediction model performance under partial observation of the selection process and target data.
-
Valid and Expressive Copulas for Irregular Multivariate Time Series
CopFITi is the first marginalization-consistent copula for irregular multivariate time series, using normalizing flows for marginals and a Gaussian mixture copula for dependencies to reach new state-of-the-art joint density modeling.
-
Generating synthetic electronic health record data using agent-based models to evaluate machine learning robustness under mass casualty incidents
Agent-based models of emergency departments generate synthetic EHR data to test whether machine learning models for length-of-stay prediction lose performance under mass casualty incident conditions.
-
Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Starling, a multi-agent LLM system, extracts ~6.3 million nuanced structured records from PubMed across six tasks with reported error rates of 0.6-7.7%, lower than several curated databases.
-
MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction
MedicalBench is a benchmark for implicit medical concept extraction and sentence-level evidence retrieval built from MIMIC-IV discharge summaries with human verification to test LLM reasoning on unstated medical ideas.
-
TiRex-2: Generalizing TiRex to Multivariate Data and Streaming
TiRex-2 is a recurrent xLSTM time series foundation model for multivariate forecasting with future covariates and constant-cost streaming that reports SOTA zero-shot results on GIFT-Eval and fev-bench.
-
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
HealthAgentBench is a new benchmark of 54 healthcare agent tasks where even the strongest frontier AI agent reaches only about 42% success rate on end-to-end clinical workflows.
-
Landmarking with Latent Class Mixed Models for Dynamic Prediction of Time-to-event Data with Heterogeneous Biomarker Trajectories
A landmarking approach using latent class mixed models for dynamic prediction of time-to-event data that accounts for latent heterogeneity in longitudinal biomarker trajectories.
-
PORTER: Language-Grounded Event Representations for Portable Structured EHR Foundation Models
PORTER is a language-grounded EHR foundation model that uses text descriptions for events and a numeric pathway, matching fixed-vocabulary performance on 74 tasks while recovering 97.1% AUROC on unseen vocabularies and outperforming on MIMIC.
-
eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian
eCREAM-MedCorpus releases ~4M anonymized Italian ED clinical notes and a 6k-note 132-item CRF annotation set, with zero-shot Gemma/MedGemma CRF-filling baselines.
-
DEM: A Distilled Explanation Model for Interpretable Anomaly Detection in Physiological Sensor Networks
DEM distills XGBoost into a residual decision tree with a new fidelity metric for interpretable anomaly detection in WBAN data, reporting AUC 0.9964 and 0.9047 with 0.17ms inference.
-
Towards end-to-end LLM-based censoring-aware survival analysis
LLMSurvival enables LLM-based survival analysis on tabular data by converting censored time-to-event tasks into pairwise comparisons, yielding small concordance gains over Cox and deep learning baselines on ICU mortality and fracture prediction.
-
The hidden risks of temporal resampling in clinical reinforcement learning
Resampling clinical time series into uniform bins for offline RL reduces performance by up to 60% and causes retrospective evaluations to overestimate returns by 1.5-3x versus unprocessed data.
-
InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
InvisibleInk achieves high-utility differentially private long-form LLM text generation at 4-8x the cost of non-private generation by isolating and clipping sensitive logits and sampling from a small superset of top-k private tokens without privacy cost.
-
WaveDetect: Robust Framework for Machine-Generated Text Detection via Wavelet Transform
WaveDetect reformulates machine-generated text detection as a time-frequency signal processing task by applying continuous wavelet transform to token probability sequences to reveal spectral fingerprints.
-
CHRONOS: Temporally-Aware Multi-Agent Coordination for Evolving Data Marketplaces
CHRONOS is a three-layer system for evolving data marketplaces that applies neural-ODE temporal decay, changepoint-aware Shapley valuation, and EXP3-IX private coordination to achieve 0.937 recall, 2.74 qps, 161 ms latency, and epsilon 4.25 at delta 10^-6.
-
Dynamical Predictive Modelling of Cardiovascular Disease Progression Post-Myocardial Infarction via ECG-Trained Artificial Intelligence Model
A contrastive-learning ECG foundation model with multitask heads predicts post-MI outcomes better than training from scratch (AUC 0.794 vs 0.608).
-
AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks
Single-agent LLM frameworks outperform naive multi-agent systems in multimodal clinical risk prediction tasks and are better calibrated.
-
Budget-Aware Routing for Long Clinical Text
RCD balances relevance, coverage, and diversity in a knapsack-constrained selection framework, with experiments showing that selector choice and budget level determine optimal unitization strategies on clinical datasets.
-
Quantum Knowledge Graph: Modeling Context-Dependent Triplet Validity
Quantum Knowledge Graphs model context-dependent triplet validity and improve LLM medical reasoning accuracy by 1.4 to 6 percentage points over baselines.
-
HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering
A 22M-parameter hyperbolic model answers structured EHR questions with accuracy close to LLM-based systems (EHRXQA 89.5%, MIMIC-Instr 76.0%).
-
CXRMate-2: Structured Multimodal Temporal Embeddings and Tractable Reinforcement Learning for Clinically Acceptable Chest X-ray Radiology Report Generation
CXRMate-2 improves chest X-ray report generation via temporal embeddings and tractable RL, delivering metric gains and 45% acceptability in radiologist review with no significant preference difference on most findings.
-
Handling and Interpreting Missing Modalities in Patient Clinical Trajectories via Autoregressive Sequence Modeling
Autoregressive transformer modeling with missingness-aware contrastive pre-training outperforms baselines on MIMIC-IV and eICU benchmarks and mitigates divergent behavior from removed modalities in clinical trajectories.
-
A Scientific Human-Agent Reproduction Pipeline
Autoregressive LLM decoders with a missingness-aware contrastive pre-training objective outperform static baselines on MIMIC-IV/eICU and reveal demographic-bias failure modes under modality ablation.
-
Representation Before Training: A Fixed-Budget Benchmark for Generative Medical Event Models
Fused code-value tokenization improves mortality AUROC from 0.891 to 0.915 and other clinical outcome predictions, while certain temporal encodings like event order match or exceed time tokens with shorter sequences.
-
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
CARIS is a new agentic LLM framework that automates clinical research workflows from planning to reporting in a coding-free and privacy-preserving manner, achieving high completeness scores on heterogeneous datasets.
-
Automatic Construction of Clinical Scoring Systems with LLM Agents
An LLM proposal loop with deterministic validation builds unit-weighted N-of-M clinical checklists that achieve AUROC comparable to flexible interpretable models on eight EHR tasks.
-
MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization
Proposes MedRLM, a recursive agent-based multimodal framework for long-context clinical reasoning, sensor-guided screening, and referral optimization using a Clinical Evidence Graph Memory.
-
Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs
Benchmarking in pediatric ICU antimicrobial stewardship shows performance depends mainly on target prevalence and dataset traits rather than model complexity, with sequence models improving precision-recall at 24-hour resolution but showing poorer calibration than tabular models.
-
DT-Transformer: A Foundation Model for Disease Trajectory Prediction on a Real-world Health System
DT-Transformer predicts next disease events with median age- and sex-stratified AUC 0.871 across 896 categories on held-out and prospective data from a 1.7M-patient multi-hospital EHR dataset.
-
Investigating Data Interventions for Subgroup Fairness: An ICU Case Study
Data addition from different sources does not reliably boost subgroup fairness in ICU models and often requires post-hoc calibration to work.
-
A Hybrid Retrieval and Reranking Framework for Evidence-Grounded Retrieval-Augmented Generation
A hybrid RAG system with retrieval, Cohere reranking, and claim-level LLM judgment achieves 100% grounding accuracy on 200 claims from 25 biomedical queries in a pilot study.
- NEURON: A Neuro-symbolic System for Grounded Clinical Explainability