Pith. sign in

REVIEW 2 cited by

Time Matters: Examine Temporal Effects on Biomedical Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.17638 v2 pith:TUGEO2NJ submitted 2024-07-24 cs.CL

classification cs.CL
keywords biomedicalmodelslanguagedataeffectstemporalperformancetasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Time roots in applying language models for biomedical applications: models are trained on historical data and will be deployed for new or future data, which may vary from training data. While increasing biomedical tasks have employed state-of-the-art language models, there are very few studies have examined temporal effects on biomedical models when data usually shifts across development and deployment. This study fills the gap by statistically probing relations between language model performance and data shifts across three biomedical tasks. We deploy diverse metrics to evaluate model performance, distance methods to measure data drifts, and statistical methods to quantify temporal effects on biomedical language models. Our study shows that time matters for deploying biomedical language models, while the degree of performance degradation varies by biomedical tasks and statistical quantification approaches. We believe this study can establish a solid benchmark to evaluate and assess temporal effects on deploying biomedical language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Large language models predict retinopathy of prematurity risk poorly from admission notes alone, over-predict medium and high risk, and positive emotional prompt framing partially corrects this bias.

  2. Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A temporal mixture-of-experts model with cluster-based shift signals improves cross-time multilingual document classification and reveals language-specific temporal performance drops.

Pith tools