Pith. sign in

REVIEW 11 cited by

Clinical ModernBERT: An efficient and long context encoder for biomedical text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.03964 v1 pith:PCM5EQGZ submitted 2025-04-04 cs.CL cs.AIcs.LG

Clinical ModernBERT: An efficient and long context encoder for biomedical text

classification cs.CL cs.AIcs.LG
keywords clinicalmodernbertbiomedicalcontextencoderlongmedicalpretrained
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce Clinical ModernBERT, a transformer based encoder pretrained on large scale biomedical literature, clinical notes, and medical ontologies, incorporating PubMed abstracts, MIMIC IV clinical data, and medical codes with their textual descriptions. Building on ModernBERT the current state of the art natural language text encoder featuring architectural upgrades such as rotary positional embeddings (RoPE), Flash Attention, and extended context length up to 8,192 tokens our model adapts these innovations specifically for biomedical and clinical domains. Clinical ModernBERT excels at producing semantically rich representations tailored for long context tasks. We validate this both by analyzing its pretrained weights and through empirical evaluation on a comprehensive suite of clinical NLP benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Causal Language Modeling Detour Improves Encoder Continued Pretraining

    cs.CL 2026-05 conditional novelty 7.0

    A temporary CLM phase followed by MLM decay during encoder continued pretraining outperforms standard MLM on biomedical tasks by 0.3-2.8pp across languages and model sizes.

  2. Expanders Meet Reed-Muller: Easy Instances of Noisy k-XOR

    cs.CC 2026-04 unverdicted novelty 7.0

    Explicit near-optimal expanders exist for which noisy k-XOR is polynomial-time solvable, falsifying conjectures that expansion implies hardness.

  3. M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

    cs.LG 2026-07 conditional novelty 6.0

    A GAN conditioned on contrastive-aligned histopathology images and clinical text generates gene expression profiles, with attention maps showing which tissue regions drove the synthesis.

  4. Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

    cs.CL 2026-06 unverdicted novelty 6.0

    A web curation recipe using medical-term density filtering and LLM signal-amplifying rephrasing produces FineMed, a French medical pretraining corpus, and DoctoBERT encoders that outperform prior approaches on medical...

  5. Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

    cs.AI 2026-06 unverdicted novelty 6.0

    Traj-Evolve combines non-parametric experience retrieval and multi-agent RL with a leave-one-out unification strategy to outperform baselines on lung cancer prediction from up to five years of multimodal EHRs, includi...

  6. AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0

    AURORA is a representation learning framework that uses contextual orthogonalization and relational alignment to create disentangled, geometrically interpretable latent spaces in healthcare foundation models.

  7. Event Fields: Learning Latent Event Structure for Waveform Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0

    Event-centric waveform foundation models are learned via self-supervised consistency on latent event structures and interactions, yielding improved performance and label efficiency over sequence-based baselines on phy...

  8. Uncertainty-Aware Foundation Models for Clinical Data

    cs.LG 2026-04 unverdicted novelty 6.0

    The work introduces uncertainty-aware foundation models for clinical data by learning set-valued patient representations that enforce consistency across partial observations and integrate multimodal self-supervised ob...

  9. MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

    eess.SP 2026-07 conditional novelty 5.0

    Morphology-aware masking plus cross-modal ECG–SpO2 pretraining on MIMIC yields stronger transfer than MAE, contrastive, Barlow Twins, and JEPA on several clinical prediction tasks.

  10. WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records

    cs.LG 2026-05 unverdicted novelty 5.0

    WISTERIA learns robust clinical representations from noisy EHR labels by enforcing consistency across multiple weak supervision views plus ontology regularization.

  11. PRIMA: Pre-training with Risk-integrated Image-Metadata Alignment for Medical Diagnosis via LLM

    cs.CV 2026-02 conditional novelty 5.0

    PRIMA aligns images and patient metadata using four contrastive-style losses plus an LLM-generated clinical knowledge corpus, reporting higher average F1 than SOTA baselines on PAD-UFES-20 and the private AQUA dataset.