Pith. sign in

REVIEW 2 cited by

LangCell: Language-Cell Pre-training for Cell Identity Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06708 v5 pith:BOGMO6KR submitted 2024-05-09 q-bio.GN cs.AIcs.CL

classification q-bio.GNcs.AIcs.CL
keywords cellidentityunderstandingdatalangcellsingle-cellinformationmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Cell identity encompasses various semantic aspects of a cell, including cell type, pathway information, disease information, and more, which are essential for biologists to gain insights into its biological characteristics. Understanding cell identity from the transcriptomic data, such as annotating cell types, has become an important task in bioinformatics. As these semantic aspects are determined by human experts, it is impossible for AI models to effectively carry out cell identity understanding tasks without the supervision signals provided by single-cell and label pairs. The single-cell pre-trained language models (PLMs) currently used for this task are trained only on a single modality, transcriptomics data, lack an understanding of cell identity knowledge. As a result, they have to be fine-tuned for downstream tasks and struggle when lacking labeled data with the desired semantic labels. To address this issue, we propose an innovative solution by constructing a unified representation of single-cell data and natural language during the pre-training phase, allowing the model to directly incorporate insights related to cell identity. More specifically, we introduce $\textbf{LangCell}$, the first $\textbf{Lang}$uage-$\textbf{Cell}$ pre-training framework. LangCell utilizes texts enriched with cell identity information to gain a profound comprehension of cross-modal knowledge. Results from experiments conducted on different benchmarks show that LangCell is the only single-cell PLM that can work effectively in zero-shot cell identity understanding scenarios, and also significantly outperforms existing models in few-shot and fine-tuning cell identity understanding scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SToFM: a Multi-scale Foundation Model for Spatial Transcriptomics

    q-bio.GN 2025-07 conditional novelty 6.0 of 10

    SToFM pretrains an SE(2) Transformer on multi-scale sub-slices of spatial transcriptomics data, using virtual cells to capture tissue structure, and reports state-of-the-art results on several ST benchmarks.

  2. Spatial Coordinates as a Cell Language: A Multi-Sentence Framework for Imaging Mass Cytometry Analysis

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Spatial2Sentence uses multi-sentence prompts built from expression-similar and spatially close cells to improve cell-type and clinical status classification in imaging mass cytometry.

Pith tools