Pith. sign in

REVIEW 5 cited by

Advancing bioinformatics with large language models: components, applications and perspectives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04155 v2 pith:VT6HXXH5 submitted 2024-01-08 q-bio.QM cs.CL

classification q-bio.QMcs.CL
keywords modelslanguagelargebioinformaticswillapplicationsartificialcomponents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are a class of artificial intelligence models based on deep learning, which have great performance in various tasks, especially in natural language processing (NLP). Large language models typically consist of artificial neural networks with numerous parameters, trained on large amounts of unlabeled input using self-supervised or semi-supervised learning. However, their potential for solving bioinformatics problems may even exceed their proficiency in modeling human language. In this review, we will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics, spanning genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, the core attention mechanism, and the pre-training processes underlying these models. Additionally, we will introduce currently available foundation models and highlight their downstream applications across various bioinformatics domains. Finally, drawing from our experience, we will offer practical guidance for both LLM users and developers, emphasizing strategies to optimize their use and foster further innovation in the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 27 citations worldwide. Full citation record

  1. Multi-granular Training Strategies for Robust Multi-hop Reasoning Over Noisy and Heterogeneous Knowledge Sources

    cs.CL 2025-02 reject novelty 2.0 of 10

    AMKOR is described as a state-of-the-art multi-hop QA system, but the paper provides no reproducible evidence and the reported numbers appear unverifiable.

  2. Generalization of Medical Large Language Models through Cross-Domain Weak Supervision

    cs.CL 2025-02 reject novelty 2.0 of 10

    A claimed curriculum-based fine-tuning framework for medical LLMs reports better question answering and response generation, but lacks reproducible evidence.

  3. Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models

    cs.CL 2025-02 reject novelty 2.0 of 10

    A vague proposal for LLM-guided diffusion with dynamic KL weighting, backed by unsupported FID/IS tables.

  4. Instruction Tuning for Story Understanding and Generation with Weak Supervision

    cs.CL 2025-01 reject novelty 2.0 of 10

    The paper claims a weak-to-strong instruction tuning curriculum improves story generation, but the method is standard sequential fine-tuning and the reported results are not reproducible.

  5. Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models

    cs.CL 2025-01 reject novelty 2.0 of 10

    The authors claim that LLM prompt refinement plus a CLIP-based weak supervision filter improves diffusion-based fashion image generation, but the evidence is unverifiable and internally inconsistent.

Pith tools