REVIEW 5 cited by
Advancing bioinformatics with large language models: components, applications and perspectives
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are a class of artificial intelligence models based on deep learning, which have great performance in various tasks, especially in natural language processing (NLP). Large language models typically consist of artificial neural networks with numerous parameters, trained on large amounts of unlabeled input using self-supervised or semi-supervised learning. However, their potential for solving bioinformatics problems may even exceed their proficiency in modeling human language. In this review, we will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics, spanning genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, the core attention mechanism, and the pre-training processes underlying these models. Additionally, we will introduce currently available foundation models and highlight their downstream applications across various bioinformatics domains. Finally, drawing from our experience, we will offer practical guidance for both LLM users and developers, emphasizing strategies to optimize their use and foster further innovation in the field.
Forward citations
Cited by 5 Pith papers
-
Multi-granular Training Strategies for Robust Multi-hop Reasoning Over Noisy and Heterogeneous Knowledge Sources
AMKOR is described as a state-of-the-art multi-hop QA system, but the paper provides no reproducible evidence and the reported numbers appear unverifiable.
-
Generalization of Medical Large Language Models through Cross-Domain Weak Supervision
A claimed curriculum-based fine-tuning framework for medical LLMs reports better question answering and response generation, but lacks reproducible evidence.
-
Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models
A vague proposal for LLM-guided diffusion with dynamic KL weighting, backed by unsupported FID/IS tables.
-
Instruction Tuning for Story Understanding and Generation with Weak Supervision
The paper claims a weak-to-strong instruction tuning curriculum improves story generation, but the method is standard sequential fine-tuning and the reported results are not reproducible.
-
Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models
The authors claim that LLM prompt refinement plus a CLIP-based weak supervision filter improves diffusion-based fashion image generation, but the evidence is unverifiable and internally inconsistent.
Discussion (0). Continue with ORCID to comment.