REVIEW 3 cited by
xVal: A Continuous Numerical Tokenization for Scientific Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency.
Forward citations
Cited by 3 Pith papers
-
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
multivariateGPT extends next-token prediction to jointly predict the class and continuous value of mixed categorical and numeric time series, with Gaussian uncertainty, and outperforms discrete-token baselines on clin...
-
POI-Enhancer: An LLM-based Semantic Enhancement Framework for POI Representation Learning
POI-Enhancer uses LLM-generated text features and attention-based fusion to improve POI embeddings from six classic models, gaining consistent accuracy on three real-world mobility datasets.
-
FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios
A four-agent LLM pipeline trained with role-specific data improves human preference on comprehensive Chinese financial analysis tasks.
Discussion (0). Sign in to comment.