Pith. sign in

REVIEW 11 cited by

xVal: A Continuous Numerical Tokenization for Scientific Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02989 v2 pith:EN7KHWA7 submitted 2023-10-04 stat.ML cs.AIcs.CLcs.LG

classification stat.MLcs.AIcs.CLcs.LG
keywords modelsscientificlanguagedatasetsxvalnumbersnumericaltext
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tokenizing 3D Molecule Structure with Quantized Spherical Coordinates

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Mol-StrucTok tokenizes 3D molecular coordinates via a spherical line notation and VQ-VAE, enabling fast GPT-2 based generation and small property-prediction improvements.

  2. multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data

    cs.LG 2025-05 conditional novelty 6.0 of 10

    multivariateGPT extends next-token prediction to jointly predict the class and continuous value of mixed categorical and numeric time series, with Gaussian uncertainty, and outperforms discrete-token baselines on clin...

  3. POI-Enhancer: An LLM-based Semantic Enhancement Framework for POI Representation Learning

    cs.AI 2025-02 conditional novelty 6.0 of 10

    POI-Enhancer uses LLM-generated text features and attention-based fusion to improve POI embeddings from six classic models, gaining consistent accuracy on three real-world mobility datasets.

  4. A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A multimodal transformer predicts ODE/PDE solutions and generates correct scientific text descriptions from numerical and symbolic inputs, with low error on in-distribution and out-of-distribution tests.

  5. 3DMolFormer: A Dual-channel Framework for Structure-based Drug Discovery

    cs.CE 2025-02 conditional novelty 6.0 of 10

    A dual-channel transformer that reads and writes 3D coordinates as continuous numbers alongside chemical tokens achieves state-of-the-art docking and pocket-aware molecule generation.

  6. Generating particle physics Lagrangians with transformers

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A BART transformer can generate gauge-invariant Lagrangians from field content with over 90% accuracy on in-distribution data, though its performance drops on realistic Standard Model benchmarks.

  7. ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A fine-tuned vision-language model converts images or text into JSON sewing-pattern descriptions, enabling estimation, generation, and editing of 3D garments.

  8. Enhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A wavelet-based tokenizer that lets an autoregressive transformer forecast quantized wavelet coefficients instead of raw values, improving accuracy and generalization on time series benchmarks.

  9. Retrofitting Large Language Models with Dynamic Tokenization

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A pretrained hypernetwork enables dynamic, batch-specific tokenization that compresses token sequences by 20% or more in multilingual models with under 2% average accuracy loss.

  10. FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios

    cs.CE 2025-07 conditional novelty 5.0 of 10

    A four-agent LLM pipeline trained with role-specific data improves human preference on comprehensive Chinese financial analysis tasks.

  11. Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges

    cs.LG 2024-12 conditional novelty 2.0 of 10

    A position paper proposing a research agenda for AI-driven scientific discovery, centered on benchmarks, science agents, multimodal representations, and unified reasoning.

Pith tools