Pith. sign in

REVIEW 7 cited by

Stable LM 2 1.6B Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17834 v1 pith:SOX4KS5V submitted 2024-02-27 cs.CL stat.ML

classification cs.CLstat.ML
keywords reportmodelstablelmbenchmarksmodelsopentechnicaladdition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce StableLM 2 1.6B, the first in a new generation of our language model series. In this technical report, we present in detail the data and training procedure leading to the base and instruction-tuned versions of StableLM 2 1.6B. The weights for both models are available via Hugging Face for anyone to download and use. The report contains thorough evaluations of these models, including zero- and few-shot benchmarks, multilingual benchmarks, and the MT benchmark focusing on multi-turn dialogues. At the time of publishing this report, StableLM 2 1.6B was the state-of-the-art open model under 2B parameters by a significant margin. Given its appealing small size, we also provide throughput measurements on a number of edge devices. In addition, we open source several quantized checkpoints and provide their performance metrics compared to the original model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Across eight zero-shot intent datasets, instruction-tuned ~3B open-weight models can match or beat larger base models, top systems are statistically tied on MASSIVE, and SNIPS is saturated.

  2. How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing

    cs.CL 2026-02 reject novelty 6.0 of 10

    Correct and incorrect single-token continuations of factual queries are separated by rotation of displacement vectors in transformer hidden states, with larger models also suppressing the correct token when forced to ...

  3. MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A skill-by-skill mixture-of-experts router lets a sub-3B vision-language model beat much larger models on autonomous-driving and robot-reasoning benchmarks.

  4. Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning

    cs.CL 2025-08 conditional novelty 5.0 of 10

    RoMed and CCL: a 144k-question perturbation benchmark for medical VQA and a consistency-plus-contrastive training method that improves LLaVA-Med's accuracy and reduces answer variation.

  5. ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Adding information-gain-selected virtual views refined by video diffusion priors to 3D Gaussian Splatting improves arbitrary-view rendering quality.

  6. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    cs.AI 2026-06 conditional novelty 4.0 of 10

    Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.

  7. Supernova: Achieving More with Less in Transformer Architectures

    cs.CL 2025-07 reject novelty 3.0 of 10

    The authors claim a 650M-parameter transformer with a custom 128k byte-level BPE tokenizer reaches 90% of 1B-model average benchmark performance using 100B training tokens.

Pith tools