REVIEW 7 cited by
Stable LM 2 1.6B Technical Report
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce StableLM 2 1.6B, the first in a new generation of our language model series. In this technical report, we present in detail the data and training procedure leading to the base and instruction-tuned versions of StableLM 2 1.6B. The weights for both models are available via Hugging Face for anyone to download and use. The report contains thorough evaluations of these models, including zero- and few-shot benchmarks, multilingual benchmarks, and the MT benchmark focusing on multi-turn dialogues. At the time of publishing this report, StableLM 2 1.6B was the state-of-the-art open model under 2B parameters by a significant margin. Given its appealing small size, we also provide throughput measurements on a number of edge devices. In addition, we open source several quantized checkpoints and provide their performance metrics compared to the original model.
Forward citations
Cited by 7 Pith papers
-
Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
Across eight zero-shot intent datasets, instruction-tuned ~3B open-weight models can match or beat larger base models, top systems are statistically tied on MASSIVE, and SNIPS is saturated.
-
How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing
Correct and incorrect single-token continuations of factual queries are separated by rotation of displacement vectors in transformer hidden states, with larger models also suppressing the correct token when forced to ...
-
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
A skill-by-skill mixture-of-experts router lets a sub-3B vision-language model beat much larger models on autonomous-driving and robot-reasoning benchmarks.
-
Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
RoMed and CCL: a 144k-question perturbation benchmark for medical VQA and a consistency-plus-contrastive training method that improves LLaVA-Med's accuracy and reduces answer variation.
-
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
Adding information-gain-selected virtual views refined by video diffusion priors to 3D Gaussian Splatting improves arbitrary-view rendering quality.
-
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.
-
Supernova: Achieving More with Less in Transformer Architectures
The authors claim a 650M-parameter transformer with a custom 128k byte-level BPE tokenizer reaches 90% of 1B-model average benchmark performance using 100B training tokens.
Discussion (0). Sign in to comment.