REVIEW 21 cited by
Orca 2: Teaching Small Language Models How to Reason
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Orca 1 learns from rich signals, such as explanation traces, allowing it to outperform conventional instruction-tuned models on benchmarks like BigBench Hard and AGIEval. In Orca 2, we continue exploring how improved training signals can enhance smaller LMs' reasoning abilities. Research on training small LMs has often relied on imitation learning to replicate the output of more capable models. We contend that excessive emphasis on imitation may restrict the potential of smaller models. We seek to teach small LMs to employ different solution strategies for different tasks, potentially different from the one used by the larger model. For example, while larger models might provide a direct answer to a complex task, smaller models may not have the same capacity. In Orca 2, we teach the model various reasoning techniques (step-by-step, recall then generate, recall-reason-generate, direct answer, etc.). More crucially, we aim to help the model learn to determine the most effective solution strategy for each task. We evaluate Orca 2 using a comprehensive set of 15 diverse benchmarks (corresponding to approximately 100 tasks and over 36,000 unique prompts). Orca 2 significantly surpasses models of similar size and attains performance levels similar or better to those of models 5-10x larger, as assessed on complex tasks that test advanced reasoning abilities in zero-shot settings. make Orca 2 weights publicly available at aka.ms/orca-lm to support research on the development, evaluation, and alignment of smaller LMs
Forward citations
Cited by 21 Pith papers
-
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
Rationale-augmented finetuning can hurt accuracy while improving calibration, with the sizes of both effects tied linearly to task difficulty.
-
PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects
PertReasonQA scores AI models on cell-state-conditioned mechanistic reasoning about perturbation effects, and PertReasonLM, trained with reasoning supervision, reaches 0.736 balanced accuracy and 0.976 edge recall ver...
-
GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
A 303,581-row Korean instruction corpus generated seedlessly from a 1,084-discipline taxonomy, with near-zero duplicates and low measured overlap with KMMLU, KoBEST, and HAE-RAE-Bench.
-
Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs
Predict-then-Diffuse predicts response lengths for diffusion LLMs via an auxiliary model and safety buffer to reduce FLOP waste while preserving output quality.
-
On the Effect of Instruction Tuning Loss on Generalization
Weighted Instruction Tuning, with low-to-moderate prompt weight and moderate-to-high response weight, beats the standard response-only instruction tuning loss in most of the 75 (model, dataset, benchmark) settings tested.
-
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
An LLM-augmented synthetic data pipeline produces the largest public eviction-focused SDoH dataset (14 categories) and fine-tuned open LLMs that outperform prompt-optimized GPT-4o on the authors' test sets.
-
CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory
A knowledge-component diagnostic pipeline, inspired by cognitive diagnosis theory, generates weakness-targeted synthetic data that improves small LLMs on math, code, and exam benchmarks by up to 13.1 percentage points.
-
Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-ray Report Generation
AM-MRG combines disease-region extraction with two Hopfield memory retrievers to improve LLM-generated chest X-ray reports on IU X-ray, MIMIC-CXR, and Chexpert Plus.
-
Efficient Knowledge Injection in LLMs via Self-Distillation
Self-distillation from a model's own in-context answers injects factual knowledge into LLM weights more efficiently than supervised fine-tuning and is competitive with RAG.
-
C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness
C3oT uses prompt-conditioned fine-tuning on both long and short chain-of-thought data to generate about 50 percent shorter reasoning traces with roughly unchanged accuracy.
-
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
A graph-guided self-training method lets video-language models generate their own reasoning training data from raw videos, improving multi-step compositional reasoning accuracy.
-
Toward Cybersecurity-Expert Small Language Models
A family of 4B–20B cybersecurity models fine-tuned on an enriched, expert-steered reasoning dataset matches or beats larger frontier models on core CTI benchmarks.
-
VERA: Variational Inference Framework for Jailbreaking Large Language Models
VERA frames black-box jailbreaking as variational inference, training a LoRA-tuned attacker that samples diverse fluent prompts; reported ASRs are high but several evaluation choices weaken the SOTA claims.
-
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
SLMEval fits a latent strength distribution to human preferences by maximum entropy and uses it to reweight LLM judge scores, reporting stronger correlation with human judgment on two production tasks.
-
Few-shot Policy (de)composition in Conversational Question Answering
A few-shot neuro-symbolic pipeline decomposes policies into logic formulas and evaluates them with three-valued logic, reaching near state-of-the-art accuracy on ShARC without task-specific fine-tuning of its decompos...
-
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
For 3B-7B LLMs, larger batch sizes with lower learning rates improve instruction-tuning benchmarks, early gradient and loss signals predict final quality, and stacked training matches phased training with fewer samples.
-
Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
Zero-shot LLMs, especially Claude 3 Sonnet and a fine-tuned 7B Mistral variant, classify semantic relations between engineering research topics with high F1 on the new IEEE-Rel-1K benchmark.
-
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.
-
Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
A comprehensive survey of cross-lingual aspect-based sentiment analysis that catalogs tasks, datasets, modeling paradigms, and cross-lingual transfer techniques, and identifies research gaps.
-
DNA 1.0 Technical Report
DNA 1.0 8B Instruct is an 8-billion-parameter Korean-English model that reports state-of-the-art results on Korean benchmarks through continual pre-training, SLERP merging, DPO, and distillation.
-
FASTNav: Fine-tuned Adaptive Small-language-models Trained for Multi-point Robot Navigation
Fine-tuned small language models, coached by a GPT-4 teacher through iterative prompting, can perform multi-point robot navigation on edge devices with success rates approaching larger models.
Discussion (0). Continue with ORCID to comment.