Pith. sign in

REVIEW 11 cited by

Explanations from Large Language Models Make Small Reasoners Better

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.06726 v1 pith:UZQXN5UT submitted 2022-10-13 cs.CL

classification cs.CL
keywords explanationsmodelsreasoningsmallbettercapabilitiesexplanationfinetuning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Integrating free-text explanations to in-context learning of large language models (LLM) is shown to elicit strong reasoning capabilities along with reasonable explanations. In this paper, we consider the problem of leveraging the explanations generated by LLM to improve the training of small reasoners, which are more favorable in real-production deployment due to their low cost. We systematically explore three explanation generation approaches from LLM and utilize a multi-task learning framework to facilitate small models to acquire strong reasoning power together with explanation generation capabilities. Experiments on multiple reasoning tasks show that our method can consistently and significantly outperform finetuning baselines across different settings, and even perform better than finetuning/prompting a 60x larger GPT-3 (175B) model by up to 9.5% in accuracy. As a side benefit, human evaluation further shows that our method can generate high-quality explanations to justify its predictions, moving towards the goal of explainable AI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Rationale-augmented finetuning can hurt accuracy while improving calibration, with the sizes of both effects tied linearly to task difficulty.

  2. Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation

    cs.CL 2025-09 conditional novelty 6.0 of 10

    CoPeD trains smaller language models with a correction task for wrong rationales and loss-based sample weighting, improving accuracy and rationale faithfulness on several reasoning benchmarks.

  3. ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    ReaLM trains small language models to learn from both right and wrong reasoning chains, then fades the chains out so the model reasons independently, improving benchmark accuracy.

  4. Main Predicate and Their Arguments as Explanation Signals For Intent Classification

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A dependency-parse-based automatic annotation creates word-level explanations for intent classification, and models trained to attend to these signals improve plausibility on held-out ATIS and SNIPS test sets.

  5. A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI

    cs.CL 2024-12 conditional novelty 6.0 of 10

    LLM-generated explanations, paired with a few human labels, produce model judgment distributions as close to human judgment distributions as human explanations do on NLI.

  6. MEGL: Multimodal Explanation-Guided Learning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multimodal explanation-guided learning framework that jointly uses visual saliency maps and textual rationales to train image classifiers, improving accuracy, visual explanation overlap, and text explanation scores ...

  7. TREK: Distill to Explore, Reinforce to Refine

    cs.LG 2026-07 conditional novelty 5.0 of 10

    TREK uses verified teacher proposals to expand a student model's exploration support before standard GRPO refinement, improving performance on hard math and agentic tasks.

  8. AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes

    cs.AI 2025-06 reject novelty 5.0 of 10

    AgentDistill distills agent capabilities without any training by having a teacher generate reusable MCP tool boxes that small-model students invoke at inference time.

  9. Enhancing Generalization in Chain of Thought Reasoning for Smaller Models

    cs.LG 2025-01 reject novelty 4.0 of 10

    PRADA combines P-Tuning and domain-adversarial training with CoT distillation and claims improved cross-domain reasoning in small models, though the evaluation is confounded by target-data access.

  10. SWSC: Shared Weight for Similar Channel in LLM

    cs.LG 2025-01 conditional novelty 4.0 of 10

    SWSC combines channel K-means clustering with an SVD low-rank error correction to compress LLM weights, and reports lower perplexity than RTN quantization on Llama-2-7B Q and K projections at 2 to 3 average bits.

  11. Enhancing the Reasoning Capabilities of Small Language Models via Solution Guidance Fine-Tuning

    cs.CL 2024-12 conditional novelty 4.0 of 10

    SGFT fine-tunes a small model to produce calculation-free solution plans and uses a second model to answer from them, outperforming CoT fine-tuning with roughly 3% of the training data.

Pith tools