Pith. sign in

REVIEW 1 major objections 2 minor 103 cited by

A Survey on Knowledge Distillation of Large Language Models

T0 review · 1 major / 2 minor · reviewed 2026-05-17 · grok-4.3

Pith's one-line read Knowledge distillation transfers advanced capabilities from proprietary LLMs like GPT-4 to open-source models such as LLaMA and Mistral.

desk verdict A useful organizing survey on KD for LLMs that groups work by algorithms, skills, and applications while stressing data augmentation, but its claims about reliable transfer of ethical alignment and semantic depth rest on thin critical review. read the letter →

arxiv 2402.13116 v4 pith:ODXWRQYN submitted 2024-02-20 cs.CL

classification cs.CL
keywords knowledgedistillationlargelanguagemodelsmodelcompressiondataaugmentationself-improvementopen-sourceLLMssurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey positions knowledge distillation as a central technique for moving sophisticated abilities from closed large language models to accessible open-source versions while also supporting compression and self-improvement loops. It organizes existing work into three main pillars covering distillation algorithms, targeted skill improvements, and domain-specific applications. The survey further examines how data augmentation generates richer training examples inside the distillation process, allowing smaller models to approach the contextual understanding and alignment seen in larger proprietary systems. A reader would care because the approach provides concrete routes to deploy powerful language capabilities without full-scale training resources or direct access to closed models.

What carries the argument

The three foundational pillars of algorithm, skill, and verticalization together with the use of data augmentation to create context-rich training data inside the KD framework.

What would settle it

Controlled experiments comparing open-source model performance on semantic depth and ethical alignment benchmarks when trained with versus without data-augmented KD, checking whether the approximated capabilities consistently appear.

Watch

Extended reading notes

Core claim

In the era of Large Language Models (LLMs), Knowledge Distillation (KD) emerges as a pivotal methodology for transferring advanced capabilities from leading proprietary LLMs, such as GPT-4, to their open-source counterparts like LLaMA and Mistral. Additionally, as open-source LLMs flourish, KD plays a crucial role in both compressing these models, and facilitating their self-improvement by employing themselves as teachers. The survey structures its examination around three foundational pillars: algorithm, skill, and verticalization, while highlighting the interplay between data augmentation and KD to bolster LLMs' performance by generating context-rich, skill-specific training data.

Load-bearing premise

That data augmentation within the KD framework can reliably enable open-source models to approximate the contextual adeptness, ethical alignment, and deep semantic insights of proprietary models.

Editorial extensions

If this is right

  • Open-source models gain the ability to approximate contextual adeptness and ethical alignment of proprietary models through data-augmented KD.
  • Large models can be compressed for more efficient deployment while retaining core capabilities.
  • Models achieve self-improvement by distilling knowledge from their own generated outputs.
  • KD techniques become practical across diverse application fields through verticalization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Stronger data augmentation strategies could accelerate closing the performance gap between open and closed models beyond what scaling laws alone predict.
  • The same distillation-plus-augmentation pattern may transfer to non-language domains such as vision or multimodal systems.
  • Legal and ethical compliance requirements noted in the survey point toward needed auditing methods for models that inherit behaviors through distillation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 2 minor

Summary. The paper surveys Knowledge Distillation (KD) techniques applied to Large Language Models (LLMs). It positions KD as central for transferring capabilities from proprietary models (e.g., GPT-4) to open-source counterparts (e.g., LLaMA, Mistral), for model compression, and for self-improvement via teacher-student setups. The survey is organized around three pillars—algorithm, skill, and verticalization—and stresses the synergy with data augmentation (DA) to generate synthetic data that lets smaller models approximate proprietary models' contextual adeptness, ethical alignment, and semantic insights. It includes a GitHub repository and calls for ethical compliance.

Significance. A well-executed survey in this area would be useful for researchers seeking an organized overview of KD methods, compression strategies, and application domains for LLMs. The explicit linkage of DA to KD and the provision of a curated repository constitute concrete strengths that could accelerate follow-on work if the coverage is balanced.

major comments (1)
  1. [Abstract / DA-KD section] Abstract and the DA-KD interplay discussion: the claim that DA-augmented KD enables open-source models to 'approximate the contextual adeptness, ethical alignment, and deep semantic insights' of proprietary models is presented without a dedicated critical review of transfer-fidelity risks (distribution shift between teacher outputs and downstream user contexts, bias amplification in synthetic data, or loss of implicit chain-of-thought structure). Because this approximation is the central justification for asserting that KD 'transcends traditional boundaries,' the absence of such analysis is load-bearing.
minor comments (2)
  1. The three-pillar structure (algorithm, skill, verticalization) is announced but the manuscript would benefit from an explicit mapping table or section numbers that link each cited work to one or more pillars.
  2. The GitHub link is given; the survey should state the last update date and the criteria used for inclusion/exclusion of papers to improve reproducibility.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the constructive review and for recognizing the survey's organization, the linkage between data augmentation and knowledge distillation, and the value of the accompanying repository. We address the single major comment below.

read point-by-point responses
  1. Referee: [Abstract / DA-KD section] Abstract and the DA-KD interplay discussion: the claim that DA-augmented KD enables open-source models to 'approximate the contextual adeptness, ethical alignment, and deep semantic insights' of proprietary models is presented without a dedicated critical review of transfer-fidelity risks (distribution shift between teacher outputs and downstream user contexts, bias amplification in synthetic data, or loss of implicit chain-of-thought structure). Because this approximation is the central justification for asserting that KD 'transcends traditional boundaries,' the absence of such analysis is load-bearing.

    Authors: We agree that the current presentation would be strengthened by an explicit discussion of the risks and limitations of DA-augmented KD. In the revised manuscript we will insert a dedicated subsection on challenges within the DA-KD interplay section. This subsection will address distribution shift between teacher-generated outputs and downstream user contexts, the potential for bias amplification in synthetic data, and the risk of losing implicit reasoning structures such as chain-of-thought. The added analysis will qualify the claim that KD 'transcends traditional boundaries' and provide readers with a more balanced view of transfer fidelity. We view this as a substantive improvement that directly responds to the load-bearing nature of the point. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity in this literature survey on LLM knowledge distillation

full rationale

This paper is a survey reviewing existing KD methods for LLMs, structured around algorithm, skill, and verticalization pillars with discussion of data augmentation interplay. It contains no new mathematical derivations, equations, fitted parameters, or predictions that could reduce to inputs by construction. All claims summarize cited prior literature without self-referential load-bearing steps or uniqueness theorems imported from the authors' own work. The central narrative on DA-augmented KD enabling approximation of proprietary capabilities is presented as a synthesis of external research rather than an internally derived result, making the work self-contained as a review with no circularity.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters, axioms, or invented entities; it aggregates existing research without postulating new mechanisms.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Knowledge Distillation of Large Language Models." pith.science (2026). https://pith.science/paper/ODXWRQYN

@misc{pith2026240213116,
  author       = {Pith},
  title        = {Pith review of: A Survey on Knowledge Distillation of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ODXWRQYN}},
  note         = {Machine review of arXiv:2402.13116}
}
read the original abstract

In the era of Large Language Models (LLMs), Knowledge Distillation (KD) emerges as a pivotal methodology for transferring advanced capabilities from leading proprietary LLMs, such as GPT-4, to their open-source counterparts like LLaMA and Mistral. Additionally, as open-source LLMs flourish, KD plays a crucial role in both compressing these models, and facilitating their self-improvement by employing themselves as teachers. This paper presents a comprehensive survey of KD's role within the realm of LLM, highlighting its critical function in imparting advanced knowledge to smaller models and its utility in model compression and self-improvement. Our survey is meticulously structured around three foundational pillars: \textit{algorithm}, \textit{skill}, and \textit{verticalization} -- providing a comprehensive examination of KD mechanisms, the enhancement of specific cognitive abilities, and their practical implications across diverse fields. Crucially, the survey navigates the intricate interplay between data augmentation (DA) and KD, illustrating how DA emerges as a powerful paradigm within the KD framework to bolster LLMs' performance. By leveraging DA to generate context-rich, skill-specific training data, KD transcends traditional boundaries, enabling open-source models to approximate the contextual adeptness, ethical alignment, and deep semantic insights characteristic of their proprietary counterparts. This work aims to provide an insightful guide for researchers and practitioners, offering a detailed overview of current methodologies in KD and proposing future research directions. Importantly, we firmly advocate for compliance with the legal terms that regulate the use of LLMs, ensuring ethical and lawful application of KD of LLMs. An associated Github repository is available at https://github.com/Tebmer/Awesome-Knowledge-Distillation-of-LLMs.

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

  • Cost.FunctionalEquation washburn_uniqueness_aczel unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    KD emerges as a pivotal methodology for transferring advanced capabilities from leading proprietary LLMs, such as GPT-4, to their open-source counterparts like LLaMA and Mistral, while also enabling model compression and self-improvement.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Forward citations

Showing 60 of 103 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 103 Pith citations

  1. Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Byte-Prefix Marginalization maps a teacher's next-token distribution onto the student's vocabulary through shared byte prefixes plus an explicit residual, giving a mass-preserving target for on-policy distillation acr...

  2. ADS-C: Antidistillation Sampling for Classification

    cs.LG 2026-07 accept novelty 7.0 of 10

    ADS-C perturbs served classification probabilities under a per-input margin budget, preserving every top-1 prediction while degrading distilled students by 13–30 percentage points.

  3. SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    SENTINEL generates targeted tasks from model failures in a Controller-Proposer-Solver loop, raising Pass^1 from 66.4 to 74.9 on Tau2-Bench Retail and outperforming standard RL.

  4. When Context Returns: Toward Robust Internalization in On-Policy Distillation

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    A stop-gradient consistency regularizer mitigates context-induced degradation in on-policy distillation, improving robustness across 12 configurations.

  5. Escaping the KL Agreement Trap in On-Policy Distillation

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    KAT detects persistent low-KL agreement traps in on-policy distillation via a dynamic threshold to filter weak supervision, improving avg@k by 2.66% and pass@k by 3.43% on four math benchmarks while shortening rollout...

  6. When Does Model Collapse Occur in Structured Interactive Learning?

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Model collapse occurs in structured interactive learning if and only if the directed interaction graph satisfies a specific topological condition, with finite-sample guarantees for linear regression and asymptotic res...

  7. Multi-Rollout On-Policy Distillation via Peer Successes and Failures

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    MOPD improves on-policy distillation for LLMs by using peer successes for positive patterns and failures for negative examples to create more informative teacher signals.

  8. Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing

    cs.LG 2026-02 conditional novelty 7.0 of 10

    In competitive ML markets, standard gradient training can drive learners into overspecialized equilibria with arbitrarily poor global performance; a proposed 'peer probing' algorithm provably escapes this under inform...

  9. Stable On-Policy Distillation through Adaptive Target Reformulation

    cs.LG 2026-01 unverdicted novelty 7.0 of 10

    Veto replaces the teacher's target distribution with a product of teacher and student probabilities, stabilizing on-policy knowledge distillation and improving small-model benchmarks.

  10. CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

    cs.SE 2025-10 conditional novelty 7.0 of 10

    CodeRL+ integrates variable-level execution trajectory inference into RLVR training to align textual code representations with execution semantics, delivering 4.6% relative pass@1 gains and generalization to code-reas...

  11. Active Data Curation Effectively Distills Large-Scale Multimodal Models

    cs.CV 2024-11 conditional novelty 7.0 of 10

    Selecting training data by a reference model's loss acts as an implicit distillation, and combining it with explicit distillation yields more FLOP-efficient vision-language models that beat prior SoTA on 27 benchmarks.

  12. Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation

    cs.LG 2026-08 conditional novelty 6.0 of 10

    SKALD shows that distilling skill-conditioned teacher predictions into a question-only student improves math reasoning more than GRPO alone, with gains concentrated on rollout groups where rewards are uniform.

  13. Reliability-Safety Trade-off in AI Distillation: A Renormalization-Group Approach

    cond-mat.stat-mech 2026-08 conditional novelty 6.0 of 10

    A two-state statistical-mechanics model of AI distillation yields a reliability-safety trade-off controlled by a single hazard-discrimination parameter, whose iterated distillation obeys a renormalization-group flow w...

  14. TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Distilling agreement-filtered multi-LRM jury annotations into Gemma-3-12B improves MQM translation quality evaluation from 52.63% to 55.03% average segment-level accuracy, approaching closed LRMs.

  15. Learning from the Future: Privileged Self-Distillation for Sequential Recommendation

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Privileged Self-Distillation turns future interactions into soft training targets for a causal sequential recommender via dual attention masks, gated KL distillation, and an EMA teacher.

  16. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  17. Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Distilling an 8B reasoning teacher into a 0.6B student recovers most summary quality at ~50× speed, but teacher type—not scale alone—determines which capabilities transfer.

  18. DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    DuoMem distills from a 72B teacher to 4B student via context and parameter space, achieving 77.9% success on ALFWorld vs 4.3% baseline.

  19. ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    ARKD uses an RL policy network to adaptively balance FKL and RKL in LLM distillation, claiming gains of 0.4-0.6 points on Rouge-L and BertScore over baselines.

  20. Labeling Training Data for Entity Matching Using Large Language Models

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    LLM-labeled training sets for entity matching produce student models with F1 scores within 2 points of benchmark-trained models on five datasets at a cost of $28-41 versus 470 hours of manual work.

  21. Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Activation steering on early layers improves diversity of synthetic data for low-resource languages and often boosts downstream classifier performance compared to non-steered prompting.

  22. Llamion Technical Report

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    A new conversion method (KEPT) transforms Orion-14B into Llama-format models while preserving benchmark performance using ~123M tokens of distillation.

  23. Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    DASD improves math reasoning in LLMs by adaptively directing self-distillation based on per-token entropy to balance exploration and step accuracy, outperforming prior self-distillation and RLVR baselines on six benchmarks.

  24. Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

    cs.SE 2026-05 unverdicted novelty 6.0 of 10

    FireFly inverts task synthesis by exploring real MCP servers first via pairwise tool graphs and sub-DAG sampling, then generates 5,144 verified tasks backward from outcomes to train a 4B model that matches Claude Sonn...

  25. OpenJarvis: Personal AI, On Personal Devices

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    OpenJarvis decomposes personal AI into Intelligence, Engine, Agents, Tools & Memory, and Learning primitives and applies LLM-guided spec search to produce on-device configurations that reach within 3.2 pp of cloud bas...

  26. SOD: Step-wise On-policy Distillation for Small Language Model Agents

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    SOD reweights on-policy distillation strength step-by-step using divergence to stabilize tool use in small language model agents, yielding up to 20.86% gains and 26.13% on AIME 2025 for a 0.6B model.

  27. SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    SimCT recovers discarded teacher signal in cross-tokenizer on-policy distillation by enlarging supervision to jointly realizable multi-token continuations, yielding consistent gains on math reasoning and code generati...

  28. UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    UniSD unifies self-distillation components for autoregressive LLMs and its full integrated version improves base models by 5.4 points and baselines by 2.8 points across six benchmarks.

  29. Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Reasoning in LLMs helps only for complex 27-class emotion recognition and hurts simple binary sentiment, across seven model families and 504 configurations.

  30. What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects

    cs.SE 2025-09 conditional novelty 6.0 of 10

    LLM-generated refactoring motivations agree with expert raters about 80% of the time, align with literature motivations in roughly half of cases, and correlate only weakly with software metrics.

  31. SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding

    eess.SP 2025-09 conditional novelty 6.0 of 10

    A semantic communication framework uses MLLM-derived scenario-aware importance labels to allocate coding resources, improving PSNR of important image regions at comparable or lower bandwidth than prior JSCC systems.

  32. Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A decoding pipeline reuses cached tokens and anneals sampling temperature to accelerate self-consistency hallucination detection by up to 3x without meaningful AUROC loss.

  33. PROVCREATOR: Synthesizing Complex Heterogenous Graphs with Node and Edge Attributes

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ProvCreator is a framework that serializes complex heterogeneous graphs into token sequences and fine-tunes LLaMA 3.2 3B to generate new graphs with structure and attributes generated jointly.

  34. GeLaCo: An Evolutionary Approach to Layer Compression

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Evolutionary search over layer-merging configurations, scored by module-wise activation similarity, yields competitive LLM compression and the first size-quality Pareto fronts.

  35. Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Intercepting intermediate features in split neural network inference lets black-box attackers build surrogate models whose adversarial examples transfer to the target far more often, e.g., 96% versus 61% success in on...

  36. Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A flipped distillation method lets a decoder-only LLM learn text-matching similarity from a smaller encoder teacher through LoRA and a margin-aware contrastive loss, improving matching accuracy and online FAQ retrieval.

  37. Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Structured moral prompts, especially first-principles reasoning, improve LLM moral classification accuracy across 12 open models and four benchmarks, and reasoning distillation transfers these gains to a 3B model.

  38. Event-Priori-Based Vision-Language Model for Efficient Visual Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EP-VLM uses event-camera motion data to sparsify image patches before a vision-language model processes them, cutting FLOPs by about half with a small accuracy drop.

  39. What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding

    cs.CL 2025-06 conditional novelty 6.0 of 10

    FoReaL-Decoding lets a strong reasoning model generate the first few tokens of each sentence and a weaker model complete the sentence, cutting theoretical FLOPs by 30-55% while retaining 86-100% of accuracy on four ma...

  40. EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EasyDistill packages established LLM knowledge-distillation techniques into a single modular toolkit with released distilled models, datasets, and Alibaba Cloud integration.

  41. Efficient Long CoT Reasoning in Small Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Binary cutting with on-policy validation prunes redundant chain-of-thought steps in teacher traces, letting 7B models keep most long-CoT accuracy while generating fewer tokens.

  42. Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A gated Shift-FFN adapter that adds the previous token's representation to the current token's before the feedforward layer reduces repetitive looping and improves math accuracy in LoRA fine-tuned models trained on lo...

  43. VERDI: VLM-Embedded Reasoning for Autonomous Driving

    cs.RO 2025-05 conditional novelty 6.0 of 10

    VERDI aligns perception, prediction, and planning outputs of end-to-end AD models with VLM-generated text features at training time to embed structured reasoning, yielding up to 11% better l2 distance and 10% higher n...

  44. DarwinLM: Evolutionary Structured Pruning of Large Language Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    DarwinLM uses evolutionary search with training-aware offspring selection to prune LLMs, beating ShearedLlama with 5x less post-training data.

  45. LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    LongReD reduces short-text performance loss after long-context extension by training the extended model to match the original model's hidden states on short texts and using skipped position indices to bridge short and...

  46. Who Taught You That? Tracing Teachers in Model Distillation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Part-of-speech template patterns in a distilled student model's outputs can identify its teacher above chance in a closed set of five candidate models, outperforming n-gram and similarity baselines on most tasks.

  47. Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A new reflection-based AI method makes LLM-generated virtual students predict real students' future quiz performance better than deep learning knowledge-tracing baselines.

  48. TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models

    cs.LG 2025-01 conditional novelty 6.0 of 10

    TAID improves knowledge distillation for language models by gradually shifting the student's training target from its own distribution to the teacher's distribution over time.

  49. StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel

    cs.LG 2025-01 conditional novelty 6.0 of 10

    StagFormer staggers transformer layers one time step apart with delayed cross-attention, enabling depth-parallel decoding at quality comparable to a deeper baseline.

  50. HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages

    cs.DC 2024-12 conditional novelty 6.0 of 10

    Fine-tuning DeepSeek-Coder on a new 122k synthetic HPC instruction dataset yields open-source models that reach 34.1 pass@1 on ParEval parallel code generation, besting other open baselines but trailing GPT-4.

  51. Densing Law of LLMs

    cs.AI 2024-12 reject novelty 6.0 of 10

    Maximum LLM capability per parameter, measured on five benchmarks, has grown exponentially, doubling about every three months.

  52. Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS

    cs.CL 2024-11 conditional novelty 6.0 of 10

    HiAR-ICL builds reusable 'thought card' reasoning templates via MCTS and shows a 7B model with these templates outperforms GPT-4o on MATH and AMC.

  53. Reassessing Layer Pruning in LLMs: New Insights and Methods

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Trimming the final 25% of layers and fine-tuning the head and last three layers outperforms sophisticated pruning metrics and LoRA-based recovery for LLM compression.

  54. Information Extraction from Clinical Notes: Are We Ready to Switch to Large Language Models?

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Instruction-tuned LLaMA models beat BERT on clinical NER and relation extraction, gaining up to 7% F1 on an unseen institution, but with much higher compute cost and lower throughput.

  55. Themes of Building LLM-based Applications for Production: A Practitioner's View

    cs.SE 2024-11 conditional novelty 6.0 of 10

    Analyzing 189 practitioner YouTube videos yields 20 topics in 8 themes for building LLM applications in production, with RAG systems the most prevalent.

  56. RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Relative Probability Association Metric (RPAM) measures LM associations via softmax-normalized continuation probabilities and correlates strongly with human associations and downstream LM behavior across three models.

  57. Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies

    eess.SY 2026-08 conditional novelty 5.0 of 10

    The authors propose ReManGPT, a conceptual orchestration framework for applying LLMs to remanufacturing, and illustrate it with case studies in disassembly planning, repair guidance, and robotic execution.

  58. Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

    cs.CL 2026-07 conditional novelty 5.0 of 10

    AdaDSF uses per-layer cosine similarity to decide which tokens skip which layers, then distills the sparse model back toward the dense one, cutting FLOPs while holding accuracy close to dense.

  59. Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    Distillation of a 688M-parameter MASt3R teacher yields up to 7x smaller students that retain most lunar reconstruction accuracy and outperform sparse-supervised baselines.

  60. Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    CRST improves ultra-low-resolution text-to-image person retrieval by 5.7% Rank-1 and 5.3% mAP on average across three datasets while stabilizing mixed-resolution galleries.

See all 103 Pith citations

Reference graph

Works this paper leans on

300 extracted references · 300 canonical work pages · cited by 103 Pith papers (see all)

  1. [1]

    Advances in Neural Information Processing Systems , volume=

    Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in Neural Information Processing Systems , volume=

  2. [2]

    arXiv preprint arXiv:2304.14233 , year=

    Large Language Models are Strong Zero-Shot Retriever , author=. arXiv preprint arXiv:2304.14233 , year=

  3. [3]

    arXiv preprint arXiv:2305.07402 , year=

    Knowledge Refinement via Interaction Between Search Engines and Large Language Models , author=. arXiv preprint arXiv:2305.07402 , year=

  4. [4]

    arXiv preprint arXiv:2212.10192 , year=

    Adam: Dense Retrieval Distillation with Adaptive Dark Examples , author=. arXiv preprint arXiv:2212.10192 , year=

  5. [5]

    The Eleventh International Conference on Learning Representations , year=

    HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization , author=. The Eleventh International Conference on Learning Representations , year=

  6. [6]

    arXiv preprint arXiv:2401.00797 , year=

    Distillation is All You Need for Practically Using Different Pre-trained Recommendation Models , author=. arXiv preprint arXiv:2401.00797 , year=

  7. [7]

    Self-Instruct: Aligning Language Models with Self-Generated Instructions

    Self-instruct: Aligning language model with self generated instructions , author=. arXiv preprint arXiv:2212.10560 , year=

  8. [8]

    WizardLM: Empowering large pre-trained language models to follow complex instructions

    Wizardlm: Empowering large language models to follow complex instructions , author=. arXiv preprint arXiv:2304.12244 , year=

Show all 300 references
  1. [9]

    Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes , booktitle =

    Cheng. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes , booktitle =

  2. [10]

    arXiv preprint arXiv:2306.02707 , year=

    Orca: Progressive learning from complex explanation traces of gpt-4 , author=. arXiv preprint arXiv:2306.02707 , year=

  3. [11]

    2023 , eprint=

    Orca 2: Teaching Small Language Models How to Reason , author=. 2023 , eprint=

  4. [12]

    arXiv preprint arXiv:2310.16944 , year=

    Zephyr: Direct distillation of lm alignment , author=. arXiv preprint arXiv:2310.16944 , year=

  5. [13]

    arXiv preprint arXiv:2310.01377 , year=

    Ultrafeedback: Boosting language models with high-quality feedback , author=. arXiv preprint arXiv:2310.01377 , year=

  6. [14]

    McAuley , title =

    Canwen Xu and Daya Guo and Nan Duan and Julian J. McAuley , title =

  7. [15]

    arXiv preprint arXiv:2303.17760 , year=

    Camel: Communicative agents for" mind" exploration of large scale language model society , author=. arXiv preprint arXiv:2303.17760 , year=

  8. [16]

    arXiv preprint arXiv:1606.07947 , year=

    Sequence-level knowledge distillation , author=. arXiv preprint arXiv:1606.07947 , year=

  9. [17]

    Yuxian Gu and Li Dong and Furu Wei and Minlie Huang , booktitle=. Mini. 2024 , url=

  10. [18]

    International Conference on Machine Learning , pages=

    Less is more: Task-aware layer-wise distillation for language model compression , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  11. [19]

    The Eleventh International Conference on Learning Representations,

    Chen Liang and Haoming Jiang and Zheng Li and Xianfeng Tang and Bing Yin and Tuo Zhao , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =

  12. [20]

    arXiv preprint arXiv:2203.10705 , year=

    Compression of generative pre-trained language models via quantization , author=. arXiv preprint arXiv:2203.10705 , year=

  13. [21]

    arXiv preprint arXiv:2305.17888 , year=

    LLM-QAT: Data-Free Quantization Aware Training for Large Language Models , author=. arXiv preprint arXiv:2305.17888 , year=

  14. [22]

    Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty

    Timiryasov, Inar and Tastet, Jean-Loup. Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty. Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning. 2023. doi:10.1...

  15. [23]

    arXiv preprint arXiv:2306.09306 , year=

    Propagating Knowledge Updates to LMs Through Distillation , author=. arXiv preprint arXiv:2306.09306 , year=

  16. [24]

    arXiv preprint arXiv:2212.10670 , year=

    In-context Learning Distillation: Transferring Few-shot Learning Ability of Pre-trained Language Models , author=. arXiv preprint arXiv:2212.10670 , year=

  17. [25]

    Ning Ding and Yulin Chen and Bokai Xu and Yujia Qin and Shengding Hu and Zhiyuan Liu and Maosong Sun and Bowen Zhou , title =

  18. [26]

    and Stoica, Ion and Xing, Eric P

    Chiang, Wei-Lin and Li, Zhuohan and Lin, Zi and Sheng, Ying and Wu, Zhanghao and Zhang, Hao and Zheng, Lianmin and Zhuang, Siyuan and Zhuang, Yonghao and Gonzalez, Joseph E. and Stoica, Ion and Xing, Eric P. , month =. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90\ url =

  19. [27]

    arXiv preprint arXiv:2305.18395 , year=

    Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks , author=. arXiv preprint arXiv:2305.18395 , year=

  20. [28]

    arXiv preprint arXiv:2305.15225 , year=

    SAIL: Search-Augmented Instruction Learning , author=. arXiv preprint arXiv:2305.15225 , year=

  21. [29]

    arXiv preprint arXiv:2310.11511 , year=

    Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection , author=. arXiv preprint arXiv:2310.11511 , year=

  22. [30]

    arXiv preprint arXiv:2304.11116 , year=

    Graph-ToolFormer: To Empower LLMs with Graph Reasoning Ability via Prompt Augmented by ChatGPT , author=. arXiv preprint arXiv:2304.11116 , year=

  23. [31]

    arXiv preprint arXiv:2312.07000 , year=

    Alignment for Honesty , author=. arXiv preprint arXiv:2312.07000 , year=

  24. [32]

    arXiv preprint arXiv:2309.00267 , year=

    Rlaif: Scaling reinforcement learning from human feedback with ai feedback , author=. arXiv preprint arXiv:2309.00267 , year=

  25. [33]

    2022 , eprint=

    Constitutional AI: Harmlessness from AI Feedback , author=. 2022 , eprint=

  26. [34]

    arXiv preprint arXiv:2312.10665 , year=

    Silkie: Preference Distillation for Large Visual Language Models , author=. arXiv preprint arXiv:2312.10665 , year=

  27. [35]

    Conference on Robot Learning , pages=

    Scaling up and distilling down: Language-guided robot skill acquisition , author=. Conference on Robot Learning , pages=. 2023 , organization=

  28. [36]

    Findings of the Association for Computational Linguistics: ACL 2023 , pages=

    Distilling reasoning capabilities into smaller language models , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=

  29. [37]

    arXiv preprint arXiv:2309.05653 , year=

    Mammoth: Building math generalist models through hybrid instruction tuning , author=. arXiv preprint arXiv:2309.05653 , year=

  30. [38]

    2023 , eprint=

    Textbooks Are All You Need , author=. 2023 , eprint=

  31. [39]

    arXiv preprint arXiv:2309.05463 , year=

    Textbooks are all you need ii: phi-1.5 technical report , author=. arXiv preprint arXiv:2309.05463 , year=

  32. [40]

    Phi-2: The surprising power of small language models , author =

  33. [41]

    arXiv preprint arXiv:2308.09583 , year=

    Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct , author=. arXiv preprint arXiv:2308.09583 , year=

  34. [42]

    2024 , url=

    Kevin Yang and Dan Klein and Asli Celikyilmaz and Nanyun Peng and Yuandong Tian , booktitle=. 2024 , url=

  35. [43]

    arXiv preprint arXiv:2305.12870 , year=

    Lion: Adversarial Distillation of Closed-Source Large Language Model , author=. arXiv preprint arXiv:2305.12870 , year=

  36. [44]

    arXiv preprint arXiv:2307.11769 , year=

    Domain Knowledge Distillation from Large Language Model: An Empirical Study in the Autonomous Driving Domain , author=. arXiv preprint arXiv:2307.11769 , year=

  37. [45]

    arXiv preprint arXiv:2304.06975 , year=

    Huatuo: Tuning llama model with chinese medical knowledge , author=. arXiv preprint arXiv:2304.06975 , year=

  38. [46]

    arXiv preprint arXiv:2303.04360 , year=

    Does synthetic data generation of llms help clinical text mining? , author=. arXiv preprint arXiv:2303.04360 , year=

  39. [47]

    arXiv preprint arXiv:2305.15062 , year=

    Lawyer LLaMA Technical Report , author=. arXiv preprint arXiv:2305.15062 , year=

  40. [48]

    GitHub repository , howpublished =

    Hongcheng Liu, Yusheng Liao, Yutong Meng, Yuhao Wang , title =. GitHub repository , howpublished =. 2023 , publisher =

  41. [49]

    International Journal of Computer Vision , volume=

    Knowledge distillation: A survey , author=. International Journal of Computer Vision , volume=. 2021 , publisher=

  42. [50]

    arXiv preprint arXiv:2307.12966 , year=

    Aligning large language models with human: A survey , author=. arXiv preprint arXiv:2307.12966 , year=

  43. [51]

    arXiv preprint arXiv:2312.10997 , year=

    Retrieval-Augmented Generation for Large Language Models: A Survey , author=. arXiv preprint arXiv:2312.10997 , year=

  44. [52]

    Hashimoto , title =

    Rohan Taori and Ishaan Gulrajani and Tianyi Zhang and Yann Dubois and Xuechen Li and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto , title =. GitHub repository , howpublished =. 2023 , publisher =

  45. [53]

    2023 , eprint=

    LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions , author=. 2023 , eprint=

  46. [54]

    arXiv preprint arXiv:2306.08568 , year=

    WizardCoder: Empowering Code Large Language Models with Evol-Instruct , author=. arXiv preprint arXiv:2306.08568 , year=

  47. [55]

    2023 , url =

    Xinyang Geng and Arnav Gudibande and Hao Liu and Eric Wallace and Pieter Abbeel and Sergey Levine and Dawn Song , title =. 2023 , url =

  48. [56]

    arXiv preprint arXiv:2305.15717 , year=

    The false promise of imitating proprietary llms , author=. arXiv preprint arXiv:2305.15717 , year=

  49. [57]

    arXiv preprint arXiv:2301.13688 , year=

    The flan collection: Designing data and methods for effective instruction tuning , author=. arXiv preprint arXiv:2301.13688 , year=

  50. [58]

    2023 , howpublished =

    Ye, Seonghyeon and Jo, Yongrae and Kim, Doyoung and Kim, Sungdong and Hwang, Hyeonbin and Seo, Minjoon , title =. 2023 , howpublished =

  51. [59]

    Wang, Guan and Cheng, Sijie and Zhan, Xianyuan and Li, Xiangang and Song, Sen and Liu, Yang , month = sep, year =

  52. [60]

    2023 , eprint=

    Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks , author=. 2023 , eprint=

  53. [61]

    2023 , eprint=

    Mixed Distillation Helps Smaller Language Model Better Reasoning , author=. 2023 , eprint=

  54. [62]

    2022 , eprint=

    Explanations from Large Language Models Make Small Reasoners Better , author=. 2022 , eprint=

  55. [63]

    Large Language Models Are Reasoning Teachers , booktitle =

    Namgyu Ho and Laura Schmid and Se. Large Language Models Are Reasoning Teachers , booktitle =

  56. [64]

    Teaching Small Language Models to Reason

    Magister, Lucie Charlotte and Mallinson, Jonathan and Adamek, Jakub and Malmi, Eric and Severyn, Aliaksei. Teaching Small Language Models to Reason. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2023. doi:10.1...

  57. [65]

    2023 , eprint=

    Specializing Smaller Language Models towards Multi-Step Reasoning , author=. 2023 , eprint=

  58. [66]

    Advances in Neural Information Processing Systems , volume=

    Principle-driven self-alignment of language models from scratch with minimal human supervision , author=. Advances in Neural Information Processing Systems , volume=

  59. [67]

    GitHub repository , howpublished =

    Sahil Chaudhary , title =. GitHub repository , howpublished =. 2023 , publisher =

  60. [68]

    Qingyi Si and Tong Wang and Zheng Lin and Xu Zhang and Yanan Cao and Weiping Wang , title =

  61. [69]

    GitHub , year=

    Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo , author=. GitHub , year=

  62. [70]

    2023 , eprint=

    Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases , author=. 2023 , eprint=

  63. [71]

    arXiv preprint arXiv:2310.16271 , year=

    CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment , author=. arXiv preprint arXiv:2310.16271 , year=

  64. [72]

    Advances in Neural Information Processing Systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in Neural Information Processing Systems , volume=

  65. [73]

    2023 , eprint=

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model , author=. 2023 , eprint=

  66. [74]

    The Twelfth International Conference on Learning Representations , year=

    On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes , author=. The Twelfth International Conference on Learning Representations , year=

  67. [75]

    arXiv preprint arXiv:1910.01108 , year=

    DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter , author=. arXiv preprint arXiv:1910.01108 , year=

  68. [76]

    OpenAI blog , volume=

    Language models are unsupervised multitask learners , author=. OpenAI blog , volume=

  69. [77]

    arXiv preprint arXiv:2308.06744 , year=

    Token-Scaled Logit Distillation for Ternary Weight Generative Language Models , author=. arXiv preprint arXiv:2308.06744 , year=

  70. [78]

    f-Divergence Minimization for Sequence-Level Knowledge Distillation

    Wen, Yuqiao and Li, Zichao and Du, Wenyu and Mou, Lili. f-Divergence Minimization for Sequence-Level Knowledge Distillation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.605

  71. [79]

    Large Language Models Can Self-Improve

    Huang, Jiaxin and Gu, Shixiang and Hou, Le and Wu, Yuexin and Wang, Xuezhi and Yu, Hongkun and Han, Jiawei. Large Language Models Can Self-Improve. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.67

  72. [80]

    Goodman , title =

    Eric Zelikman and Yuhuai Wu and Jesse Mu and Noah D. Goodman , title =. NeurIPS , year =

  73. [81]

    f -Divergence Inequalities , year=

    Sason, Igal and Verdú, Sergio , journal=. f -Divergence Inequalities , year=

  74. [82]

    Advances in neural information processing systems , volume=

    Policy gradient methods for reinforcement learning with function approximation , author=. Advances in neural information processing systems , volume=

  75. [83]

    2019 , eprint=

    Patient Knowledge Distillation for BERT Model Compression , author=. 2019 , eprint=

  76. [84]

    M obile BERT : a Compact Task-Agnostic BERT for Resource-Limited Devices

    Sun, Zhiqing and Yu, Hongkun and Song, Xiaodan and Liu, Renjie and Yang, Yiming and Zhou, Denny. M obile BERT : a Compact Task-Agnostic BERT for Resource-Limited Devices. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.1865...

  77. [85]

    T iny BERT : Distilling BERT for Natural Language Understanding

    Jiao, Xiaoqi and Yin, Yichun and Shang, Lifeng and Jiang, Xin and Chen, Xiao and Li, Linlin and Wang, Fang and Liu, Qun. T iny BERT : Distilling BERT for Natural Language Understanding. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. doi:10.18653/v...

  78. [86]

    Advances in Neural Information Processing Systems , volume=

    Dynabert: Dynamic bert with adaptive width and depth , author=. Advances in Neural Information Processing Systems , volume=

  79. [87]

    arXiv preprint arXiv:2204.07675 , year=

    Moebert: from bert to mixture-of-experts via importance-guided adaptation , author=. arXiv preprint arXiv:2204.07675 , year=

  80. [88]

    Liang and Weituo Hao and Dinghan Shen and Yufan Zhou and Weizhu Chen and Changyou Chen and Lawrence Carin , title =

    Kevin J. Liang and Weituo Hao and Dinghan Shen and Yufan Zhou and Weizhu Chen and Changyou Chen and Lawrence Carin , title =. 9th International Conference on Learning Representations,. 2021 , url =

  81. [89]

    2017 , eprint=

    Proximal Policy Optimization Algorithms , author=. 2017 , eprint=

  82. [90]

    arXiv preprint arXiv:2306.17492 , year=

    Preference ranking optimization for human alignment , author=. arXiv preprint arXiv:2306.17492 , year=

  83. [91]

    2023 , eprint=

    Instruction Tuning with GPT-4 , author=. 2023 , eprint=

  84. [92]

    arXiv preprint arXiv:2304.05302 , year=

    Rrhf: Rank responses to align language models with human feedback without tears , author=. arXiv preprint arXiv:2304.05302 , year=

  85. [93]

    2023 , eprint=

    Making Large Language Models Better Reasoners with Alignment , author=. 2023 , eprint=

  86. [94]

    2023 , eprint=

    Adapting Large Language Models via Reading Comprehension , author=. 2023 , eprint=

  87. [95]

    2023 , eprint=

    Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering , author=. 2023 , eprint=

  88. [96]

    2022 , eprint=

    PEER: A Collaborative Language Model , author=. 2022 , eprint=

  89. [97]

    2023 , eprint=

    Self-Refine: Iterative Refinement with Self-Feedback , author=. 2023 , eprint=

  90. [98]

    NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=

    Reflection-Tuning: Recycling Data for Better Instruction-Tuning , author=. NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=

  91. [99]

    2024 , url=

    Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning , author=. 2024 , url=

  92. [100]

    2024 , eprint=

    Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements , author=. 2024 , eprint=

  93. [101]

    2022 , eprint=

    Self-critiquing models for assisting human evaluators , author=. 2022 , eprint=

  94. [102]

    arXiv preprint arXiv:2204.05862 , year=

    Training a helpful and harmless assistant with reinforcement learning from human feedback , author=. arXiv preprint arXiv:2204.05862 , year=

  95. [103]

    Advances in Neural Information Processing Systems , volume=

    Learning to summarize with human feedback , author=. Advances in Neural Information Processing Systems , volume=

  96. [104]

    arXiv preprint arXiv:1909.08593 , year=

    Fine-tuning language models from human preferences , author=. arXiv preprint arXiv:1909.08593 , year=

  97. [105]

    2023 , eprint=

    OpenChat: Advancing Open-source Language Models with Mixed-Quality Data , author=. 2023 , eprint=

  98. [106]

    2021 , eprint=

    Recursively Summarizing Books with Human Feedback , author=. 2021 , eprint=

  99. [107]

    2023 , eprint=

    OpenAssistant Conversations -- Democratizing Large Language Model Alignment , author=. 2023 , eprint=

  100. [108]

    Minae Kwon and Sang Michael Xie and Kalesha Bullard and Dorsa Sadigh , title =

  101. [109]

    2023 , eprint=

    Training Language Models with Language Feedback at Scale , author=. 2023 , eprint=

  102. [110]

    Aligning Large Language Models through Synthetic Feedback

    Kim, Sungdong and Bae, Sanghwan and Shin, Jamin and Kang, Soyoung and Kwak, Donghyun and Yoo, Kang and Seo, Minjoon. Aligning Large Language Models through Synthetic Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10....

  103. [111]

    2023 , eprint=

    Training Socially Aligned Language Models on Simulated Social Interactions , author=. 2023 , eprint=

  104. [112]

    Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback

    Roit, Paul and Ferret, Johan and Shani, Lior and Aharoni, Roee and Cideron, Geoffrey and Dadashi, Robert and Geist, Matthieu and Girgin, Sertan and Hussenot, Leonard and Keller, Orgad and Momchev, Nikola and Ramos Garea, Sabela and Stanczyk, Piotr and Vieillard, Nino and Bache...

  105. [113]

    2021 , eprint=

    Ethical and social risks of harm from Language Models , author=. 2021 , eprint=

  106. [114]

    2021 , eprint=

    A General Language Assistant as a Laboratory for Alignment , author=. 2021 , eprint=

  107. [115]

    Aligning Generative Language Models with Human Values

    Liu, Ruibo and Zhang, Ge and Feng, Xinyu and Vosoughi, Soroush. Aligning Generative Language Models with Human Values. Findings of the Association for Computational Linguistics: NAACL 2022. 2022. doi:10.18653/v1/2022.findings-naacl.18

  108. [116]

    2023 , eprint=

    GPT-4 Technical Report , author=. 2023 , eprint=

  109. [117]

    2024 , eprint=

    TrustLLM: Trustworthiness in Large Language Models , author=. 2024 , eprint=

  110. [118]

    2023 , eprint=

    BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset , author=. 2023 , eprint=

  111. [119]

    Advances in Neural Information Processing Systems , volume=

    Process for adapting language models to society (palms) with values-targeted datasets , author=. Advances in Neural Information Processing Systems , volume=

  112. [120]

    arXiv preprint arXiv:2209.14375 , year=

    Improving alignment of dialogue agents via targeted human judgements , author=. arXiv preprint arXiv:2209.14375 , year=

  113. [121]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Valuenet: A new dataset for human value driven dialogue system , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  114. [122]

    Identifying the Human Values behind Arguments

    Kiesel, Johannes and Alshomary, Milad and Handke, Nicolas and Cai, Xiaoni and Wachsmuth, Henning and Stein, Benno. Identifying the Human Values behind Arguments. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 20...

  115. [123]

    M oral D ial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral Discussions

    Sun, Hao and Zhang, Zhexin and Mi, Fei and Wang, Yasheng and Liu, Wei and Cui, Jianwei and Wang, Bin and Liu, Qun and Huang, Minlie. M oral D ial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral Discussions. Proceedings of the 61st Annual Meeting of the Asso...

  116. [124]

    arXiv preprint arXiv:2308.05374 , year=

    Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment , author=. arXiv preprint arXiv:2308.05374 , year=

  117. [125]

    2023 , eprint=

    From Instructions to Intrinsic Human Values -- A Survey of Alignment Goals for Big Models , author=. 2023 , eprint=

  118. [126]

    2023 , eprint=

    AugGPT: Leveraging ChatGPT for Text Data Augmentation , author=. 2023 , eprint=

  119. [127]

    ChatGPT outperforms crowd workers for text-annotation tasks , volume=

    Gilardi, Fabrizio and Alizadeh, Meysam and Kubli, Maël , year=. ChatGPT outperforms crowd workers for text-annotation tasks , volume=. Proceedings of the National Academy of Sciences , publisher=. doi:10.1073/pnas.2305016120 , number=

  120. [128]

    Targeted Data Generation: Finding and Fixing Model Weaknesses

    He, Zexue and Ribeiro, Marco Tulio and Khani, Fereshte. Targeted Data Generation: Finding and Fixing Model Weaknesses. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.474

  121. [129]

    Advances in neural information processing systems , volume=

    Attention is all you need , author=. Advances in neural information processing systems , volume=

  122. [130]

    2019 , eprint=

    RoBERTa: A Robustly Optimized BERT Pretraining Approach , author=. 2019 , eprint=

  123. [131]

    Bosheng Ding and Chengwei Qin and Linlin Liu and Yew Ken Chia and Boyang Li and Shafiq Joty and Lidong Bing , title =

  124. [132]

    arXiv preprint arXiv:2303.16854 , year=

    Annollm: Making large language models to be better crowdsourced annotators , author=. arXiv preprint arXiv:2303.16854 , year=

  125. [133]

    2021 , eprint=

    Towards Zero-Label Language Learning , author=. 2021 , eprint=

  126. [134]

    Xuanli He and Islam Nassar and Jamie Kiros and Gholamreza Haffari and Mohammad Norouzi , title =. Trans. Assoc. Comput. Linguistics , volume =. 2022 , url =

  127. [135]

    Jiacheng Ye and Jiahui Gao and Qintong Li and Hang Xu and Jiangtao Feng and Zhiyong Wu and Tao Yu and Lingpeng Kong , title =

  128. [136]

    Yu Meng and Jiaxin Huang and Yu Zhang and Jiawei Han , title =. Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 , year =

  129. [137]

    QUILL : Query Intent with Large Language Models using Retrieval Augmentation and Multi-stage Distillation

    Srinivasan, Krishna and Raman, Karthik and Samanta, Anupam and Liao, Lingrui and Bertelli, Luca and Bendersky, Michael. QUILL : Query Intent with Large Language Models using Retrieval Augmentation and Multi-stage Distillation. Proceedings of the 2022 Conference on Empirical Me...

  130. [138]

    Query Rewriting in Retrieval-Augmented Large Language Models

    Ma, Xinbei and Gong, Yeyun and He, Pengcheng and Zhao, Hai and Duan, Nan. Query Rewriting in Retrieval-Augmented Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.322

  131. [139]

    CoRR , volume =

    Luiz Henrique Bonifacio and Hugo Queiroz Abonizio and Marzieh Fadaee and Rodrigo Frassetto Nogueira , title =. CoRR , volume =. 2022 , url =. 2202.05144 , timestamp =

  132. [140]

    Zhao and Ji Ma and Yi Luan and Jianmo Ni and Jing Lu and Anton Bakalov and Kelvin Guu and Keith B

    Zhuyun Dai and Vincent Y. Zhao and Ji Ma and Yi Luan and Jianmo Ni and Jing Lu and Anton Bakalov and Kelvin Guu and Keith B. Hall and Ming. Promptagator: Few-shot Dense Retrieval From 8 Examples , booktitle =. 2023 , url =

  133. [141]

    Tri Nguyen and Mir Rosenberg and Xia Song and Jianfeng Gao and Saurabh Tiwary and Rangan Majumder and Li Deng , title =. Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co-located with the 30th Annual Conference on Neural I...

  134. [142]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,

    Jon Saad. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,. 2023 , url =

  135. [143]

    Zhao and Yanping Huang and Andrew M

    Hyung Won Chung and Le Hou and Shayne Longpre and Barret Zoph and Yi Tay and William Fedus and Eric Li and Xuezhi Wang and Mostafa Dehghani and Siddhartha Brahma and Albert Webson and Shixiang Shane Gu and Zhuyun Dai and Mirac Suzgun and Xinyun Chen and Aakanksha Chowdhery and...

  136. [144]

    2023 , eprint=

    Zero-Shot Listwise Document Reranking with a Large Language Model , author=. 2023 , eprint=

  137. [145]

    2023 , eprint=

    Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents , author=. 2023 , eprint=

  138. [146]

    2023 , eprint=

    Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting , author=. 2023 , eprint=

  139. [147]

    , title =

    Raffel, Colin and Shazeer, Noam and Roberts, Adam and Lee, Katherine and Narang, Sharan and Matena, Michael and Zhou, Yanqi and Li, Wei and Liu, Peter J. , title =. J. Mach. Learn. Res. , month =. 2020 , issue_date =

  140. [148]

    Improving Passage Retrieval with Zero-Shot Question Generation

    Sachan, Devendra and Lewis, Mike and Joshi, Mandar and Aghajanyan, Armen and Yih, Wen-tau and Pineau, Joelle and Zettlemoyer, Luke. Improving Passage Retrieval with Zero-Shot Question Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Proce...

  141. [149]

    Proceedings of the 2019

    Sebastian Bruch and Xuanhui Wang and Michael Bendersky and Marc Najork , title =. Proceedings of the 2019. 2019 , url =. doi:10.1145/3341981.3344221 , timestamp =

  142. [150]

    Proceedings of the 22nd International Conference on Machine Learning , pages =

    Burges, Chris and Shaked, Tal and Renshaw, Erin and Lazier, Ari and Deeds, Matt and Hamilton, Nicole and Hullender, Greg , title =. Proceedings of the 22nd International Conference on Machine Learning , pages =. 2005 , isbn =. doi:10.1145/1102351.1102363 , abstract =

  143. [151]

    Proceedings of the 27th ACM International Conference on Information and Knowledge Management , pages =

    Wang, Xuanhui and Li, Cheng and Golbandi, Nadav and Bendersky, Michael and Najork, Marc , title =. Proceedings of the 27th ACM International Conference on Information and Knowledge Management , pages =. 2018 , isbn =. doi:10.1145/3269206.3271784 , abstract =

  144. [152]

    2023 , eprint=

    RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models , author=. 2023 , eprint=

  145. [153]

    2023 , eprint=

    RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! , author=. 2023 , eprint=

  146. [154]

    Questions Are All You Need to Train a Dense Passage Retriever

    Sachan, Devendra Singh and Lewis, Mike and Yogatama, Dani and Zettlemoyer, Luke and Pineau, Joelle and Zaheer, Manzil. Questions Are All You Need to Train a Dense Passage Retriever. Transactions of the Association for Computational Linguistics. 2023. doi:10.1162/tacl_a_00564

  147. [155]

    2023 , eprint=

    Mistral 7B , author=. 2023 , eprint=

  148. [156]

    arXiv preprint arXiv:2307.08303 , year=

    Soft prompt tuning for augmenting dense retrieval with large language models , author=. arXiv preprint arXiv:2307.08303 , year=

  149. [157]

    Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =

    Ferraretto, Fernando and Laitz, Thiago and Lotufo, Roberto and Nogueira, Rodrigo , title =. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =. 2023 , isbn =. doi:10.1145/3539618.3592067 , abstract =

  150. [158]

    Generating Datasets with Pretrained Language Models

    Schick, Timo and Sch. Generating Datasets with Pretrained Language Models. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.555

  151. [159]

    arXiv preprint arXiv:2301.01820 , year=

    InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval , author=. arXiv preprint arXiv:2301.01820 , year=

  152. [160]

    2023 , eprint=

    AugTriever: Unsupervised Dense Retrieval by Scalable Data AugAmentation , author=. 2023 , eprint=

  153. [161]

    LLM 4 V is: Explainable Visualization Recommendation using C hat GPT

    Wang, Lei and Zhang, Songheng and Wang, Yun and Lim, Ee-Peng and Wang, Yong. LLM 4 V is: Explainable Visualization Recommendation using C hat GPT. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2023. doi:10.18653/v1/2023...

  154. [162]

    2023 , eprint=

    Representation Learning with Large Language Models for Recommendation , author=. 2023 , eprint=

  155. [163]

    2024 , eprint=

    LLMRec: Large Language Models with Graph Augmentation for Recommendation , author=. 2024 , eprint=

  156. [164]

    2023 , eprint=

    Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach , author=. 2023 , eprint=

  157. [165]

    Proceedings of the 17th ACM Conference on Recommender Systems , pages =

    Mysore, Sheshera and Mccallum, Andrew and Zamani, Hamed , title =. Proceedings of the 17th ACM Conference on Recommender Systems , pages =. 2023 , isbn =. doi:10.1145/3604915.3608829 , abstract =

  158. [166]

    2023 , eprint=

    ONCE: Boosting Content-based Recommendation with Both Open- and Closed-source Large Language Models , author=. 2023 , eprint=

  159. [167]

    2022 , eprint=

    M6-Rec: Generative Pretrained Language Models are Open-Ended Recommender Systems , author=. 2022 , eprint=

  160. [168]

    2023 , eprint=

    Pre-train, Prompt and Recommendation: A Comprehensive Survey of Language Modelling Paradigm Adaptations in Recommender Systems , author=. 2023 , eprint=

  161. [169]

    2023 , eprint=

    Generative Recommendation: Towards Next-generation Recommender Paradigm , author=. 2023 , eprint=

  162. [170]

    Proceedings of the 17th ACM Conference on Recommender Systems , pages =

    Dai, Sunhao and Shao, Ninglu and Zhao, Haiyuan and Yu, Weijie and Si, Zihua and Xu, Chen and Sun, Zhongxiang and Zhang, Xiao and Xu, Jun , title =. Proceedings of the 17th ACM Conference on Recommender Systems , pages =. 2023 , isbn =. doi:10.1145/3604915.3610646 , abstract =

  163. [171]

    2023 , eprint=

    Towards Open-World Recommendation with Knowledge Augmentation from Large Language Models , author=. 2023 , eprint=

  164. [172]

    The Eleventh International Conference on Learning Representations,

    Jiahui Gao and Renjie Pi and Yong Lin and Hang Xu and Jiacheng Ye and Zhiyong Wu and Weizhong Zhang and Xiaodan Liang and Zhenguo Li and Lingpeng Kong , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =

  165. [173]

    Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , pages=

    Unleashing the Power of Large Language Models for Legal Applications , author=. Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , pages=

  166. [174]

    arXiv preprint arXiv:2303.09136 , year=

    A short survey of viewing large language models in legal aspect , author=. arXiv preprint arXiv:2303.09136 , year=

  167. [175]

    arXiv preprint arXiv:2312.03718 , year=

    Large Language Models in Law: A Survey , author=. arXiv preprint arXiv:2312.03718 , year=

  168. [176]

    Want To Reduce Labeling Cost? GPT -3 Can Help

    Wang, Shuohang and Liu, Yang and Xu, Yichong and Zhu, Chenguang and Zeng, Michael. Want To Reduce Labeling Cost? GPT -3 Can Help. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021. doi:10.18653/v1/2021.findings-emnlp.354

  169. [177]

    2024 , url=

    Fangyuan Xu and Weijia Shi and Eunsol Choi , booktitle=. 2024 , url=

  170. [178]

    I nherit S umm: A General, Versatile and Compact Summarizer by Distilling from GPT

    Xu, Yichong and Xu, Ruochen and Iter, Dan and Liu, Yang and Wang, Shuohang and Zhu, Chenguang and Zeng, Michael. I nherit S umm: A General, Versatile and Compact Summarizer by Distilling from GPT. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi...

  171. [179]

    2023 , eprint=

    Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing , author=. 2023 , eprint=

  172. [180]

    Data Augmentation for Radiology Report Simplification

    Yang, Ziyu and Cherian, Santhosh and Vucetic, Slobodan. Data Augmentation for Radiology Report Simplification. Findings of the Association for Computational Linguistics: EACL 2023. 2023. doi:10.18653/v1/2023.findings-eacl.144

  173. [181]

    2023 , eprint=

    Improving Small Language Models on PubMedQA via Generative Data Augmentation , author=. 2023 , eprint=

  174. [182]

    2023 , eprint=

    Neural Machine Translation Data Generation and Augmentation using ChatGPT , author=. 2023 , eprint=

  175. [183]

    UMASS \_ B io NLP at MEDIQA -Chat 2023: Can LLM s generate high-quality synthetic note-oriented doctor-patient conversations?

    Wang, Junda and Yao, Zonghai and Mitra, Avijit and Osebe, Samuel and Yang, Zhichao and Yu, Hong. UMASS \_ B io NLP at MEDIQA -Chat 2023: Can LLM s generate high-quality synthetic note-oriented doctor-patient conversations?. Proceedings of the 5th Clinical Natural Language Proc...

  176. [184]

    2023 , eprint=

    Tailoring Self-Rationalizers with Multi-Reward Distillation , author=. 2023 , eprint=

  177. [185]

    Symbolic Chain-of-Thought Distillation: Small Models Can Also `` Think '' Step-by-Step

    Li, Liunian Harold and Hessel, Jack and Yu, Youngjae and Ren, Xiang and Chang, Kai-Wei and Choi, Yejin. Symbolic Chain-of-Thought Distillation: Small Models Can Also `` Think '' Step-by-Step. Proceedings of the 61st Annual Meeting of the Association for Computational Linguisti...

  178. [186]

    2024 , eprint=

    Leveraging Large Language Models for NLG Evaluation: A Survey , author=. 2024 , eprint=

  179. [187]

    2023 , eprint=

    PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization , author=. 2023 , eprint=

  180. [188]

    The Twelfth International Conference on Learning Representations , year=

    Prometheus: Inducing Evaluation Capability in Language Models , author=. The Twelfth International Conference on Learning Representations , year=

  181. [189]

    INSTRUCTSCORE : Towards Explainable Text Generation Evaluation with Automatic Feedback

    Xu, Wenda and Wang, Danqing and Pan, Liangming and Song, Zhenqiao and Freitag, Markus and Wang, William and Li, Lei. INSTRUCTSCORE : Towards Explainable Text Generation Evaluation with Automatic Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Langu...

  182. [190]

    2023 , eprint=

    TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks , author=. 2023 , eprint=

  183. [191]

    The Twelfth International Conference on Learning Representations , year=

    Generative Judge for Evaluating Alignment , author=. The Twelfth International Conference on Learning Representations , year=

  184. [192]

    Proceedings of the 40th Annual Meeting on Association for Computational Linguistics , pages =

    Papineni, Kishore and Roukos, Salim and Ward, Todd and Zhu, Wei-Jing , title =. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics , pages =. 2002 , publisher =. doi:10.3115/1073083.1073135 , abstract =

  185. [193]

    ROUGE : A Package for Automatic Evaluation of Summaries

    Lin, Chin-Yew. ROUGE : A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out. 2004

  186. [194]

    2023 , eprint=

    Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias , author=. 2023 , eprint=

  187. [195]

    2023 , eprint=

    Magicoder: Source Code Is All You Need , author=. 2023 , eprint=

  188. [196]

    2024 , eprint=

    Genie: Achieving Human Parity in Content-Grounded Datasets Generation , author=. 2024 , eprint=

  189. [197]

    Jiazheng Li and Lin Gui and Yuxiang Zhou and David West and Cesare Aloisi and Yulan He , title =

  190. [198]

    2023 , eprint=

    LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention , author=. 2023 , eprint=

  191. [199]

    2023 , eprint=

    Code Llama: Open Foundation Models for Code , author=. 2023 , eprint=

  192. [200]

    2023 , eprint=

    Llama 2: Open Foundation and Fine-Tuned Chat Models , author=. 2023 , eprint=

  193. [201]

    Personalized Distillation: Empowering Open-Sourced

    Hailin Chen and Amrita Saha and Steven Hoi and Shafiq Joty , booktitle=. Personalized Distillation: Empowering Open-Sourced. 2023 , url=

  194. [202]

    2023 , eprint=

    MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning , author=. 2023 , eprint=

  195. [203]

    2023 , eprint=

    Instruction Fusion: Advancing Prompt Evolution through Hybridization , author=. 2023 , eprint=

  196. [204]

    2024 , eprint=

    WaveCoder: Widespread And Versatile Enhanced Instruction Tuning with Refined Data Generation , author=. 2024 , eprint=

  197. [205]

    6th International Conference on Learning Representations,

    Ozan Sener and Silvio Savarese , title =. 6th International Conference on Learning Representations,. 2018 , url =

  198. [206]

    LLM-Assisted Code Cleaning For Training Accurate Code Generators , journal =

    Naman Jain and Tianjun Zhang and Wei. LLM-Assisted Code Cleaning For Training Accurate Code Generators , journal =. 2023 , url =. doi:10.48550/ARXIV.2311.14904 , eprinttype =. 2311.14904 , timestamp =

  199. [207]

    Distilled

    Chia. Distilled. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2308.14731 , eprinttype =. 2308.14731 , timestamp =

  200. [208]

    CoRR , volume =

    Weidong Guo and Jiuding Yang and Kaitong Yang and Xiangyang Li and Zhuwei Rao and Yu Xu and Di Niu , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2312.15692 , eprinttype =. 2312.15692 , timestamp =

  201. [209]

    2022 , eprint=

    Limitations of Language Models in Arithmetic and Symbolic Induction , author=. 2022 , eprint=

  202. [210]

    2023 , eprint=

    Pitfalls in Language Models for Code Intelligence: A Taxonomy and Survey , author=. 2023 , eprint=

  203. [211]

    2023 , eprint=

    Language models are weak learners , author=. 2023 , eprint=

  204. [212]

    2023 , eprint=

    Toolformer: Language Models Can Teach Themselves to Use Tools , author=. 2023 , eprint=

  205. [213]

    2022 , eprint=

    OPT: Open Pre-trained Transformer Language Models , author=. 2022 , eprint=

  206. [214]

    Advances in neural information processing systems , volume=

    Language models are few-shot learners , author=. Advances in neural information processing systems , volume=

  207. [215]

    2023 , eprint=

    ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs , author=. 2023 , eprint=

  208. [216]

    2023 , eprint=

    CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets , author=. 2023 , eprint=

  209. [217]

    2024 , eprint=

    MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning , author=. 2024 , eprint=

  210. [218]

    2024 , eprint=

    Small LLMs Are Weak Tool Learners: A Multi-LLM Agent , author=. 2024 , eprint=

  211. [219]

    2024 , eprint=

    EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction , author=. 2024 , eprint=

  212. [220]

    2023 , eprint=

    ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases , author=. 2023 , eprint=

  213. [221]

    2023 , eprint=

    HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face , author=. 2023 , eprint=

  214. [222]

    2023 , eprint=

    Gorilla: Large Language Model Connected with Massive APIs , author=. 2023 , eprint=

  215. [223]

    2022 , eprint=

    WebGPT: Browser-assisted question-answering with human feedback , author=. 2022 , eprint=

  216. [224]

    W eb CPM : Interactive Web Search for C hinese Long-form Question Answering

    Qin, Yujia and Cai, Zihan and Jin, Dian and Yan, Lan and Liang, Shihao and Zhu, Kunlun and Lin, Yankai and Han, Xu and Ding, Ning and Wang, Huadong and Xie, Ruobing and Qi, Fanchao and Liu, Zhiyuan and Sun, Maosong and Zhou, Jie. W eb CPM : Interactive Web Search for C hinese ...

  217. [225]

    2024 , eprint=

    ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings , author=. 2024 , eprint=

  218. [226]

    CREATOR : Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models

    Qian, Cheng and Han, Chi and Fung, Yi and Qin, Yujia and Liu, Zhiyuan and Ji, Heng. CREATOR : Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/...

  219. [227]

    2023 , eprint=

    RestGPT: Connecting Large Language Models with Real-World RESTful APIs , author=. 2023 , eprint=

  220. [228]

    2022 , eprint=

    TALM: Tool Augmented Language Models , author=. 2022 , eprint=

  221. [229]

    2023 , eprint=

    Large Language Models as Tool Makers , author=. 2023 , eprint=

  222. [230]

    2023 , eprint=

    Confucius: Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum , author=. 2023 , eprint=

  223. [231]

    2023 , eprint=

    TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs , author=. 2023 , eprint=

  224. [232]

    2023 , eprint=

    Augmented Language Models: a Survey , author=. 2023 , eprint=

  225. [233]

    arXiv preprint arXiv:2311.11608 , year=

    Taiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical Tasks , author=. arXiv preprint arXiv:2311.11608 , year=

  226. [234]

    arXiv preprint arXiv:2311.06025 , year=

    ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences , author=. arXiv preprint arXiv:2311.06025 , year=

  227. [235]

    H uatuo GPT , Towards Taming Language Model to Be a Doctor

    Zhang, Hongbo and Chen, Junying and Jiang, Feng and Yu, Fei and Chen, Zhihong and Chen, Guiming and Li, Jianquan and Wu, Xiangbo and Zhiyi, Zhang and Xiao, Qingying and Wan, Xiang and Wang, Benyou and Li, Haizhou. H uatuo GPT , Towards Taming Language Model to Be a Doctor. Fin...

  228. [236]

    arXiv preprint arXiv:2304.01097 , year=

    Doctorglm: Fine-tuning your chinese doctor is not a herculean task , author=. arXiv preprint arXiv:2304.01097 , year=

  229. [237]

    arXiv preprint arXiv:2310.14558 , year=

    Alpacare: Instruction-tuned large language models for medical application , author=. arXiv preprint arXiv:2310.14558 , year=

  230. [238]

    arXiv preprint arXiv:2308.14346 , year=

    Disc-medllm: Bridging general large language models and real-world medical consultation , author=. arXiv preprint arXiv:2308.14346 , year=

  231. [239]

    arXiv preprint arXiv:2305.10415 , volume=

    Pmc-llama: Towards building open-source language models for medicine , author=. arXiv preprint arXiv:2305.10415 , volume=

  232. [240]

    Cureus , volume=

    ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge , author=. Cureus , volume=. 2023 , publisher=

  233. [241]

    arXiv preprint arXiv:2310.14151 , year=

    PromptCBLUE: A Chinese Prompt Tuning Benchmark for the Medical Domain , author=. arXiv preprint arXiv:2310.14151 , year=

  234. [242]

    arXiv preprint arXiv:2308.08833 , year=

    Cmb: A comprehensive medical benchmark in chinese , author=. arXiv preprint arXiv:2308.08833 , year=

  235. [243]

    International Conference on Machine Learning , pages=

    Language models as zero-shot planners: Extracting actionable knowledge for embodied agents , author=. International Conference on Machine Learning , pages=. 2022 , organization=

  236. [244]

    2022 , eprint=

    ProgPrompt: Generating Situated Robot Task Plans using Large Language Models , author=. 2022 , eprint=

  237. [245]

    2023 , eprint=

    Least-to-Most Prompting Enables Complex Reasoning in Large Language Models , author=. 2023 , eprint=

  238. [246]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    On grounded planning for embodied tasks with language models , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  239. [247]

    Thirty-seventh Conference on Neural Information Processing Systems , year=

    On the Planning Abilities of Large Language Models - A Critical Investigation , author=. Thirty-seventh Conference on Neural Information Processing Systems , year=

  240. [248]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    Llm-planner: Few-shot grounded planning for embodied agents with large language models , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  241. [249]

    arXiv preprint arXiv:2302.01560 , year=

    Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents , author=. arXiv preprint arXiv:2302.01560 , year=

  242. [250]

    arXiv preprint arXiv:2305.10601 , year=

    Tree of thoughts: Deliberate problem solving with large language models , author=. arXiv preprint arXiv:2305.10601 , year=

  243. [251]

    arXiv preprint arXiv:2304.11477 , year=

    Llm+ p: Empowering large language models with optimal planning proficiency , author=. arXiv preprint arXiv:2304.11477 , year=

  244. [252]

    arXiv preprint arXiv:2305.14992 , year=

    Reasoning with language model is planning with world model , author=. arXiv preprint arXiv:2305.14992 , year=

  245. [253]

    arXiv preprint arXiv:2310.08582 , year=

    Tree-Planner: Efficient Close-loop Task Planning with Large Language Models , author=. arXiv preprint arXiv:2310.08582 , year=

  246. [254]

    arXiv preprint arXiv:2308.09442 , year=

    Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine , author=. arXiv preprint arXiv:2308.09442 , year=

  247. [255]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    JEC-QA: a legal-domain question answering dataset , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  248. [256]

    arXiv preprint arXiv:2306.16092 , year=

    Chatlaw: Open-source legal large language model with integrated external knowledge bases , author=. arXiv preprint arXiv:2306.16092 , year=

  249. [257]

    arXiv preprint arXiv:2309.11325 , year=

    Disc-lawllm: Fine-tuning large language models for intelligent legal services , author=. arXiv preprint arXiv:2309.11325 , year=

  250. [258]

    Wu, Shiguang and Liu, Zhongkun and Zhang, Zhen and Chen, Zheng and Deng, Wentao and Zhang, Wenhao and Yang, Jiyuan and Yao, Zhitao and Lyu, Yougang and Xin, Xin and Gao, Shen and Ren, Pengjie and Ren, Zhaochun and Chen, Zhumin , year = 2023, journal =

  251. [259]

    2023 , eprint=

    FireAct: Toward Language Agent Fine-tuning , author=. 2023 , eprint=

  252. [260]

    2023 , eprint=

    AgentTuning: Enabling Generalized Agent Abilities for LLMs , author=. 2023 , eprint=

  253. [261]

    2023 , eprint=

    Lumos: Learning Agents with Unified Data, Modular Design, and Open-Source LLMs , author=. 2023 , eprint=

  254. [262]

    2024 , eprint=

    AUTOACT: Automatic Agent Learning from Scratch via Self-Planning , author=. 2024 , eprint=

  255. [263]

    2023 , eprint=

    TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems , author=. 2023 , eprint=

  256. [264]

    2023 , eprint=

    Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld , author=. 2023 , eprint=

  257. [265]

    2023 , eprint=

    GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , author=. 2023 , eprint=

  258. [266]

    NeurIPS , year =

    Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae , title =. NeurIPS , year =

  259. [267]

    2023 , eprint=

    Improved Baselines with Visual Instruction Tuning , author=. 2023 , eprint=

  260. [268]

    2023 , eprint=

    GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , author=. 2023 , eprint=

  261. [269]

    2023 , eprint=

    SVIT: Scaling up Visual Instruction Tuning , author=. 2023 , eprint=

  262. [270]

    2023 , eprint=

    To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning , author=. 2023 , eprint=

  263. [271]

    2023 , eprint=

    Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic , author=. 2023 , eprint=

  264. [272]

    Proceedings of the IEEE international conference on computer vision , pages=

    Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models , author=. Proceedings of the IEEE international conference on computer vision , pages=

  265. [273]

    2023 , eprint=

    Localized Symbolic Knowledge Distillation for Visual Commonsense Models , author=. 2023 , eprint=

  266. [274]

    2023 , eprint=

    LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding , author=. 2023 , eprint=

  267. [275]

    2023 , url=

    GPT-4V(ision) System Card , author=. 2023 , url=

  268. [276]

    arXiv preprint arXiv:2306.09093 , year=

    Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration , author=. arXiv preprint arXiv:2306.09093 , year=

  269. [277]

    2023 , eprint=

    MIMIC-IT: Multi-Modal In-Context Instruction Tuning , author=. 2023 , eprint=

  270. [278]

    2023 , eprint=

    ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst , author=. 2023 , eprint=

  271. [279]

    D et GPT : Detect What You Need via Reasoning

    Pi, Renjie and Gao, Jiahui and Diao, Shizhe and Pan, Rui and Dong, Hanze and Zhang, Jipeng and Yao, Lewei and Han, Jianhua and Xu, Hang and Kong, Lingpeng and Zhang, Tong. D et GPT : Detect What You Need via Reasoning. Proceedings of the 2023 Conference on Empirical Methods in...

  272. [280]

    2023 , eprint=

    ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning , author=. 2023 , eprint=

  273. [281]

    2023 , eprint=

    Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning , author=. 2023 , eprint=

  274. [282]

    2023 , eprint=

    NExT-GPT: Any-to-Any Multimodal LLM , author=. 2023 , eprint=

  275. [283]

    2023 , eprint=

    Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare , author=. 2023 , eprint=

  276. [284]

    2023 , eprint=

    Valley: Video Assistant with Large Language model Enhanced abilitY , author=. 2023 , eprint=

  277. [285]

    2023 , eprint=

    ILuvUI: Instruction-tuned LangUage-Vision modeling of UIs from Machine Conversations , author=. 2023 , eprint=

  278. [286]

    2023 , eprint=

    StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data , author=. 2023 , eprint=

  279. [287]

    2023 , eprint=

    PointLLM: Empowering Large Language Models to Understand Point Clouds , author=. 2023 , eprint=

  280. [288]

    2023 , url=

    Chunting Zhou and Pengfei Liu and Puxin Xu and Srini Iyer and Jiao Sun and Yuning Mao and Xuezhe Ma and Avia Efrat and Ping Yu and LILI YU and Susan Zhang and Gargi Ghosh and Mike Lewis and Luke Zettlemoyer and Omer Levy , booktitle=. 2023 , url=

  281. [289]

    2023 , eprint=

    \#InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models , author=. 2023 , eprint=

  282. [290]

    2023 , eprint=

    AlpaGasus: Training A Better Alpaca with Fewer Data , author=. 2023 , eprint=

  283. [291]

    ArXiv , year=

    From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning , author=. ArXiv , year=

  284. [292]

    2023 , eprint=

    Reinforced Self-Training (ReST) for Language Modeling , author=. 2023 , eprint=

  285. [293]

    2024 , eprint=

    Self-Rewarding Language Models , author=. 2024 , eprint=

  286. [294]

    2023 , eprint=

    What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning , author=. 2023 , eprint=

  287. [295]

    2023 , eprint=

    Rethinking the Instruction Quality: LIFT is What You Need , author=. 2023 , eprint=

  288. [296]

    arXiv preprint arXiv:1503.02531 , year=

    Distilling the knowledge in a neural network , author=. arXiv preprint arXiv:1503.02531 , year=

  289. [297]

    2024 , url=

    Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning , author=. 2024 , url=

  290. [298]

    2023 , eprint=

    MoDS: Model-oriented Data Selection for Instruction Tuning , author=. 2023 , eprint=

  291. [299]

    2023 , eprint=

    A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment , author=. 2023 , eprint=

  292. [300]

    2023 , eprint=

    ExpertPrompting: Instructing Large Language Models to be Distinguished Experts , author=. 2023 , eprint=

Pith tools

Reviewed May 17, 2026 · model on record in the stance chip above.