Magpie synthesizes 300K high-quality alignment instructions from Llama-3-Instruct via auto-regressive prompting on partial templates, enabling fine-tuned models to match official instruct performance on AlpacaEval, ArenaHard, and WildBench.
Pandora's white-box: Increased training data leakage in open llms
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 3verdicts
UNVERDICTED 3representative citing papers
LLMSurgeon recovers pretraining domain mixtures from LLM-generated text by estimating a calibrated soft confusion matrix and solving a constrained inverse problem under the label-shift assumption.
Set-level data entropy estimators show linear correlation with LLM memorization scores, forming the Entropy-Memorization Linearity.
citing papers explorer
-
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Magpie synthesizes 300K high-quality alignment instructions from Llama-3-Instruct via auto-regressive prompting on partial templates, enabling fine-tuned models to match official instruct performance on AlpacaEval, ArenaHard, and WildBench.
-
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
LLMSurgeon recovers pretraining domain mixtures from LLM-generated text by estimating a calibrated soft confusion matrix and solving a constrained inverse problem under the label-shift assumption.
-
Data Compressibility Quantifies LLM Memorization
Set-level data entropy estimators show linear correlation with LLM memorization scores, forming the Entropy-Memorization Linearity.