Self-generated QA supervision for language models is fragile due to non-uniform question selection and instruction compliance during answering, with mitigations that reduce compliance from 88% to 13%.
arXiv preprint arXiv:2505.14212 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
DoRA generates synthetic RAG training and evaluation data from 40 defense documents, halving hallucination rates in a LoRA-adapted Llama3.1-8B compared to 8 baselines.
citing papers explorer
-
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
Self-generated QA supervision for language models is fragile due to non-uniform question selection and instruction compliance during answering, with mitigations that reduce compliance from 88% to 13%.
-
A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents
DoRA generates synthetic RAG training and evaluation data from 40 defense documents, halving hallucination rates in a LoRA-adapted Llama3.1-8B compared to 8 baselines.