Self-generated QA supervision for language models is fragile due to non-uniform question selection and instruction compliance during answering, with mitigations that reduce compliance from 88% to 13%.
InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Min- ing (KDD 2024), pages 5351–5362
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
EVE is the first open-source end-to-end system with a domain-adapted 24B LLM that outperforms peers on new Earth Intelligence benchmarks while adding RAG and hallucination detection in a production deployment.
citing papers explorer
-
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
Self-generated QA supervision for language models is fragile due to non-uniform question selection and instruction compliance during answering, with mitigations that reduce compliance from 88% to 13%.
-
EVE: A Domain-Specific LLM Framework for Earth Intelligence
EVE is the first open-source end-to-end system with a domain-adapted 24B LLM that outperforms peers on new Earth Intelligence benchmarks while adding RAG and hallucination detection in a production deployment.