MultiSynt/MT supplies 4.8 trillion translated tokens in 36 languages from 100B English tokens, letting LLMs match native-data baselines with 72% fewer tokens and beat them by 15% at equal budget.
Advances in Neural Information Processing Systems , volume =
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
PRISM weights target examples by model preference to build an improved direction for influence-based data selection in LLM fine-tuning.
Interpreting harmful Discord messages requires integrating external knowledge and extended context, not just local message-level classification; LLMs leverage local context better than humans but still fail on coded language and community-specific references.
A monograph-length survey claiming deep learning theory can be told as one narrative, from approximation guarantees to emergence.
citing papers explorer
-
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
MultiSynt/MT supplies 4.8 trillion translated tokens in 36 languages from 100B English tokens, letting LLMs match native-data baselines with 72% fewer tokens and beat them by 15% at equal budget.
-
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
PRISM weights target examples by model preference to build an improved direction for influence-based data selection in LLM fine-tuning.
-
Understanding Interpretation Difficulty in Harmful Online Communication: Insights from Cybercrime Communities
Interpreting harmful Discord messages requires integrating external knowledge and extended context, not just local message-level classification; LLMs leverage local context better than humans but still fail on coded language and community-specific references.
-
From Approximation to Emergence: A Theory of Deep Learning
A monograph-length survey claiming deep learning theory can be told as one narrative, from approximation guarantees to emergence.