Introduces AMALIA-VL, the first open-source instruction-tuned LVLM for European Portuguese, using a high-resolution vision encoder, pt-PT language model, learned connector, and three-stage training on a custom data mix.
2409.16235 , archiveprefix =
4 Pith papers cite this work, alongside 3 external citations. Polarity classification is still indexing.
years
2026 4representative citing papers
A translated-English German corpus (725B tokens) produced higher point estimates on German HellaSwag and ARC-C than native German web corpora in matched 12B-token pretraining runs, though the differences are not statistically robust.
Merging any combination of monolingual pre-trained models leads to performance collapse due to interference, indicating that merging flexibility from fine-tuning does not extend to pre-training.
Combines GRPO with teacher-guided on-policy distillation and introduces LongBlocks dataset to yield more stable long-context reasoning than either method alone.
citing papers explorer
-
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
Introduces AMALIA-VL, the first open-source instruction-tuned LVLM for European Portuguese, using a high-resolution vision encoder, pt-PT language model, learned connector, and three-stage training on a custom data mix.
-
KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
A translated-English German corpus (725B tokens) produced higher point estimates on German HellaSwag and ARC-C than native German web corpora in matched 12B-token pretraining runs, though the differences are not statistically robust.
-
On the Limits of Model Merging for Multilinguality in Pre-Training
Merging any combination of monolingual pre-trained models leads to performance collapse due to interference, indicating that merging flexibility from fine-tuning does not extend to pre-training.
-
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
Combines GRPO with teacher-guided on-policy distillation and introduces LongBlocks dataset to yield more stable long-context reasoning than either method alone.