Pith. sign in

REVIEW 18 cited by

UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09151 v1 pith:QGYRHMFW submitted 2023-04-18 cs.CL

classification cs.CL
keywords samplinglanguagelanguagesmultilingualunimaxmodelacrosscorpus
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pretrained multilingual large language models have typically used heuristic temperature-based sampling to balance between different languages. However previous work has not systematically evaluated the efficacy of different pretraining language distributions across model scales. In this paper, we propose a new sampling method, UniMax, that delivers more uniform coverage of head languages while mitigating overfitting on tail languages by explicitly capping the number of repeats over each language's corpus. We perform an extensive series of ablations testing a range of sampling strategies on a suite of multilingual benchmarks, while varying model scale. We find that UniMax outperforms standard temperature-based sampling, and the benefits persist as scale increases. As part of our contribution, we release: (i) an improved and refreshed mC4 multilingual corpus consisting of 29 trillion characters across 107 languages, and (ii) a suite of pretrained umT5 model checkpoints trained with UniMax sampling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

    cs.CL 2026-08 conditional novelty 7.0 of 10

    MameLoshnLM, trained by continuing pretraining Llama 3.1 8B on a curated Yiddish corpus, outperforms similar-scale open models on a new multi-task Yiddish benchmark and better retains Yiddish-specific loshn-koydesh vo...

  2. FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A DiT-based portrait animation model transfers implicit facial expressions to one or more characters using a masked cross-attention mechanism, supported by a new multi-face dataset and benchmark.

  3. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A hybrid NoPE visual diffusion backbone with HeteroP scaling yields ~7× pretraining compute efficiency versus matched full attention and near-stable zero-shot 5s→30s video extrapolation.

  4. StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Coupling explicit state prediction with MoT-style video generation raises mechanics fidelity of Street Fighter 3 rollouts by about 18.6% over stateless game world models.

  5. Wonder: Video World Model Done Better

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Wonder generates minute-scale, real-time camera-controllable video worlds from a single image or video at 16 FPS, using a rendered coordinate-field control signal, sparse full-fidelity memory, and stage-specialized di...

  6. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  7. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

    cs.CV 2026-07 conditional novelty 6.0 of 10

    HOMIE unifies inter- and intra-subject video personalization by injecting MLLM-derived relational features into DiT self-attention (GMG) and tagging tokens with modality/reference embeddings (MRE), reporting SOTA on a...

  8. Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.

  9. OpenCoF: Learning to Reason Through Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fine-tuning a video generator on a new 17K reasoning-video dataset improves Chain-of-Frame reasoning, and adding learnable visual/textual reasoning tokens yields further gains on external benchmarks.

  10. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  11. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.

  12. UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.

  13. Yume: An Interactive World Generation Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-based video model generates extendable, keyboard-controlled walkthroughs from a single input image, using quantized camera actions as text prompts.

  14. QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    QUENCH introduces a generation-based quiz benchmark with masked entities and rationales, and documents a consistent Indic-versus-non-Indic performance gap across seven LLMs.

  15. JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A dual-branch diffusion transformer with joint video-audio self-attention and a keypoint-based mouth-area loss reports top lip-sync and speech metrics on two benchmarks.

  16. MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Multi-domain RLVR data mixing, guided by a quadratic surrogate fitted to 11 pilot runs, improves a Qwen2-VL-2B model's out-of-distribution accuracy by about 5 points over uniform mixing.

  17. IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment

    cs.AI 2025-05 reject novelty 5.0 of 10

    IDEAL tunes SFT data mixture proportions per domain with influence-function gradients, claiming about 7% average benchmark improvement over uniform mixing.

  18. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

Pith tools