Pith. sign in

REVIEW 37 cited by

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14818 v1 pith:F5KIEAQK submitted 2025-01-20 cs.CV cs.AIcs.LG

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

classification cs.CV cs.AIcs.LG
keywords modelsdatafrontieropen-sourcepost-trainingstrategyvlmsbuilding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most open-source models only publish their final model weights, leaving the critical details of data strategies and implementation largely opaque. In this work, we address VLM post-training from a data-centric perspective, showing the key role of data strategy in developing frontier VLMs. By studying and building our post-training data strategy from scratch, we share detailed insights into the development processes, aiming to benefit the development of competitive models for the open-source community. Our introduced data strategy, together with training recipes and model design, leads to a family of performant VLMs named Eagle2. Specifically, Eagle2-9B achieves state-of-the-art results across various multimodal benchmarks, matching certain competitive models with up to 70B parameters.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 37 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

    cs.CV 2026-07 conditional novelty 7.0

    A taxonomy of eight foundation-model priors organizes HOI reconstruction, generation, and embodied transfer, mapping what knowledge large models inject and where.

  2. AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

    cs.CV 2026-06 unverdicted novelty 7.0

    AMALIA-VL introduces the first open-source instruction-tuned LVLM natively optimized for European Portuguese via vision-language alignment, instruction tuning, preference optimization, and a pt-PT-centric data mix.

  3. Argus-Retriever: Vision-LLM Late-Interaction Retrieval with Region-Aware Query-Conditioned MoE for Visual Document Retrieval

    cs.IR 2026-06 unverdicted novelty 7.0

    Argus achieves the highest reported NDCG scores among open late-interaction models on ViDoRe V1 and combined V1+V2 by introducing query-dependent document representations via a region-aware MoE on Qwen3.5-VL, trained ...

  4. Being-H0.7: A Latent World-Action Model from Egocentric Videos

    cs.RO 2026-04 unverdicted novelty 7.0

    Being-H0.7 adds future-aware latent reasoning to direct VLA policies via dual-branch alignment on latent queries, matching world-model benefits at VLA efficiency.

  5. EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training

    cs.CV 2026-04 unverdicted novelty 7.0

    EmbodiedMidtrain mid-trains VLMs on curated VLA-aligned data subsets to improve downstream performance on robot manipulation benchmarks.

  6. High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

    cs.CV 2025-07 conditional novelty 7.0

    MGPO elicits grounding in LMMs via multi-turn RL with binary rewards, yielding 5.4% and 5.2% gains on MME-Realworld and V* Bench and surpassing GPT-4o on the latter after training on 21K samples.

  7. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  8. AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

    cs.CV 2026-06 unverdicted novelty 6.0

    AMALIA-VL is the first open-source LVLM natively optimized for European Portuguese via three-stage training on a pt-PT-centric data mix combining curated, translated, and novel datasets.

  9. AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces AMALIA-VL, the first open-source instruction-tuned LVLM for European Portuguese, using a high-resolution vision encoder, pt-PT language model, learned connector, and three-stage training on a custom data mix.

  10. Contrastive Action-Image Pre-training for Visuomotor Control

    cs.RO 2026-06 unverdicted novelty 6.0

    CAIP learns action-aligned visual representations via contrastive pre-training on human hand keypoints from egocentric video, outperforming DINOv2, SigLIP, MVP, and R3M with >30% gains on real dexterous manipulation tasks.

  11. Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

    cs.RO 2026-06 conditional novelty 6.0

    A full learning stack—10K-hour UMI data, a flow-matching VLA, preference-optimization RL, and asynchronous deployment—reports SOTA RoboTwin results and cross-embodiment transfer to four real robots.

  12. RhinoVLA Technical Report

    cs.RO 2026-06 conditional novelty 6.0

    RhinoVLA matches π0.5-scale VLA task performance while reaching 11.69 Hz end-to-end inference on the Huixi R1 edge SoC via token-efficient Qwen3-VL, a 72D unified action interface, and hardware co-optimization.

  13. LARA: Latent Action Representation Alignment for Vision-Language-Action Models

    cs.CV 2026-06 unverdicted novelty 6.0

    LARA jointly optimizes LAM and VLA models via representation alignment to improve robotic manipulation performance using human videos.

  14. Cambrian-P: Pose-Grounded Video Understanding

    cs.CV 2026-05 conditional novelty 6.0

    Adding per-frame camera-pose supervision to a video MLLM improves spatial and general video question answering by 2–6% and yields SOTA streaming pose estimates on ScanNet.

  15. Cambrian-P: Pose-Grounded Video Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.

  16. $M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

    cs.RO 2026-04 unverdicted novelty 6.0

    M²-VLA shows that generalized VLMs can serve as direct backbones for robotic manipulation by selectively extracting task-critical features via Mixture of Layers and adding Meta Skill Modules for efficient trajectory learning.

  17. $M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

    cs.RO 2026-04 conditional novelty 6.0

    Freezing a VLM backbone and routing its layers through Mixture-of-Layers plus a Meta-Skill memory yields higher success and stronger zero-shot generalization than fine-tuned VLAs on LIBERO and real robots.

  18. Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

    cs.RO 2026-04 unverdicted novelty 6.0

    State-of-the-art vision-language-action models catastrophically fail dynamic embodied reasoning due to lexical-kinematic shortcuts, behavioral inertia, and semantic feature collapse caused by architectural bottlenecks...

  19. MLLM-as-a-Judge Exhibits Model Preference Bias

    cs.CV 2026-04 unverdicted novelty 6.0

    MLLMs show self-preference bias and family-level mutual bias when judging captions; Philautia-Eval quantifies it and Pomms ensemble reduces it.

  20. Adaptive Action Chunking at Inference-time for Vision-Language-Action Models

    cs.RO 2026-04 unverdicted novelty 6.0

    Adaptive Action Chunking uses action entropy to dynamically adjust chunk sizes in VLA models, improving performance on simulated and real robotic manipulation tasks.

  21. RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

    cs.CV 2025-12 conditional novelty 6.0

    RefBench-PRO organizes REC into attribute, position, interaction, relation, commonsense, and reject tasks; no tested MLLM exceeds 72%, and Ref-R1 raises Qwen2.5-VL-7B from 57.6 to 69.4 on it.

  22. FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

    cs.CV 2025-10 conditional novelty 6.0

    A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.

  23. FLARE: Robot Learning with Implicit World Modeling

    cs.RO 2025-05 unverdicted novelty 6.0

    FLARE integrates predictive latent world modeling into diffusion transformer policies for robots, delivering up to 26% gains on multitask manipulation benchmarks and enabling co-training with action-free human videos.

  24. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

    cs.CV 2025-04 conditional novelty 6.0

    InternVL3-78B sets a new open-source SOTA of 72.2 on MMMU via native joint multimodal pre-training, V2PE, MPO, and test-time scaling while remaining competitive with proprietary models.

  25. SmolVLM: Redefining small and efficient multimodal models

    cs.AI 2025-04 unverdicted novelty 6.0

    SmolVLM-256M outperforms a 300-times larger model using under 1 GB GPU memory, while the 2.2B version matches state-of-the-art VLMs at half the memory cost.

  26. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

    cs.RO 2025-03 unverdicted novelty 6.0

    GR00T N1 is a new open VLA foundation model for humanoid robots that outperforms imitation learning baselines in simulation and shows strong performance on real-world bimanual manipulation tasks.

  27. SigLIP-HD by Fine-to-Coarse Supervision

    cs.CV 2026-07 conditional novelty 5.5

    Fine-to-coarse L1 supervision lets a standard-resolution SigLIP 2 encoder produce better visual tokens for MLLMs without higher-resolution inference.

  28. Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

    cs.CV 2026-07 accept novelty 5.0

    Foundation-model HOI work is organized into eight geometric, semantic, and visual sub-priors that enter six reconstruction/generation tasks and three robot-transfer routes.

  29. LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

    cs.RO 2026-07 conditional novelty 5.0

    LeapBot-WA shows robot policies can be trained with latent world-model predictions instead of pixel video generation, hitting state-of-the-art for predictive action models and staying competitive with generative WAMs.

  30. RhinoVLA Technical Report

    cs.RO 2026-06 conditional novelty 5.0

    RhinoVLA runs a ~2.5B-parameter robot policy end-to-end at 11.69 Hz on the Huixi R1 edge SoC with LIBERO accuracy within a few points of π0.5 and mixed real-robot results.

  31. RhinoVLA Technical Report

    cs.RO 2026-06 unverdicted novelty 5.0

    RhinoVLA uses a token-efficient Qwen3-VL backbone, continuous Action Expert, and unified cross-robot interface to match π0.5 performance while hitting 11.69 Hz on Huixi R1 edge SoC.

  32. LARA: Latent Action Representation Alignment for Vision-Language-Action Models

    cs.CV 2026-06 unverdicted novelty 5.0

    LARA jointly optimizes LAM and VLA models via representation alignment, reporting average gains of ~10%, ~5%, and ~15% on simulation and real robotic manipulation tasks.

  33. Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

    cs.RO 2026-05 unverdicted novelty 5.0

    Gaze2Act conditions VLA policies on mapped human gaze for precise object and interaction specification, reporting SOTA intent accuracy and success across 16 real-robot tasks on a Unitree G1 humanoid.

  34. ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

    cs.CV 2025-07 unverdicted novelty 5.0

    ThinkAct introduces reinforced visual latent planning in a dual VLA system to enable better long-horizon reasoning and adaptation for embodied tasks.

  35. Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning

    cs.CV 2026-06 unverdicted novelty 4.0

    ConsistRoll enforces cross-view consistency during RLVR training for MLLMs by joint rewards on grouped original and augmented views, yielding robustness gains on math, general, and hallucination benchmarks.

  36. PLaMo 2.1-VL Technical Report

    cs.CV 2026-04 unverdicted novelty 4.0

    PLaMo 2.1-VL reports 61.5 ROUGE-L on JA-VG-VQA-500, 85.2% on Japanese Ref-L4, 53.9% zero-shot factory accuracy, and raises anomaly detection F1 from 39.7 to 64.9 after fine-tuning.

  37. RhinoVLA Technical Report

    cs.RO 2026-06 unverdicted novelty 3.0

    RhinoVLA cuts VLM tokens with a Qwen3-VL backbone and continuous action expert, adds a unified cross-robot interface, and reaches real-time 11.69 Hz on Huixi R1 while matching π0.5 downstream performance.