REVIEW 3 major objections 5 minor 1 cited by
A systematic review of 435 papers claims the field of foundation-model robotics divides into five phases and can be classified along six criteria.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 05:26 UTC pith:K27DUO3T
load-bearing objection A useful, carefully organized survey of 435 FM-robotics papers, with a six-criteria taxonomy and five-phase periodization; the main soft spots are corpus transparency and a verifiable citation error. the 3 major comments →
Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that robotic foundation-model research has evolved through five distinct phases: (1) integration of native NLP and computer-vision models (2018-2021); (2) grounded planning with vision-language representations (2021-2022); (3) embodied vision-language-action policies (2022-2023); (4) memory, autonomous task composition, and web-to-robot transfer (2023-2024); and (5) multisensory generalization and real-world deployment (2024-present). Across these phases, the paper classifies 435 works along six criteria — FM type (LLM, VFM, VLM, VLA), neural-network architecture, learning paradigm, learning stage, robotic task, and application domain — and provides per-criterion compara
What carries the argument
The carrying mechanism is the taxonomy itself: six analysis criteria plus the five-phase chronology. Each criterion has a small set of categories (e.g., four FM types, five architecture families, nine learning paradigms, four learning stages, five robotic tasks, nine application domains) and is used to sort the 435 papers, with comparative tables and illustrative figures for every criterion. The phase periodization supplies the historical narrative; the taxonomy supplies the granular comparison.
Load-bearing premise
The whole taxonomy rests on the 435-paper corpus being representative of the field; if the search and screening filters (notably the post-2020 search window and the priority given to prominent venues) biased which papers were included, the phase boundaries and per-criterion comparisons could shift.
What would settle it
Re-run the literature search with an explicit pre-2021 strategy and without the prominence-priority filter, then check whether the five phase boundaries and the per-criterion proportions among the 435 papers survive; if the phase or category shares move materially, the review's central mapping is an artifact of the selection process.
If this is right
- Any new robotic foundation-model paper can be positioned in the five-phase scheme and classified along the six criteria, giving the field a shared vocabulary.
- The public-dataset report gives practitioners a practical map of available training and evaluation resources across robotic tasks.
- The comparative analysis identifies where the field has matured (e.g., perception and planning) and where bottlenecks remain, notably the lack of large-scale physical-world training data.
- The hierarchical challenges-and-directions discussion provides a roadmap that can be used to prioritize research efforts.
Where Pith is reading between the lines
- The five-phase periodization is likely to become the default way the field describes its own history, even though the phase boundaries depend on the search window and inclusion filters.
- Because the screening gave priority to prominent venues and excluded non-English or paywalled work, the per-criterion counts should be read as directional rather than exact.
- The taxonomy could be reused as a coding scheme for future bibliometric or living-review updates, providing a test of whether the five phases and six criteria remain stable as the field grows.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a systematic review of foundation models in robotics, claiming to be a holistic and highly granular map of the field. The authors describe a structured methodology: a multi-database search (IEEE Xplore, Google Scholar, Scopus, DBLP, arXiv, Web of Science), a Scopus example query with PUBYEAR>2020, iterative screening, and selection of 435 articles. The review organizes the literature into five research phases (2018–2021 through 2024–present) and analyzes it along six taxonomic criteria: FM type, NN architecture, learning paradigm, learning stage, robotic task, and application domain. It also includes comparative tables, a dataset report, and hierarchical challenges/future-directions discussion. The central claim is that this is the most complete and granular existing map of robotic FM research.
Significance. If the corpus is representative and the taxonomy is accurate, this survey would be a valuable reference resource: it is unusually systematic compared to narrative surveys, applies a consistent six-criterion framework across §§5–10, provides comparative summary tables (Tables 3–7), and consolidates a large body of recent work including datasets and application domains. The concrete Scopus query and the explicit screening steps are strengths that improve reproducibility relative to typical surveys. However, the significance is conditional on corpus representativeness and factual accuracy; the current search-window mismatch and a verifiable model misattribution in Table 2 undermine the reliability of the central map until corrected.
major comments (3)
- [§2.2 and §3.1.1] The Scopus example query in §2.2 uses 'PUBYEAR>2020', yet Phase 1 of the evolution (§3.1.1) covers 2018–2021. Pre-2021 papers can therefore enter only through the undocumented 'list of references' backward-chasing step. The manuscript does not provide a PRISMA-style flow diagram, search logs, or an excluded-study list. This creates a differential sampling mechanism by publication year: the earliest phase is not systematically searched, so the phase boundaries and per-criterion counts (e.g., the distribution of VLA papers across phases) may be artifacts of the search window rather than properties of the literature. This is load-bearing for the survey's central 'five phases' claim and for the quantitative taxonomic conclusions. Please provide a full search protocol with per-stage counts, or restrict/relabel Phase 1 as a backward-chased seminal-work overview and temper the phase-trend claim
- [§2.3] The screening step states that 'priority was given to research works originating from prominent robotics and AI/ML publication venues,' but 'prominent' is not operationalized, and non-English papers and paywalled full texts are excluded. Without a definition of prominence, an excluded-study list, or an analysis of how the prominence filter correlates with the taxonomy axes (e.g., FM type, venue, task), the 435-paper set cannot be distinguished from a convenience sample. This is a load-bearing limitation because the entire taxonomic distribution, including Tables 3–7, rests on the representativeness of this corpus. Please provide the excluded-study list, define the prominence criterion, and report the number of papers screened/excluded at each stage.
- [§3.2, Table 2] The DINO row in Table 2 cites 'Zhang et al., 2023' and describes a VFM that creates 'object attention maps for facilitating robot manipulation tasks' with 86M parameters and year 2021. This conflates the self-supervised vision transformer DINO by Caron et al. (2021) with the DETR-based detector DINO by Zhang et al. (2023). Since Table 2 is the paper's compendium of 'most common and widely adopted robotic FMs,' this misattribution is a concrete accuracy error that reduces confidence in the other model entries. Please correct the citation to the appropriate DINO paper (or clarify which DINO is meant) and verify the remaining rows for similar source inconsistencies.
minor comments (5)
- [Table 1] The 'Current survey' row lists limitations as '–'. Every review has limitations; adding the methodological limitations discussed in §2 would strengthen credibility.
- [§2.2] The text says the search 'primarily focused on research works published within the last five years' while the query uses PUBYEAR>2020. For a 2026 submission this is roughly consistent, but for clarity please state the exact date window used when the search was executed.
- [Table 6] In the table caption, the listed columns end with 'h) Indicative models'; the preceding column is likely 'g) Indicative models' or another letter. Please correct the numbering.
- [Table 2] The 'Param.' column uses '–' for several entries (e.g., SayCan, RT-2, SayPlan, Eureka). Clarify whether this means 'not reported' or 'not applicable' and add a footnote.
- [References] Several references are dated 2026 (e.g., Sun et al., 2026; Wang et al., 2026b) and may be preprints. Please verify their publication status or mark them as preprints consistently.
Circularity Check
No circularity: the survey's taxonomy is an external literature classification, not a derivation from its own input.
full rationale
This paper is a narrative and taxonomic literature review, not a derivation chain. Its central claims—five research phases, six classification criteria, per-criterion comparative tables, and a dataset report—are interpretive summaries of 435 external primary sources. No equation is fit to a subset of data and then reported as a prediction of a closely related quantity, and no parameter is defined in terms of the outcome it is said to explain. The screening methodology (PUBYEAR>2020 querying, exclusion of non-English and paywalled papers, and priority to prominent venues) raises legitimate representativeness and coverage concerns, and the mismatch between the Scopus search window and Phase 1 (2018–2021) may affect how the earliest phase is sampled. However, that is a methodological limitation about corpus selection, not circularity: the taxonomy does not claim to derive the corpus from itself, nor does it assert that its categories are forced by a self-citation. The differentiation claims against prior surveys in Table 1 are comparative judgments, not constructions that reduce to the paper's own inputs. No load-bearing self-citation chain was identified. Accordingly, the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption The six taxonomic criteria (FM type, NN architecture, learning paradigm, learning stage, robotic task, application domain) are the right axes for organizing robotic FM literature.
- domain assumption The five-phase periodization (2018-2021 ... 2024-present) marks real discontinuities in research practice.
- domain assumption The 435-paper corpus is representative of the field.
- domain assumption Model descriptions in the summary tables match their cited sources.
read the original abstract
Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towards adaptive, multi-function, generalpurpose agents, capable of operating in complex, open-world, and dynamic environments. This tremendous advancement is primarily driven by the emergence of Foundation Models (FMs), i.e., large-scale neural-network architectures trained on massive, heterogeneous datasets that provide unprecedented capabilities in multi-modal understanding and reasoning, long-horizon planning, and cross-embodiment generalization. In this context, the current study provides a holistic, systematic, and in-depth review of the research landscape of FMs in robotics. In particular, the evolution of the field is initially delineated through five distinct research phases, spanning from the early incorporation of Natural Language Processing (NLP) and Computer Vision (CV) models to the current frontier of multi-sensory generalization and real-world deployment. Subsequently, a highly-granular taxonomic investigation of the literature is performed, examining the following key aspects: a) the employed FM types, including LLMs, VFMs, VLMs, and VLAs, b) the underlying neural-network architectures, c) the adopted learning paradigms, d) the different learning stages of knowledge incorporation, e) the major robotic tasks, and f) the main real-world application domains. For each aspect, comparative analysis and critical insights are provided. Moreover, a report on the publicly available datasets used for model training and evaluation across the considered robotic tasks is included. Furthermore, a hierarchical discussion on the current open challenges and promising future research directions in the field is incorporated.
Figures
Forward citations
Cited by 1 Pith paper
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
Reference graph
Works this paper leans on
-
[2024]
arXiv preprint arXiv:2409.08249
doi: 10.48550/arXiv.2409.08249. arXiv preprint arXiv:2409.08249. Tomas Berriel Martins, Martin R Oswald, and Javier Civera. Open-vocabulary online semantic mapping for slam.IEEE Robotics and Automation Letters, 2025a. Tomas Berriel Martins, Martin R Oswald, and Javier Civera. Open-vocabulary online semantic mapping for slam.IEEE Robotics and Automation Le...
-
[2025]
doi: 10.15607/RSS.2025.XXI.028. Liang Xu, Ziyang Song, Dongliang Wang, Jing Su, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin, Xiaokang Yang, et al. Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2228–2238, 2023...
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.