{"id":"2dfa36ab-b5c9-4c8b-94e2-dc8f5e8fca45","arxiv_id":"2502.08556","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey proposing a four-part taxonomy for human-centric foundation models and reviewing representative methods in each.","lead":"This survey organizes recent \"human-centric foundation models\" into four groups: perception, content generation, unified perception-generation, and agentic models. It is a useful map for researchers tracking how large AI models are being specialized for human understanding, synthesis, and humanoid robots.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The roadmap's value depends on the taxonomy being exhaustive and genuinely first, but the survey provides no search protocol and never compares itself to the prior Feng et al. 2024 survey it cites, leaving novelty and coverage unverified.","rationale":"The reader's conditional verdict already identifies the absence of a systematic selection protocol and the need to benchmark the first-survey claim against Feng et al. 2024. I agree with that assessment, and my stress-test confirms that this is the most load-bearing weakness because the survey's contribution is organizational rather than empirical: its value is precisely that readers can rely on the taxonomy and its completeness. The paper gives no formal verification, no reproducible code, and no parameter-free derivation that could independently support the taxonomy; the only evidence is the list of representative works, which is asserted rather than derived from a stated process. I do not see an internal inconsistency in the individual model summaries; they are broadly consistent with the literature as cited. The concern is therefore not about correctness of any single description but about the central generalization that the four categories are exhaustive and that the survey is the first of its kind. This is a verification gap, not a disagreement with community consensus, and it can be settled by the concrete test above. Since the reader already conditioned acceptance on addressing this gap, my recommendation is UNCHANGED rather than a new verdict. I marked agreement as partial because the reader's stated weakest assumption emphasizes selection bias and coverage, while I also foreground the closely related but distinct first-survey claim; both are part of the same underlying need for a documented, reproducible survey methodology.","tokens_in":13376,"tokens_out":3715,"duration_ms":42880,"concrete_test":"Run a systematic scope test: query arXiv and OpenAlex for papers up to 2025-02-12 using a fixed string such as 'human-centric foundation model' OR 'human foundation model' OR 'foundation model for 3D humans' in cs.CV, independently screen abstracts against the paper's own HcFM definition, classify every eligible paper into the four proposed categories or 'other,' and compare the result to Fig. 1. Also compare the resulting survey landscape to Feng et al. 2024 to check whether any prior survey already covers comparable scope. If eligible papers fall outside the four categories, or if a prior survey overlaps the claimed scope, the taxonomy and first-survey claims require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 1—'to our best knowledge, this is the first survey about human-centric foundation models with a novel taxonomy'—plus the abstract's promise to 'serve as a roadmap' requires two things to be true: no prior survey covers the same scope, and the four categories in Section 2/Fig. 1 fairly and exhaustively represent the field. Neither is established. The paper cites Feng et al. 2024, 'Foundation Models for 3D Humans,' and even references it in Section 1 as [Feng and others, 2024], but it never states how the present taxonomy differs from that prior survey. If that earlier survey already organizes human perception, generation, and agentic 3D-human foundation models, the first-survey claim fails. Separately, Section 2 says models are grouped 'according to their supported downstream tasks,' but no search, screening, or selection protocol is reported; Fig. 1 lists only about 29 exemplars and includes several author-affiliated systems (UniHCP, PATH, Hulk, MotionGPT-2). Without a completeness argument, the reader cannot tell whether the four-part structure reflects the field or a curated subset. The agentic category in Section 6 is especially thin—only HumanVLA, SuperPADL, and GR00T—and the text itself says VLA humanoid models are 'largely unexplored'; if this is meant to be a major fourth pillar, the representative choice needs explicit justification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a survey of human-centric foundation models (HcFMs), proposing a taxonomy that divides the field into four categories: (1) human-centric perception foundation models, (2) human-centric AIGC foundation models, (3) unified perception and generation models, and (4) human-centric agentic foundation models. For each category, the authors review representative methods, describe their underlying learning frameworks (e.g., contrastive learning, masked image modeling, diffusion transformers, LLM-integrated vocabularies), and discuss challenges and future directions in data, representation, interactivity, and ethics. The paper claims to be the first survey on HcFMs with a novel taxonomy and positions itself as a roadmap for researchers and practitioners working on digital humans and humanoid embodiments.","tokens_in":13673,"tokens_out":2718,"duration_ms":28371,"significance":"If the proposed taxonomy is accepted and its coverage is representative, the survey would provide a useful organizational structure for a rapidly growing but fragmented field. The paper includes several strengths: it gives clear conceptual distinctions among perception, generation, unified, and agentic models; it illustrates each category with framework diagrams (Figs. 2–5); it discusses both self-supervised and supervised paradigms within perception and generation; and it explicitly enumerates open challenges in data, representation, interactivity, and ethics. The survey also draws attention to the emerging area of humanoid agentic models. However, the significance of the survey as a 'roadmap' depends on two unverified premises: that no prior survey covers the same scope, and that the four-category taxonomy and the selected representative works fairly capture the field. These premises are not established in the manuscript, which limits the current contribution to a potentially useful but incomplete organizational proposal.","major_comments":[{"comment":"The claim that 'this is the first survey about human-centric foundation models with a novel taxonomy' is not supported by evidence. The paper cites [Feng and others, 2024] ('Foundation Models for 3D Humans') in Section 1 but never states how the present taxonomy differs from or improves upon that prior survey. Without an explicit comparison to earlier surveys—including [Feng and others, 2024] and any other relevant overviews—the novelty claim remains an assertion. The authors should either provide such a comparison or soften the claim to avoid an unverifiable first-survey statement.","section":"Section 1 and References"},{"comment":"The taxonomy is introduced with no inclusion criteria or search protocol. The text states that models are grouped 'according to their supported downstream tasks,' but it does not specify how works were identified, screened, or selected for inclusion. Fig. 1 lists only about 29 representative examples, which is far from the full landscape of human-centric foundation models. The absence of a completeness argument means the four categories cannot be verified as exhaustive, and the roadmap promise in the abstract is not justified. The authors should add a methodology subsection describing the literature search and selection process, and clearly state which works are representative rather than exhaustive.","section":"Section 2 and Fig. 1"},{"comment":"The agentic category is disproportionately thin compared to the other three categories. Section 6 discusses only HumanVLA, SuperPADL, and GR00T, and the text itself admits that applying VLA models to humanoid robotics 'remains largely unexplored.' If agentic modeling is meant to be a major fourth pillar of the taxonomy, the representative choice needs explicit justification, and the survey should explain why this category is included at the same level as perception, generation, and unified models despite its limited maturity. This is not a fatal flaw, but it weakens the claim of a balanced and exhaustive taxonomy.","section":"Section 6 and Fig. 1"},{"comment":"The selection of representative examples appears to be skewed toward works with which the authors are affiliated. For example, PATH (Tang et al., 2023), UniHCP (Ci et al., 2023), Hulk (Wang et al., 2023), MotionGPT (Jiang et al., 2023), and MotionGPT-2 (Wang et al., 2024b) are all associated with the authors of this survey, and several are given prominent placement in Fig. 1 and the main text. This is not inherently improper, but the survey provides no statement of how representatives were chosen or whether conflicts of interest influenced selection. A brief note on selection neutrality or a statement of author contributions to surveyed works would address the concern.","section":"Fig. 1 and Sections 3, 5"}],"minor_comments":[{"comment":"There are numerous typographical and formatting errors: 'NeruIPS' for NeurIPS (reference [Jiang et al., 2023]), 'labled' instead of 'labeled' (Section 3.2), 'Yoshikawaet al.' missing space (reference [Yoshikawa et al., 2023]), and 'V ocabulary' with an extra space in the Fig. 4 caption. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The sentence 'The four categories of foundation models are interconnected rather than mutually exclusive' is in tension with the earlier statement that models are classified 'according to their supported downstream tasks.' The authors should clarify how a model that spans multiple categories (e.g., a unified perception-generation model that also has agentic capabilities) would be classified.","section":"Section 2"},{"comment":"In the text introducing contrastive learning methods, the phrase 'instead of commonly used momentum encoders, multiple encoders were used' is unclear. It is not obvious what is being contrasted with what; consider rewriting for clarity.","section":"Section 3.1"},{"comment":"The caption says 'Parameters in modules with are used in downstream tasks,' which appears to be missing a symbol or a word. This obscures the intended meaning and should be fixed.","section":"Fig. 2 caption"},{"comment":"The 'Ethics' paragraph mentions that 'anonymization methods should be applied to all training data,' but it does not discuss the trade-offs between anonymization and data utility, nor does it reference specific existing work on privacy-preserving human-centric learning. Adding a pointer to relevant literature would strengthen this discussion.","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a more rigorous positioning relative to prior surveys, especially the cited [Feng and others, 2024] work. The lack of a methodology section for the literature search and the heavy representation of the authors' own papers in the selected examples may raise concerns among readers about the survey's objectivity and completeness. These issues are fixable within the manuscript's scope, but they require substantive additions rather than local edits. The journal may also want to consider whether the survey's coverage is sufficiently broad for its claimed 'roadmap' scope, particularly in the agentic category."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What's new: the taxonomy itself. Grouping HcFMs into perception, AIGC, unified perception-generation, and agentic is a sensible way to slice a very fast-moving literature. Fig. 1 gives a newcomer a quick map, and the short descriptions of individual models are accurate as far as I can check them. The agentic category is forward-looking even if thin.\n\nThe soft spots are real but fixable. The \"first survey\" claim in Section 1 is asserted \"to our best knowledge,\" yet the paper cites Feng et al. 2024, \"Foundation Models for 3D Humans,\" and never says how the present scope differs. That needs to be addressed head-on. Section 2 says models are grouped by supported tasks, but there is no search strategy, inclusion criteria, or completeness argument. Without that, the roadmap is at risk of being a curated subset rather than a map of the field. The self-citation pattern (PATH, UniHCP, Hulk, MotionGPT-2) becomes harder to wave off precisely because the selection protocol is missing. And the agentic section is thin—three models, with the text admitting VLA humanoid models are \"largely unexplored.\" As a \"fourth pillar\" it needs either more scaffolding or a clear statement that it is a nascent category.\n\nNone of this sinks the survey. For a practitioner or newcomer wanting a structured entry point into human-centric foundation models, the paper delivers. It is not a new scientific result, but it is a serviceable organizational contribution. With a methodology paragraph, a direct comparison to prior surveys like Feng et al., and a recalibrated agentic section, it would be a solid community resource.\n\nI'd send it to peer review. It deserves referee time, with the expectation of a revision. I'd probably cite it in my own intro if I worked in this area.","headline":"Useful organizational survey of human-centric foundation models, but the first-survey claim and missing methodology keep it from being a fully trustworthy roadmap.","tokens_in":14167,"tokens_out":1831,"would_cite":true,"duration_ms":18958,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes that human-centric foundation models divide into four families: perception, generation, unified perception-and-generation, and agentic models.","keywords":["human-centric foundation models","taxonomy","perception","AIGC generation","unified perception-generation","agentic models","multimodal large language models","humanoid embodied AI"],"falsifier":"Take the most recent 50 human-centric foundation models from top computer-vision and robotics venues and ask two independent human annotators to place each into exactly one of the four categories. If a substantial share cannot be placed, or if annotators disagree on which category a model belongs to, the taxonomy's exhaustiveness and clarity fail. A sharper check: find one model that performs perception, generation, and physical action control in one shared parameter set without treating any of these as a foreign language in an LLM; that model would sit outside all four categories as stated.","tokens_in":13175,"feed_emoji":"🧍","tokens_out":7304,"duration_ms":72625,"temperature":0.7,"pith_summary":"Human-centric foundation models are generalist models trained to handle many tasks involving human bodies, faces, motion, and behavior within one framework. This paper proposes that those models fall into four families: perception, AI-generated content, unified perception-and-generation, and agentic models, grouped by the downstream tasks they support. The authors claim this is the first survey to organize the area with this taxonomy, and they support it by mapping representative methods, learning paradigms, and open challenges onto each family. The payoff, if the taxonomy holds, is a shared map that lets researchers compare models, spot missing capabilities, and choose starting points for new work.","feed_headline":"Survey sorts human-centric AI into four families","feed_subtitle":"Perception, generation, unified understanding-plus-synthesis, and embodied agents give researchers one map of the field.","key_machinery":"The load-bearing structure is the taxonomy itself, a two-level classification: four task-based families, each with paradigm-based subcategories. The organizing signal is downstream task support, so perception, generation, unified perception-generation, and agentic behavior are the four buckets. Subcategories include contrastive learning and masked image modeling under perception; GANs with style modulation and diffusion models under generation; fixed-vocabulary and extended-vocabulary LLM integration under unified models; and vision-language versus vision-language-action architectures under agents. The taxonomy does the work: it turns a scattered literature into a coordinate system for comparison.","core_discovery":"The central claim is that the sprawling set of human-centric foundation models can be organized by what they do with people: perceiving them, generating them, doing both through a language-model hub, or acting in the world with human-like embodiment. Perception models learn fine-grained human representations for tasks like re-identification, parsing, pose estimation, mesh recovery, and action recognition; AIGC models synthesize human-focused images, videos, and avatars with high fidelity; unified models treat human-centric cues such as skeletons, body-model parameters, motion tokens, or audio as 'foreign languages' attached to large language models, so one model can both understand and generate; agentic models take vision, language, and other sensor signals and map them to humanoid motor behavior. The paper further splits each family by training paradigm and presents the four families as an interconnected roadmap for future digital-human and humanoid-embodiment research.","pith_inferences":["Editorial extension: the four families are likely to blur, since a capable agentic model may soon include its own perceptual and generative components, so future taxonomies may need fewer or overlapping categories rather than four disjoint buckets.","Editorial extension: the fixed-versus-extended-vocabulary distinction predicts a testable trade-off—frozen-vocabulary tool use is cheaper and safer to extend, while grown vocabularies give finer control over human-centric outputs at higher training cost.","Editorial extension: a concrete way to stress-test the taxonomy is to classify the full set of recent humanoid robotics models; many may straddle the vision-language and vision-language-action split, suggesting the agentic category needs its own finer-grained subcategories.","Editorial extension: if the 'first survey' claim becomes influential, later papers should increasingly cite this taxonomy as their organizing frame, which is a checkable bibliographic prediction."],"forward_implications":["New or existing human-centric models can be located on the taxonomy by asking which downstream tasks they support, making comparisons and gap analysis more systematic.","The unified perception-and-generation family indicates that LLMs and multimodal LLMs are becoming the default hub, with human-centric signals treated as foreign languages.","The agentic family marks embodiment, interaction, and humanoid control as a frontier distinct from perception and generation.","The survey's challenges section implies that progress hinges on solving human-data scarcity, holistic body-face-hand representation, interactivity, and privacy ethics.","If the taxonomy is accepted, it can serve as a shared reference for structuring future human-centric foundation-model research and evaluation."],"supporting_citations":[{"why":"SOLIDER anchors the contrastive self-supervised branch of human-centric perception by adding semantic priors to learned features.","marker":"[Chen et al., 2023a]"},{"why":"PATH anchors multitask supervised pretraining for perception, using task-specific projectors with shared weights.","marker":"[Tang et al., 2023]"},{"why":"UniHCP anchors the unified modeling branch, handling five human-centric tasks with a task-guided interpreter.","marker":"[Ci et al., 2023]"},{"why":"StyleGAN-Human anchors the GAN-with-style-modulation branch of human-centric generation and downstream applications.","marker":"[Fu et al., 2022]"},{"why":"HumanSD anchors conditional latent diffusion for human image generation with skeleton-guided denoising.","marker":"[Ju et al., 2023]"},{"why":"MotionGPT anchors extended-vocabulary unified models by treating motion as a foreign language appended to an LLM.","marker":"[Jiang et al., 2023]"},{"why":"ChatHuman anchors fixed-vocabulary unified models by wiring 22 domain-specific human tools into a multimodal LLM.","marker":"[Lin et al., 2024]"},{"why":"HumanVLA anchors vision-language-based agentic models for humanoid embodied control.","marker":"[Xu et al., 2024]"},{"why":"Project GR00T anchors the vision-language-action branch of agentic models for humanoid robots.","marker":"[Dong et al., 2024]"}],"fun_headline_variants":["Four families of human-centric AI: perceive, generate, unify, act","A map for human-aware AI: from perception to embodiment","Human-centric foundation models: one framework, four roles","Digital-human AI gets a four-part taxonomy","See, make, combine, embody: the new AI roadmap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The four-category taxonomy carries the whole survey, so it must be both exhaustive and fairly representative; the paper does not report a systematic search or selection protocol, so if a significant class of human-centric foundation models falls outside the four families or the representative list skews, the roadmap misleads readers.","fun_headline_variants_meta":{"raw":{"variants":["Four families of human-centric AI: perceive, generate, unify, act","A map for human-aware AI: from perception to embodiment","Human-centric foundation models: one framework, four roles","Digital-human AI gets a four-part taxonomy","See, make, combine, embody: the new AI roadmap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000811,"raw_usage":{"total_tokens":3543,"prompt_tokens":918,"completion_tokens":2625,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":2543}},"tokens_in":534,"tokens_out":2625,"duration_ms":20247,"temperature":1.0,"reasoning_tokens":2543,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T04:37:29.434757+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the most recent 50 human-centric foundation models from top computer-vision and robotics venues and ask two independent human annotators to place each into exactly one of the four categories. If a substantial share cannot be placed, or if annotators disagree on which category a model belongs to, the taxonomy's exhaustiveness and clarity fail. A sharper check: find one model that performs perception, generation, and physical action control in one shared parameter set without treating any of these as a foreign language in an LLM; that model would sit outside all four categories as stated.","supporting_citations":[{"cited_title":"Humanbench: To- wards general human-centric perception with projector as- sisted pretraining","cited_arxiv_id":null,"evidence_quote":"PATH anchors multitask supervised pretraining for perception, using task-specific projectors with shared weights."},{"cited_title":"Motiongpt: Human motion as a foreign language","cited_arxiv_id":null,"evidence_quote":"MotionGPT anchors extended-vocabulary unified models by treating motion as a foreign language appended to an LLM."},{"cited_title":"Bringing Robots Home: The Rise of AI Robots in Consumer Electronics","cited_arxiv_id":"2403.14449","evidence_quote":"Project GR00T anchors the vision-language-action branch of agentic models for humanoid robots."}],"review_version":1}