Pith. sign in

REVIEW 4 cited by

Hulk: A Universal Knowledge Translator for Human-Centric Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01697 v5 pith:ZN6XFZX2 submitted 2023-12-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords human-centrictaskshulkheadstask-specificachievingbenchmarksemph
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Human-centric perception tasks, e.g., pedestrian detection, skeleton-based action recognition, and pose estimation, have wide industrial applications, such as metaverse and sports analysis. There is a recent surge to develop human-centric foundation models that can benefit a broad range of human-centric perception tasks. While many human-centric foundation models have achieved success, they did not explore 3D and vision-language tasks for human-centric and required task-specific finetuning. These limitations restrict their application to more downstream tasks and situations. To tackle these problems, we present Hulk, the first multimodal human-centric generalist model, capable of addressing 2D vision, 3D vision, skeleton-based, and vision-language tasks without task-specific finetuning. The key to achieving this is condensing various task-specific heads into two general heads, one for discrete representations, \emph{e.g.,} languages, and the other for continuous representations, \emph{e.g.,} location coordinates. The outputs of two heads can be further stacked into four distinct input and output modalities. This uniform representation enables Hulk to treat diverse human-centric tasks as modality translation, integrating knowledge across a wide range of tasks. Comprehensive evaluations of Hulk on 12 benchmarks covering 8 human-centric tasks demonstrate the superiority of our proposed method, achieving state-of-the-art performance in 11 benchmarks. The code will be available on https://github.com/OpenGVLab/Hulk.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios

    cs.CV 2024-12 conditional novelty 6.0 of 10

    One transformer handles referring expression comprehension, keypoint detection, human parsing, and human captioning, and a new benchmark tests reasoning about people.

  2. Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An iterative detector-mask-pose loop with a mask-conditioned pose model, MaskPose, sets new state-of-the-art results on OCHuman while matching top-down COCO pose accuracy.

  3. Enhancing Sports Strategy with Video Analytics and Data Mining: Automated Video-Based Analytics Framework for Tennis Doubles

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A tennis doubles annotation framework is built and evaluated, showing transfer-learned CNNs outperform pose-only GCNs for automated shot and formation labeling.

  4. Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A survey proposing a four-part taxonomy for human-centric foundation models and reviewing representative methods in each.

Pith tools