Pith. sign in

REVIEW 2 cited by

Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09306 v2 pith:3R4N67BH submitted 2024-09-14 cs.CV

classification cs.CV
keywords humanunderstandingdatamodelsactionbenchmarkmodelmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision-language instruction-following data. We introduce a method for generating such data by integrating human keypoints with traditional visual features such as captions and bounding boxes, enabling more precise understanding of human-centric scenes. Our approach constructs a dataset comprising 200,328 samples tailored to fine-tune models for human-centric tasks, focusing on three areas: conversation, detailed description, and complex reasoning. We establish a benchmark called Human Pose and Action Understanding Benchmark (HPAUB) to assess model performance on human pose and action understanding. We fine-tune the LLaVA-1.5-7B model using this dataset and evaluate it on the benchmark, achieving significant improvements. Experimental results show an overall improvement of 21.18% compared to the original LLaVA-1.5-7B model. These findings highlight the effectiveness of keypoint-integrated data in enhancing multimodal models. Code is available at https://github.com/Ody-trek/Keypoint-Instruction-Tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Feeding human keypoints into GPT-4o prompts to create pose-focused instruction data and fine-tuning LLaVA-1.5 with it raises scores on the authors' self-generated E-HPAUB benchmark.

  2. PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment

    cs.CV 2025-07 conditional novelty 3.0 of 10

    PoseLLM swaps LocLLM's linear vision-language projector for a two-layer MLP with GELU, reporting +0.4 AP on COCO (77.8) with comparable zero-shot transfer.

Pith tools