Pith. sign in

GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Large vision-language models (LVLMs), such as the Generative Pre-trained Transformer 4-omni (GPT-4o), are emerging multi-modal foundation models which have great potential as powerful artificial-intelligence (AI) assistance tools for a myriad of applications, including healthcare, industrial, and academic sectors. Although such foundation models perform well in a wide range of general tasks, their capability without fine-tuning is often limited in specialized tasks. However, full fine-tuning of large foundation models is challenging due to enormous computation/memory/dataset requirements. We show that GPT-4o can decode hand gestures from forearm ultrasound data even with no fine-tuning, and improves with few-shot, in-context learning.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Human Re-ID Meets LVLMs: What can we expect?

cs.CV · 2025-01-30 · conditional · novelty 4.0

On a curated 20-query subset of Market1501, PersonViT strongly outperforms ChatGPT-4o, Gemini-2.0-Flash, Claude 3.5 Sonnet, and Qwen-VL-Max in separating genuine from impostor person matches.

citing papers explorer

Showing 1 of 1 citing paper.

  • Human Re-ID Meets LVLMs: What can we expect? cs.CV · 2025-01-30 · conditional · none · ref 7 · internal anchor

    On a curated 20-query subset of Market1501, PersonViT strongly outperforms ChatGPT-4o, Gemini-2.0-Flash, Claude 3.5 Sonnet, and Qwen-VL-Max in separating genuine from impostor person matches.