Pith. sign in

REVIEW 4 cited by

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09781 v1 pith:E3V6DOHR submitted 2024-06-14 cs.CV

classification cs.CV
keywords animalmultimodalperceptionvideobehaviorclipsunderstandingvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is complex, subjective, and multimodal. With the rapid development of multimodal large language models(LLMs), new application have emerged for animal behavior understanding tasks in livestock scenarios. This study evaluates the visual perception capabilities of multimodal LLMs in animal activity recognition. To achieve this, we created piglet test data comprising close-up video clips of individual piglets and annotated full-shot video clips. These data were used to assess the performance of four multimodal LLMs-Video-LLaMA, MiniGPT4-Video, Video-Chat2, and GPT-4 omni (GPT-4o)-in piglet activity understanding. Through comprehensive evaluation across five dimensions, including counting, actor referring, semantic correspondence, time perception, and robustness, we found that while current multimodal LLMs require improvement in semantic correspondence and time perception, they have initially demonstrated visual perception capabilities for animal activity recognition. Notably, GPT-4o showed outstanding performance, with Video-Chat2 and GPT-4o exhibiting significantly better semantic correspondence and time perception in close-up video clips compared to full-shot clips. The initial evaluation experiments in this study validate the potential of multimodal large language models in livestock scene video understanding and provide new directions and references for future research on animal behavior video understanding. Furthermore, by deeply exploring the influence of visual prompts on multimodal large language models, we expect to enhance the accuracy and efficiency of animal behavior recognition in livestock scenarios through human visual processing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A two-stage pipeline, RAMQA, using a LLaVA pointwise ranker and a LLaMA multi-task generative re-ranker with document permutations, achieves state-of-the-art QA and retrieval scores on WebQA and MultiModalQA.

  2. TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A dual-codebook VQ tokenizer with shared indices reports GenEval 0.55 autoregressive generation, 0.63 reconstruction FID, and a 7.2% average understanding gain over LLaVA-1.5, though the gain is partly inherited from ...

  3. VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new graph-image benchmark generator shows that six large vision-language models are sensitive to layout, labeling, and visual defects across seven graph tasks.

  4. Leveraging Large Language Models for Generating Labeled Mineral Site Record Linkage Data

    cs.IR 2024-11 conditional novelty 4.0 of 10

    LLM-generated labels can fine-tune a RoBERTa model to nearly match the LLM's own linkage quality on mineral site data at 18x lower inference cost, but the evaluation uses only a handful of positive test pairs and a de...

Pith tools