Pith. sign in

cs.MM

Multimedia

Roughly includes material in ACM Subject Class H.5.1.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Home activity benchmark shows AI question-answering gaps

HOME-KGQA tests multimodal KGQA on daily household tasks, where LLM methods lag behind their encyclopedic results.

· “HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities”

open re-runnable review →
Figure from the paper

Story-driven game embeds design rules for ADHD transition support

Literature on ADHD and LD challenges is turned into concrete game features that target academic, social, and organizational hurdles during t

· “From Design Principles to Prototype: A Game for Students with ADHD and Learning Disabilities Transitioning to Post-Secondary Education”

open re-runnable review →
Figure from the paper

Paired counterfactuals expose omni models at 32-41% binding accuracy

3600-item benchmark from 256 hours of video shows calibration without retraining lifts open models by up to 10 points and aids other tasks.

· “OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination”

open re-runnable review →
Figure from the paper

Automated matching creates region labels for better person search

ROGLE mines pseudo region-sentence pairs from existing data to add local alignment, raising accuracy on detailed natural-language queries.

· “ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search”

open re-runnable review →
Figure from the paper

Semantic codebook creates style-matched co-speech gestures

By organizing motion codes according to gesture semantics and applying reference prompts, the system produces motions faithful to both word,

· “PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation”

open re-runnable review →
Figure from the paper

This paper builds a fashion knowledge graph from textbooks and combines it with a…

A textbook-grounded fashion knowledge graph plus pruning and grounding retrieval steps improves fashion QA accuracy over non-RAG and KG-RAG…

· “FashionKG-RAG: Knowledge Graph-Enhanced Retrieval-Augmented Generation for Fashion Question Answering”

open re-runnable review →
Figure from the paper

0.7057 final: speech-to-mask pipeline takes second on MeViS-Audio

Delayed commitment, trajectory ranking, and verified empty recovery let absent queries stay empty.

· “Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026”

open re-runnable review →
Figure from the paper

Fake-news detector makes LLMs argue both sides before ruling

RoE-FND mines past reasoning mistakes into reusable guidelines and outperforms trained detectors on five benchmarks.

· “RoE-FND: Synergizing LLMs with Experiential Learning for Effective and Generalizable Evidence-Based Fake News Detection”

open re-runnable review →
Figure from the paper

Best AI model scores 76 percent on long Japanese video benchmark

NARU pairs narrative tracking with cultural subtext like “reading the air,” where even leading models struggle most.

· “NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video”

open re-runnable review →
Figure from the paper

Vibration feedback boosts realism in dynamic XR scenes

A 29-actuator glove translating motion and edges into touch raised perceived realism for video; static textures saw no gain.

· “Visual-to-Haptic Augmentation in XR: A Wearable Glove for Perceptual Grounding in Multimodal Interaction”

open re-runnable review →
Figure from the paper

Zero-shot drone retrieval beats fine-tuned rivals at 39.47 mR

Region alignment and semantic perturbations sharpen fine-grained attribute matching at zero added inference cost.

· “GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views”

open re-runnable review →
Figure from the paper

browse all of cs.MM → full archive · search · sub-categories