Home activity benchmark shows AI question-answering gaps
HOME-KGQA tests multimodal KGQA on daily household tasks, where LLM methods lag behind their encyclopedic results.
Multimedia
Roughly includes material in ACM Subject Class H.5.1.
sort pith recommended most recent
HOME-KGQA tests multimodal KGQA on daily household tasks, where LLM methods lag behind their encyclopedic results.
Literature on ADHD and LD challenges is turned into concrete game features that target academic, social, and organizational hurdles during t
X-ray microtomography and refined algorithms recover full text from unopened ancient papyrus without damage.
· “Complete virtual unwrapping and reading of a rolled Herculaneum papyrus”
3600-item benchmark from 256 hours of video shows calibration without retraining lifts open models by up to 10 points and aids other tasks.
ROGLE mines pseudo region-sentence pairs from existing data to add local alignment, raising accuracy on detailed natural-language queries.
By organizing motion codes according to gesture semantics and applying reference prompts, the system produces motions faithful to both word,
Multi-scale probing and phase-matched aggregation cut reconstruction error versus standard LDP methods at tight privacy budgets.
· “Period-conscious Time-series Reconstruction under Local Differential Privacy”
Separating expression and pose features while highlighting intense motion frames produces more accurate talking-head video from audio.
· “KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation”
Suite training transfers to seven external benchmarks; rule-based scorers beat VLM judges and power multi-task RL.
· “VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning”
Dual-brain retrieval hides inside the silence gap, adding facts and empathy without slowing the conversation down.
· “VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction”
Regional 2-m temperature forecasting as query-conditioned field evaluation beats fixed-grid baselines in skill.
· “Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries”
A spectral-local-wavelet backbone lifts CSI at 35 dBZ from 0.094 to 0.146.
DART-I routes spatial and color priors into hidden states, beating LoRA fine-tuning without retraining the backbone.
Semantic-aware density control puts Gaussian detail where eyes go, not evenly, cutting storage to a fraction.
· “SACHA: Semantic-Aware Compression for 3D Gaussian Head Avatars”
Naming helps some models spot a concept; spotting it rarely helps them locate the moment in time.
· “Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia”
All four teams beat chance on 185 questions over 600+ hours of egocentric video; nine questions stumped everyone.
Per-layer shared-private experts keep text, vision, and audio cues at their natural semantic depth.
· “Adaptive Hierarchical Representation Alliance for Multimodal Learning”
A textbook-grounded fashion knowledge graph plus pruning and grounding retrieval steps improves fashion QA accuracy over non-RAG and KG-RAG…
Delayed commitment, trajectory ranking, and verified empty recovery let absent queries stay empty.
Per-field reference retrieval plus targeted repair turns vague captions into aligned art, ranking 2nd in the challenge.
· “ReART: Reference-Guided Retrieval and Refinement for Emotion-Aware Art Generation”
Even top models stay below 63% accuracy on scenes whose look contradicts their physical state.
· “Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models”
Q2IT lifts Gemini's joint accuracy from 0.507 to 0.722 on the ITJoint benchmark.
· “Query-Driven Multimodal Information Extraction from Long Documents”
On corrupted MJPEG and H.264 videos, injecting byte-derived action priors raises PSNR by 2.51 dB, CIDEr by 0.20, and PCK by 0.18.
Splitting recognition from explanation in frozen models lifts fine-grained reasoning by 28 points.
A dual loss forces vision-language models to prefer intact images, cutting hallucination rates on three benchmarks.
· “PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment”
Humanoid performers let researchers isolate acoustic, visual, and social cues in music-evoked emotion.
· “Humanoid Musical Robots as Experimental Interfaces for Music-Evoked Emotion”
With the backbone frozen, routing small adapters by subject, grade, and visual cues adds 2.6 points of accuracy.
A frozen model trained only on synthetic episodes matches task-specific rivals; light adapters top most benchmarks.
· “Pretraining Reusable Inference Across Views with Synthetic Task Priors”
Images, video, and audio become first-class nodes, so SPARQL can search and segment them directly.
· “MediaGraph: A Content-Aware Data Model and Query Framework for Multimodal Knowledge Graphs”
A new prosody-only benchmark shows tone, not words, carries confident, sarcastic, and impatient attitudes.
· “SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis”
Model-generated training data plus latent audio alignment lifts Qwen2.5-Omni by 14.9 points on folk music.
Music-computing survey finds single-modality studies and biased sample corpora; it maps three data-driven fixes.
Anchored rollouts plus motion rewards break the freeze that plagues distilled streaming avatars.
White-box attack steers an ODE flow toward target identities, beating published baselines on six face models.
· “Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching”
One universal image codec steers bits to task-relevant regions at runtime, matching task-specific codecs within 1.9 points.
· “UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures”
Identical retinal input still yields non-human scanpaths, so outcome metrics miss the process gap.
· “Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans”
Draft the motion, face-swap a few anchors, interpolate the rest — no fine-tuning needed.
· “KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation”
AnyTalk lifts 2D talking-head video to blendshape animation, no per-character training data required.
· “AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model”
CLARA aligns speech, audio, and visuals inside short clips to catch hate that builds across a video.
· “CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection”
Training on future-frame embeddings matches explicit visual chain-of-thought while cutting inference latency by over 5x.
· “Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning”
Synthetic data from appliance manuals lets a small open model outplan zero-shot giants on real robots.
· “Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning”
Map-free path, turn, and goal overlays drive frozen-follower success to 63.3 percent.
RoE-FND mines past reasoning mistakes into reusable guidelines and outperforms trained detectors on five benchmarks.
A single set of weights switches between rates 2, 4, 8, and 16, cutting storage needs by 61%.
· “Flexible Deep Joint Source-Channel Coding: A Vibrotactile Example”
Prepared once and stored at the edge, each request sends only the parameters it needs, cutting rate and latency.
· “ParaJSCC: A Parameterized Framework for Reusable Multimodal Joint Source-Channel Coding”
Tool-augmented meta-editing reaches 56.1% on a hard exam benchmark, beating a 48.7% proprietary baseline.
A 210-bit codebook stream keeps source meaning and matches receiver history, even at low SNR.
· “Personalized Digital Semantic Communication for Image Transmission with Vision-Language Models”
A metric that splits similarity into 20 explainable facial attributes matches human choices and per-attribute reasoning
· “AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations”
Six hours of duo improvisation, annotated from both sides, show perception lags intention.
· “H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation”
NARU pairs narrative tracking with cultural subtext like “reading the air,” where even leading models struggle most.
New framework DRUF drops correlated-field leakage from 64% to near zero while preserving extraction quality.
· “Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs”
Draft attention drifts up to 21 frames; verification attention recenters it to 2, restoring speed.
· “Alignment Drift in Single-Model Speculative Decoding for ASR: Mechanism, Correction, and Cost”
When narrowing to one place or hour, MASCOT keeps R@10 at 0.94 where the DPP baseline drops to 0.49.
· “MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval”
A two-stage optimizer tunes prompt, seed, and guidance scale, winning up to 69% of human preference tests.
· “Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence”
Per-layer bit allocation keeps compression quality while killing cross-platform decode failures.
· “HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression”
A single loss-rate-aware model scatters the damage, restores missing latent features, and beats standard codecs on ShapeNet and…
· “ResPCC: A Loss-Resilient Neural Point Cloud Codec over Lossy Networks”
Where static checks fix only 2 of 11, real inference feedback drives 100% recovery on the CMLR system.
Same pseudo-masks, new losses: open-vocabulary segmentation improves through captions and CLIP-verified synonyms.
Speech-driven adapter beats baselines on a new 246-hour public-domain film benchmark.
· “Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections”
Frozen Whisper frames with a linear projector reach 97.3% on music AVQA, beating larger omni-modal fine-tuning.
· “Qwen-MusicAVQA-7B: A Multimodal Model for Music Audio-Visual QA”
Frozen embeddings from score-conditioned JEPA training beat separate baselines on quality, ranking, technique, and mistake tasks.
· “MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space”
A relative 4D scene graph lifts object-question accuracy 6.7 points and when-questions 12.5 points on EgoLifeQA.
Multi-hop recall, temporal reasoning, preference drift, and abstention all get measured on synthesized year-scale trajectories.
A thin rule layer between Media over QUIC and multipath QUIC is the only setup to meet the 150 ms interactive-latency budget.
· “Media-over-Multipath-QUIC for Realtime Video Applications”
74,234 paintings each get narrative, formal, emotional, and historical captions — none works alone.
· “MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding”
Moves plenoptic light-field attributes into G-PCC, reporting 1–3 dB gains over earlier schemes.
A 29-actuator glove translating motion and edges into touch raised perceived realism for video; static textures saw no gain.
New objective stops models leaning on text cues and improves diagnosis on glaucoma and chest X-ray benchmarks.
When 95% of items lose a modality, the model keeps 61% of its accuracy; without the trick, only 22%.
· “Sequential Modality Dropout for Robust Multi-Modal Sequential Recommendation”
Giving the biggest slice of a 3 Mbps uplink to the camera the operator needs lifts success from 48% to 71%.
· “TAMS: Task-Aware Multi-View Adaptive Streaming for Wireless Telerobotic Manipulation”
Region alignment and semantic perturbations sharpen fine-grained attribute matching at zero added inference cost.
A shared tokenizer splits semantic and motion detail, letting one language model translate and produce signs from pose data.
It writes a time-aligned structural plan before audio and beats a matched no-layout baseline on boundary scores.
· “MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation”
Reports 1.3x adaptive audio playback, 53% shorter lecture viewing, and pronunciation gains.
A-PACK thins video first, then prunes low-relevance audio and video inside the LLM, lifting decoding throughput 2.21x
· “Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs”
Shuffling structural and motion tokens across packets lets receivers rebuild lost frames from survivors.
· “Loss-Resilient Wireless Video Token Communication over Block Fading Channels”
Causal sparse attention and a user smoothness dial keep quality high at 18.6 FPS on 768×1408 clips.
· “MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling”
A relay that shares base content and selectively adds enhancement cuts link load 27% versus clustering.
· “MD2G-Cast: Relay-Coordinated Multicast for Scalable Volumetric Streaming over MoQ”
One change to a character or scene ripples through every linked sound asset automatically.
· “Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books”
Swapping timestamp regression for a binary-question scan lifts R@0.5 by 28–50 points with zero training.
· “Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No”