By drawing object boxes and motion trails visually on video frames instead of serializing coordinates as text, BoxTuning reduces token costs dramatically and improves accuracy on video question answering benchmarks.
In: ICLR (2022)
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.
A contextual multimodal document retrieval benchmark (CMDR-Bench) and embedding model (CMDR-Embed) that jointly encodes multiple document pages and splits them into page-level representations, trained with a context-aware contrastive objective, outperforming non-contextual baselines by 13–16 nDCG@5.
ECoSim adds multi-modal controllability to pretrained diffusion and autoregressive traffic models via identity-initialized FiLM layers while using less than 1% paired control data on Waymo Open Sim Agents Challenge.
A 14B model trained on synthetic data from Brazilian clinical guidelines outperforms larger LLMs on new benchmarks for Brazilian healthcare protocols.
citing papers explorer
-
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
By drawing object boxes and motion trails visually on video frames instead of serializing coordinates as text, BoxTuning reduces token costs dramatically and improves accuracy on video question answering benchmarks.
-
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.
-
CMDR: Contextual Multimodal Document Retrieval
A contextual multimodal document retrieval benchmark (CMDR-Bench) and embedding model (CMDR-Embed) that jointly encodes multiple document pages and splits them into page-level representations, trained with a context-aware contrastive objective, outperforming non-contextual baselines by 13–16 nDCG@5.
-
ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation
ECoSim adds multi-modal controllability to pretrained diffusion and autoregressive traffic models via identity-initialized FiLM layers while using less than 1% paired control data on Waymo Open Sim Agents Challenge.
-
Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines
A 14B model trained on synthetic data from Brazilian clinical guidelines outperforms larger LLMs on new benchmarks for Brazilian healthcare protocols.