ChronoPhyBench is a new benchmark and dataset for chronological physical dynamics reasoning that combines video-conditioned next-state prediction with VQA to reduce language bias in MLLM evaluation.
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
6 Pith papers cite this work. Polarity classification is still indexing.
abstract
Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intrinsic heterogeneity of diverse modalities presents substantial challenges to achieve effective cross-modal collaboration and integration. To address this, we introduce DecAlign, a novel hierarchical cross-modal alignment framework designed to decouple multimodal representations into modality-unique (heterogeneous) and modality-common (homogeneous) features. For handling heterogeneity, we employ a prototype-guided optimal transport alignment strategy leveraging gaussian mixture modeling and multi-marginal transport plans, thus mitigating distribution discrepancies while preserving modality-unique characteristics. To reinforce homogeneity, we ensure semantic consistency across modalities by aligning latent distribution matching with Maximum Mean Discrepancy regularization. Furthermore, we incorporate a multimodal transformer to enhance high-level semantic feature fusion, thereby further reducing cross-modal inconsistencies. Our extensive experiments on four widely used multimodal benchmarks demonstrate that DecAlign consistently outperforms existing state-of-the-art methods across five metrics. These results highlight the efficacy of DecAlign in enhancing superior cross-modal alignment and semantic consistency while preserving modality-unique features, marking a significant advancement in multimodal representation learning scenarios. Our project page is at https://taco-group.github.io/DecAlign.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
DFPL introduces prototype-based disentanglement and alignment modules to preserve fine-grained consistency across heterogeneous modalities for better performance under missing data conditions.
People apply stricter, more rule-based moral standards to AI systems and their engineers when the AI's human programming is made explicit, while judging the same AI and a human actor similarly when programming is invisible.
IndusAgent achieves state-of-the-art zero-shot performance on industrial anomaly benchmarks by using a custom Indus-CoT dataset, dynamic tool orchestration, and gated RL to optimize anomaly classification, localization, and reasoning.
Introduces MAF framework and DeepModal-Bench to capture universal cross-modal forgery traces for better generalization in multimodal deepfake detection.
DVP-MVS++ combines depth-normal-edge alignment via erosion-dilation and harmonized visibility priors with geometry consistency checks to achieve state-of-the-art multi-view stereo results on ETH3D, Tanks & Temples, and Strecha.
citing papers explorer
-
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?
ChronoPhyBench is a new benchmark and dataset for chronological physical dynamics reasoning that combines video-conditioned next-state prediction with VQA to reduce language bias in MLLM evaluation.
-
Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification
DFPL introduces prototype-based disentanglement and alignment modules to preserve fine-grained consistency across heterogeneous modalities for better performance under missing data conditions.
-
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
People apply stricter, more rule-based moral standards to AI systems and their engineers when the AI's human programming is made explicit, while judging the same AI and a human actor similarly when programming is invisible.
-
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
IndusAgent achieves state-of-the-art zero-shot performance on industrial anomaly benchmarks by using a custom Indus-CoT dataset, dynamic tool orchestration, and gated RL to optimize anomaly classification, localization, and reasoning.
-
Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities
Introduces MAF framework and DeepModal-Bench to capture universal cross-modal forgery traces for better generalization in multimodal deepfake detection.
-
DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
DVP-MVS++ combines depth-normal-edge alignment via erosion-dilation and harmonized visibility priors with geometry consistency checks to achieve state-of-the-art multi-view stereo results on ETH3D, Tanks & Temples, and Strecha.