MLLMs drop from over 85% accuracy on action presence to under 50% on matched action-denial videos, exposing a causal verification gap that causal graph prompts partially close.
hub
International Journal of Computer Vision 131(1), 284–301 (2023)
15 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
A variational autoencoder learns quantum embeddings compressing ImageNet into 13 qubits and achieving 98.5% accuracy on MNIST 3-vs-5 classification with a quantum circuit, close to classical baselines and far above naive amplitude embeddings.
An ILP-based oracle applied to seven VIS methods on YouTube-VIS and OVIS shows tracking instability as the dominant bottleneck, producing gaps exceeding 20 AP under occlusion while classification impact is secondary.
DHCNet improves ultra-fine-grained visual categorization by progressively building holistic cognition from local discrepancies using self-shuffling and refinement on limited data.
CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.
A variational latent bottleneck with KL regularization and a dynamic binary mask based on saliency produces model-specific features that keep high accuracy for one classifier but drop others below 2% on CIFAR-100 with over 45x suppression.
Inference-time α-entmax sparsification of CLIP’s final self-attention denoises diffuse mass and lifts dense open-vocabulary segmentation and region retrieval in proportion to baseline off-class spread.
MapDreamer synthesizes lane-level maps from aerial imagery via VAE latent encoding, transformer latent diffusion conditioned on aerial features, a lane cardinality module with ghost latents, and sliding-window graph aggregation, showing improved fidelity on UrbanLaneGraph data.
A3 is a learnable activation scaling module that trains on amplified adversarial signals via contrastive losses to improve robustness when the same parameters are used in attenuation mode.
Empirical study finds OV-OD robustness driven by vision backbone and image domain via layer-wise feature collapse analysis, validated with a low-parameter robustness improvement on real data.
ReMATF proposes a lightweight recurrent multi-scale network for atmospheric turbulence mitigation in dynamic videos that uses two-frame recurrent processing with motion-adaptive per-pixel fusion to enhance efficiency and coherence.
Frozen DINOv2-L features with k-NN classification and PCA/ICA refinement achieve state-of-the-art few-shot performance on four benchmarks without any backpropagation or fine-tuning.
A literature survey that categorizes high-level abstract concept image classification tasks in CV into semantic clusters and identifies persistent challenges and opportunities for hybrid AI approaches.
INAR-VL routes 36% of visual question answering requests to the edge using lightweight complexity signals, cutting latency 24% and energy 26% while retaining 97% of cloud accuracy.
A systematic evaluation of GPU memory and utilization estimators across analytical, library-based, and ML paradigms identifies key limitations in generalization, integration overhead, and hardware variability for training-aware resource management.
citing papers explorer
-
Learning to Deny: Action Denial in Multimodal Large Language Models
MLLMs drop from over 85% accuracy on action presence to under 50% on matched action-denial videos, exposing a causal verification gap that causal graph prompts partially close.
-
Tailor Made Embeddings for Quantum Machine Learning
A variational autoencoder learns quantum embeddings compressing ImageNet into 13 qubits and achieving 98.5% accuracy on MNIST 3-vs-5 classification with a quantum circuit, close to classical baselines and far above naive amplitude embeddings.
-
Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation
An ILP-based oracle applied to seven VIS methods on YouTube-VIS and OVIS shows tracking instability as the dominant bottleneck, producing gaps exceeding 20 AP under occlusion while classification impact is secondary.
-
Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data
DHCNet improves ultra-fine-grained visual categorization by progressively building holistic cognition from local discrepancies using self-shuffling and refinement on limited data.
-
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.
-
Variational Feature Compression for Model-Specific Representations
A variational latent bottleneck with KL regularization and a dynamic binary mask based on saliency produces model-specific features that keep high accuracy for one classifier but drop others below 2% on CIFAR-100 with over 45x suppression.
-
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP
Inference-time α-entmax sparsification of CLIP’s final self-attention denoises diffuse mass and lifts dense open-vocabulary segmentation and region retrieval in proportion to baseline off-class spread.
-
MapDreamer: Aerial Imagery Conditioned Latent Diffusion for Lane-Level Map Generation
MapDreamer synthesizes lane-level maps from aerial imagery via VAE latent encoding, transformer latent diffusion conditioned on aerial features, a lane cardinality module with ghost latents, and sliding-window graph aggregation, showing improved fidelity on UrbanLaneGraph data.
-
Improving Adversarial Robustness via Activation Amplification and Attenuation
A3 is a learnable activation scaling module that trains on amplified adversarial signals via contrastive losses to improve robustness when the same parameters are used in attenuation mode.
-
Robust Onion: Peeling Open Vocab Object Detectors Under Noise
Empirical study finds OV-OD robustness driven by vision backbone and image domain via layer-wise feature collapse analysis, validated with a low-parameter robustness improvement on real data.
-
ReMATF: Recurrent Motion-Adaptive Multi-scale Turbulence Mitigation for Dynamic Scenes
ReMATF proposes a lightweight recurrent multi-scale network for atmospheric turbulence mitigation in dynamic videos that uses two-frame recurrent processing with motion-adaptive per-pixel fusion to enhance efficiency and coherence.
-
Rethinking the Good Enough Embedding for Easy Few-Shot Learning
Frozen DINOv2-L features with k-NN classification and PCA/ICA refinement achieve state-of-the-art few-shot performance on four benchmarks without any backpropagation or fine-tuning.
-
Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories
A literature survey that categorizes high-level abstract concept image classification tasks in CV into semantic clusters and identifies persistent challenges and opportunities for hybrid AI approaches.
-
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
INAR-VL routes 36% of visual question answering requests to the edge using lightweight complexity signals, cutting latency 24% and energy 26% while retaining 97% of cloud accuracy.
-
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
A systematic evaluation of GPU memory and utilization estimators across analytical, library-based, and ML paradigms identifies key limitations in generalization, integration overhead, and hardware variability for training-aware resource management.