Text-guided class-agnostic counting models exhibit significant weaknesses in grounding textual prompts to visual objects, as demonstrated by new negative-label and distractor tests on a multi-category dataset.
hub
Generating 3d adversarial point clouds, in: 2019 IEEE/CVF Conference on Computer Vision and PatternRecognition(CVPR)
18 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
CROWD is a new global dataset of 51,753 continuous urban dashcam segments spanning over 20,000 hours from 238 countries, with manual labels and automated object detections for routine driving analysis.
A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.
TCG-AR is a real-time multi-view AR system for trading card games using only commodity RGB cameras and synthetic training data.
Mahalanobis PatchCore adds covariance-aware whitening and incremental streaming aggregation to PatchCore, preserving benchmark performance while cutting peak memory from 5.41 GB to 2.78 GB and raising mean industrial AUC from 0.981 to 0.986.
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
MAPR improves adversarial robustness in 3D point cloud networks by aligning latent predictions with intrinsic manifold geometry via curvature/diffusion features and a consistency loss.
ProtoFair uses prototype clustering to form pseudo-counterfactual pairs for an additive fairness contrastive loss that improves fairness metrics on CelebA and UTKFace while preserving accuracy in SimCLR and SupCon.
A multimodal adversarial attack using stage-sampled image nullification and cross-attention flattening degrades lip-sync and facial dynamics in Hallo-based talking-head generation.
Recent LiDAR 3D detectors remain as vulnerable to adversarial attacks as predecessors, with voxel-based and non-anchor-based models showing greater susceptibility under a multi-factor robustness framework.
Aligning optical-flow and motion-magnified features to hierarchical AU text anchors, then reliability-aware complementary fusion, yields strong multimodal micro-expression recognition across five benchmarks.
The paper releases SignNet-1M, a 1M-scale augmented dataset for ASL, CSL and DGS with 3DGS and diffusion-based variations, plus benchmarks showing improved cross-shift generalization.
Derives an asymptotic equivalent for the Representation Gap in equivariant diffusion models, showing it depends primarily on the intrinsic dimension of the task.
Zebrafish tectal subcircuits are dissociated into spike-efficient information gating and feedback-like robustness stabilization, then transferred to improve ResNet efficiency and noise tolerance.
SA-HGNN with contrastive learning improves power outage prediction by modeling spatial effects of extreme weather on infrastructure across multiple utility territories.
TMVA4D uses CNN and ConvLSTM encoders on multi-view 2D projections of 4D radar point clouds for semantic segmentation of people, reporting Dice 75.9% and IoU 61.2% in field tests.
MAPE combines a channel-attention U-Net (SAPE) trained on multi-model adversarial examples scheduled by PPSA to eliminate perturbations, reporting over 95.1% average defense on CIFAR-10 and 71.5% on Mini-ImageNet against black-box transferable attacks.
Adding a P2 branch to YOLOX-Nano raises small-object AP by 31.10% on VisDrone; QIEA screens structures balancing accuracy, FLOPs, latency, memory and recall.
citing papers explorer
-
Does it Really Count? Assessing Semantic Grounding in Text-Guided Class-Agnostic Counting
Text-guided class-agnostic counting models exhibit significant weaknesses in grounding textual prompts to visual objects, as demonstrated by new negative-label and distractor tests on a multi-category dataset.
-
A global dataset of continuous urban dashcam driving
CROWD is a new global dataset of 51,753 continuous urban dashcam segments spanning over 20,000 hours from 238 countries, with manual labels and automated object detections for routine driving analysis.
-
`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation
A frozen SAM2 backbone with adaptive token selection and symmetric KL clustering achieves competitive self-supervised video object segmentation by aligning soft part assignments across time.
-
TCG-AR: Real-Time Multi-View Augmented Reality for Trading Card Game Streaming
TCG-AR is a real-time multi-view AR system for trading card games using only commodity RGB cameras and synthetic training data.
-
Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection
Mahalanobis PatchCore adds covariance-aware whitening and incremental streaming aggregation to PatchCore, preserving benchmark performance while cutting peak memory from 5.41 GB to 2.78 GB and raising mean industrial AUC from 0.981 to 0.986.
-
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
-
Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness
MAPR improves adversarial robustness in 3D point cloud networks by aligning latent predictions with intrinsic manifold geometry via curvature/diffusion features and a consistency loss.
-
ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs
ProtoFair uses prototype clustering to form pseudo-counterfactual pairs for an additive fairness contrastive loss that improves fairness metrics on CelebA and UTKFace while preserving accuracy in SimCLR and SupCon.
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
A multimodal adversarial attack using stage-sampled image nullification and cross-attention flattening degrades lip-sync and facial dynamics in Hallo-based talking-head generation.
-
Comprehensive Robustness Analysis of LiDAR-based 3D Object Detection in Autonomous Driving
Recent LiDAR 3D detectors remain as vulnerable to adversarial attacks as predecessors, with voxel-based and non-anchor-based models showing greater susceptibility under a multi-factor robustness framework.
-
SAC$^2$-Net: Semantic Anchoring and Complementary-Consensus Fusion for Multimodal Micro-Expression Recognition
Aligning optical-flow and motion-magnified features to hierarchical AU text anchors, then reliability-aware complementary fusion, yields strong multimodal micro-expression recognition across five benchmarks.
-
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
The paper releases SignNet-1M, a 1M-scale augmented dataset for ASL, CSL and DGS with 3DGS and diffusion-based variations, plus benchmarks showing improved cross-shift generalization.
-
Representation Gap: Explaining the Unreasonable Effectiveness of Neural Networks from a Geometric Perspective
Derives an asymptotic equivalent for the Representation Gap in equivariant diffusion models, showing it depends primarily on the intrinsic dimension of the task.
-
Dual-axis attribution of zebrafish tectal microcircuits for energy-efficient and robust neurocomputing
Zebrafish tectal subcircuits are dissociated into spike-efficient information gating and feedback-like robustness stabilization, then transferred to improve ResNet efficiency and noise tolerance.
-
Empowering Power Outage Prediction with Spatially Aware Hybrid Graph Neural Networks and Contrastive Learning
SA-HGNN with contrastive learning improves power outage prediction by modeling spatial effects of extreme weather on infrastructure across multiple utility territories.
-
4D Radar Semantic Segmentation of People in Field Conditions Using Temporal Multi-View Networks
TMVA4D uses CNN and ConvLSTM encoders on multi-view 2D projections of 4D radar point clouds for semantic segmentation of people, reporting Dice 75.9% and IoU 61.2% in field tests.
-
MAPE: Defending Against Transferable Adversarial Attacks Using Multi-Source Adversarial Perturbations Elimination
MAPE combines a channel-attention U-Net (SAPE) trained on multi-model adversarial examples scheduled by PPSA to eliminate perturbations, reporting over 95.1% average defense on CIFAR-10 and 71.5% on Mini-ImageNet against black-box transferable attacks.
-
Edge-Constrained UAV Small-Object Detection with P2 Enhancement and Quantum-Inspired Lightweight Structure Search
Adding a P2 branch to YOLOX-Nano raises small-object AP by 31.10% on VisDrone; QIEA screens structures balancing accuracy, FLOPs, latency, memory and recall.