WildBox provides over 237k 3D wildlife annotations from drone video and benchmarks reveal zero-shot 3D detection at 0 AP but fine-tuned performance of 8.68 AP-BEV and 13.17 AP3D, with depth estimation causing most errors.
super hub Mixed citations
PoseNet: A convolutional network for real-time 6-dof camera relocalization
Mixed citation behavior. Most common role is background (67%).
hub tools
citation-role summary
citation-polarity summary
claims ledger
- method single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distribu
- dataset (C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative
- background Available: https://api.semanticscholar.org/CorpusID:15559857 [84] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, "Recognizing text with perspective distortion in natural scenes," in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569-576. [Online]. Available: https://doi.org/10.1109/ICCV .2013.76 [85] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, "Toward understanding wordart: Corner-guided transformer for scene text recognition," in Pr
- method . Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulat
- method To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridiss
- background Class Activation Mapping methods were introduced to ad- dress the deployment trust gap by making CNN spatial rea- soning visible and auditable [12]. However, a series of foundational studies has revealed that CAM methods them- selves suffer from reliability failures that are independent of - and invisible to - classification performance met- rics. Model Parameter Randomization[13]: Several widely used explanation methods produce nearly identical heatmaps for a fully trained model and for a model
authors
co-cited works
representative citing papers
MoHallBench is a new benchmark evaluating motion hallucination in VideoLLMs from co-occurrence priors, sequential inference, and similarity confusion, revealing decoupling from action recognition performance.
MATCH is the first flow matching method for multi-view anomaly detection, reporting SOTA results on Real-IAD and the first comprehensive evaluation on MANTA-Tiny while enabling real-time use by omitting the divergence term.
Signed pairwise interaction scores conflate U/R/S; Stochastic Hi-Fi uses interventional masked inference to recover per-feature uniqueness, redundancy, and synergy profiles.
AdaVoMP predicts accurate dense spatially-varying Young's modulus, Poisson's ratio and density for 3D objects using an adaptive sparse voxel structure generated by a sparse transformer encoder-decoder at 16^3 higher resolution than prior fixed-voxel methods.
SpikeTAD proposes the first SNN-based end-to-end TAD model, reporting 67.2% mAP on THUMOS14 and 37.42% on ActivityNet-1.3 with extremely low power consumption.
An ILP-based oracle applied to seven VIS methods on YouTube-VIS and OVIS shows tracking instability as the dominant bottleneck, producing gaps exceeding 20 AP under occlusion while classification impact is secondary.
Smaller self-supervised ViTs localize objects better via attention than larger ViTs, enabling A² to decouple localization from feature extraction for competitive performance on distribution-shifted benchmarks.
A new quality-guided approach for semi-supervised medical image segmentation that trains a predictor on synthetic errors to enhance pseudolabel handling.
A 3D-aware framework uses SAM3D geometry and pose estimation plus geodesic filtering to supervise a lightweight adapter on DINO and Stable Diffusion features, improving semantic correspondence with less manual supervision.
Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.
Morpheus learns morphable category-level shape priors to produce implicit 3D correspondences in camera space without explicit supervision and releases the HouseCorr3D benchmark with amodal and symmetry annotations.
Introduces NMCA-aligned L1/L2 LULC schemes and the Loosdorf-MSL benchmark dataset, with Point Transformer V3 reaching 79.4% mIoU on 8 classes and 58.9% on 20 classes, plus gains from multispectral inputs.
HyperDn is a configuration-conditioned predictor that transfers oracle supervision across denoising paradigms to achieve near-oracle hyperparameter prediction with few or zero target labels.
MulTaBench is a new collection of 40 image-tabular and text-tabular datasets designed to test target-aware representation tuning in multimodal tabular models.
AnomalyClaw turns single-step VLM anomaly judgments into a multi-round tool-grounded refutation process, delivering consistent macro-AUROC gains of 3.5-7.9 percentage points over direct inference across 12 cross-domain datasets.
Urban-ImageNet is a 2-million-image multi-modal dataset with HUSIC 10-class taxonomy enabling benchmarks for urban scene classification, cross-modal retrieval, and instance segmentation.
Cross3R performs feed-forward 3D reconstruction and 6-DoF pose estimation from any combination of satellite, UAV, and ground images, outperforming baselines on a new 278K-image tri-view dataset.
Text-guided class-agnostic counting models exhibit significant weaknesses in grounding textual prompts to visual objects, as demonstrated by new negative-label and distractor tests on a multi-category dataset.
NeuralEmu uses machine learning trained on real 5G telemetry to predict resource blocks and modulation for multiple users, cutting emulation error by 51-57% versus prior tools for web, video, and gaming metrics.
A fitted iso-depth scaling law measures that one recurrence in looped transformers is worth r^0.46 unique blocks in validation loss.
Concept Graph Convolutions perform message passing on node concepts to increase interpretability of graph neural networks without losing task performance.
DHCNet improves ultra-fine-grained visual categorization by progressively building holistic cognition from local discrepancies using self-shuffling and refinement on limited data.
The C-Score quantifies intra-class explanation consistency for CAM methods via confidence-weighted pairwise soft IoU and detects AUC-consistency dissociation as an early warning for model instability on chest X-ray classification.
citing papers explorer
-
WildBox: A Dataset and Benchmark for Aerial Monocular 3D Detection of African Savanna Wildlife
WildBox provides over 237k 3D wildlife annotations from drone video and benchmarks reveal zero-shot 3D detection at 0 AP but fine-tuned performance of 8.68 AP-BEV and 13.17 AP3D, with depth estimation causing most errors.
-
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
MoHallBench is a new benchmark evaluating motion hallucination in VideoLLMs from co-occurrence priors, sequential inference, and similarity confusion, revealing decoupling from action recognition performance.
-
MATCH: Flow Matching for Multi-View Anomaly Detection
MATCH is the first flow matching method for multi-view anomaly detection, reporting SOTA results on Real-IAD and the first comprehensive evaluation on MANTA-Tiny while enabling real-time use by omitting the divergence term.
-
The Representational Limit of Scalar Interactions: An Interventional Decomposition
Signed pairwise interaction scores conflate U/R/S; Stochastic Hi-Fi uses interventional masked inference to recover per-feature uniqueness, redundancy, and synergy profiles.
-
Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
AdaVoMP predicts accurate dense spatially-varying Young's modulus, Poisson's ratio and density for 3D objects using an adaptive sparse voxel structure generated by a sparse transformer encoder-decoder at 16^3 higher resolution than prior fixed-voxel methods.
-
SpikeTAD: Spiking Neural Networks for End-to-End Temporal Action Detection
SpikeTAD proposes the first SNN-based end-to-end TAD model, reporting 67.2% mAP on THUMOS14 and 37.42% on ActivityNet-1.3 with extremely low power consumption.
-
Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation
An ILP-based oracle applied to seven VIS methods on YouTube-VIS and OVIS shows tracking instability as the dominant bottleneck, producing gaps exceeding 20 AP under occlusion while classification impact is secondary.
-
$A^2$: Smaller Self-Supervised ViTs Localize Better than Larger Ones
Smaller self-supervised ViTs localize objects better via attention than larger ViTs, enabling A² to decouple localization from feature extraction for competitive performance on distribution-shifted benchmarks.
-
Quality-Guided Semi-Supervised Learning for Medical Image Segmentation
A new quality-guided approach for semi-supervised medical image segmentation that trains a predictor on synthetic errors to enhance pseudolabel handling.
-
Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
A 3D-aware framework uses SAM3D geometry and pose estimation plus geodesic filtering to supervise a lightweight adapter on DINO and Stable Diffusion features, improving semantic correspondence with less manual supervision.
-
Brain-IT-VQA: From Brain Signals to Answers
Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.
-
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
Morpheus learns morphable category-level shape priors to produce implicit 3D correspondences in camera space without explicit supervision and releases the HouseCorr3D benchmark with amodal and symmetry annotations.
-
3D LULC classification using multispectral LiDAR and deep learning: current and prospective schemes
Introduces NMCA-aligned L1/L2 LULC schemes and the Loosdorf-MSL benchmark dataset, with Point Transformer V3 reaching 79.4% mIoU on 8 classes and 58.9% on 20 classes, plus gains from multispectral inputs.
-
Oracle Supervision Transfers for Hyperparameter Prediction in Model-Based Image Denoising
HyperDn is a configuration-conditioned predictor that transfers oracle supervision across denoising paradigms to achieve near-oracle hyperparameter prediction with few or zero target labels.
-
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
MulTaBench is a new collection of 40 image-tabular and text-tabular datasets designed to test target-aware representation tuning in multimodal tabular models.
-
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation
AnomalyClaw turns single-step VLM anomaly judgments into a multi-round tool-grounded refutation process, delivering consistent macro-AUROC gains of 3.5-7.9 percentage points over direct inference across 12 cross-domain datasets.
-
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
Urban-ImageNet is a 2-million-image multi-modal dataset with HUSIC 10-class taxonomy enabling benchmarks for urban scene classification, cross-modal retrieval, and instance segmentation.
-
Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images
Cross3R performs feed-forward 3D reconstruction and 6-DoF pose estimation from any combination of satellite, UAV, and ground images, outperforming baselines on a new 278K-image tri-view dataset.
-
Does it Really Count? Assessing Semantic Grounding in Text-Guided Class-Agnostic Counting
Text-guided class-agnostic counting models exhibit significant weaknesses in grounding textual prompts to visual objects, as demonstrated by new negative-label and distractor tests on a multi-category dataset.
-
NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation
NeuralEmu uses machine learning trained on real 5G telemetry to predict resource blocks and modulation for multiple users, cutting emulation error by 51-57% versus prior tools for web, video, and gaming metrics.
-
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
A fitted iso-depth scaling law measures that one recurrence in looped transformers is worth r^0.46 unique blocks in validation loss.
-
Concept Graph Convolutions: Message Passing in the Concept Space
Concept Graph Convolutions perform message passing on node concepts to increase interpretability of graph neural networks without losing task performance.
-
Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data
DHCNet improves ultra-fine-grained visual categorization by progressively building holistic cognition from local discrepancies using self-shuffling and refinement on limited data.
-
Quantifying Explanation Consistency: The C-Score Metric for CAM-Based Explainability in Medical Image Classification
The C-Score quantifies intra-class explanation consistency for CAM methods via confidence-weighted pairwise soft IoU and detects AUC-consistency dissociation as an early warning for model instability on chest X-ray classification.
-
Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning
Multimodal contrastive learning using multilinear products is fragile to single bad modalities, and a gated version improves top-1 retrieval accuracy on synthetic and real trimodal data.
-
Contour Refinement using Discrete Diffusion in Low Data Regime
A CNN-based discrete diffusion method refines sparse contours from segmentation masks using simplified denoising steps and minimal post-processing, outperforming baselines on small medical and environmental datasets while running 3.5 times faster.
-
Transfer-learned Kolosov-Muskhelishvili Informed Neural Networks for Fracture Mechanics
A Kolosov-Muskhelishvili informed neural network satisfies plane elasticity equations by construction, achieves sub-1% errors on benchmarks, and uses transfer learning to predict crack paths under multiple criteria with over 70% less training time.
-
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
BadVSFM is the first effective backdoor attack on prompt-driven video segmentation foundation models, using a two-stage encoder-decoder strategy to achieve high attack success rates with limited clean performance loss.
-
Combined Hyperbolic and Euclidean Soft Triple Loss Beyond the Single Space Deep Metric Learning
CHEST loss combines proxy-based soft triple losses in hyperbolic and Euclidean spaces with hyperbolic hierarchical clustering regularization, improving deep metric learning accuracy and stability while achieving new state-of-the-art results on four benchmark datasets.
-
Effective Model Pruning: Measure The Redundancy of Model Components
EMP maps importance scores to effective sample size N_eff and prunes the lowest N - N_eff components, with a derived lower bound on retained effective mass and upper bound on loss increase.
-
BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations
BEVCALIB performs LiDAR-camera calibration from raw data by fusing camera and LiDAR bird's-eye view features with a novel feature selector and reports state-of-the-art accuracy on KITTI and NuScenes.
-
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
OCRBench v2 is a new benchmark with four times more tasks than prior versions that reveals most large multimodal models score below 50 out of 100 on visual text tasks and share five specific weaknesses.
-
Pose Estimation for Non-Cooperative Rendezvous Using Neural Networks
SPN is a CNN that detects a spacecraft bounding box, classifies then regresses attitude, and optimizes position via Gauss-Newton, achieving degree-level attitude and cm-level position errors on real images after training only on synthetic data.
-
BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Adversaries can create backdoored neural networks during outsourced training that maintain high accuracy on normal data but misbehave on attacker-chosen triggers.
-
Computer vision-based neural networks for radioisotope identification in urban environments
A CNN trained on time-stacked waterfall spectrograms beats non-negative matrix factorization for radioisotope identification at one false alarm per hour, but not at stricter false-alarm rates.
-
Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference
A confidence-feedback-weighted graph matching network achieves 96.36% F1-score on damage site matching by using matchability confidence to weight edge features and applying geometric consistency and hard-example mining.
-
Drag, Infer, Reproject: Grounding LLMs through Spatial Interaction for Image Clustering
CriterionSI infers clustering criteria from sequential user drags via LLMs to produce progressively aligned image cluster layouts.
-
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
OSOR is a one-step diffusion inpainting method using an occupancy-guided discriminator, alpha head, and semantic-anchored verification pipeline to achieve effect-aware object removal, outperforming multi-step baselines in quality at 4-30x speed.
-
Improving Richardson--Lucy Deconvolution with Diffusion Priors for Fluorescence Microscopy
Integrates a diffusion prior into Richardson-Lucy deconvolution to reduce noise amplification and better preserve filamentous and punctate structures in low-photon fluorescence microscopy.
-
From Pixels to Concepts: Growing Rich 3D Semantic Scene Graph Forests utilizing Foundation Models
Uses VLMs to detect instance concepts and LLMs to infer abstract relationships, assembling them into 3D scene graph forests that are evaluated on uHumans2 and ScanNet and tested in open-vocabulary retrieval on a Spot robot.
-
Physics-Guided Spatiotemporal State Space Modeling for Lookahead Molten Pool Segmentation in Laser Wire-Feed Welding
WeldMamba achieves 74.63% mIoU for 500 ms lookahead segmentation of keyhole, wire, and molten pool using spatiotemporal state space modeling conditioned on welding signals and physics-based losses on a 43-sequence dataset.
-
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
Moebius introduces a compressed diffusion inpainting model using Local-λ Mix Interaction blocks and latent-space multi-granularity distillation to reach 10B-level quality with 0.22B parameters.
-
APT: Atomic Physical Transitions for Causal Video-Language Understanding
Introduces APT chains as ordered causal transition sequences and APT-Tune to improve VLM transition detection while preserving event-level performance.
-
MotionPyramid: Hierarchical Motion Representation and Residual Interfaces
MotionPyramid learns a stack of latent decoders from motion tracking data to create multi-resolution action interfaces for RL policies in humanoid control, with residual interfaces allowing coarse programs and fine corrections to coexist.
-
AdaCodec: A Predictive Visual Code for Video MLLMs
AdaCodec introduces a predictive visual code that cuts visual token use in video MLLMs by sending full frames only on high predictive cost and otherwise encoding inter-frame changes as P-tokens, yielding better benchmark scores at lower budgets.
-
Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization
Sampling pairs directly with auxiliary information for higher inclusion probabilities on informative pairs yields near-full pairwise loss performance at reduced computational cost.
-
RefDiffNet: Learning to Expose Subtle PCB Defects Before Detection
RefDiffNet is a lightweight input enhancement block that uses reference image comparison to expose PCB defects, delivering up to 18% relative mAP50:95 gains across YOLO, RT-DETR, and Faster R-CNN detectors with 0.004-0.005M extra parameters.
-
NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving
NTR adds a self-distillation masked latent reconstruction objective that uses only scene tokens to reconstruct masked patch features, improving visual representation quality and planning performance in end-to-end autonomous driving.
-
MuNet: A Mutualistic Network for Joint 3D Human Mesh Recovery and 3D Clothed Human Reconstruction from Single Images
MuNet is an end-to-end graph convolutional network using 2-manifold graphs and a mutualistic training mechanism that jointly optimizes 3D human mesh recovery and clothed reconstruction, reporting state-of-the-art results on six benchmarks.
-
ARCANE-PedSynth: Synthetic Multi-Pedestrian Datasets with Behavioural Crossing Annotations
ARCANE-PedSynth is a CARLA-based framework that generates synthetic multi-pedestrian datasets with behavioral crossing annotations by using hybrid AI-manual control to raise crossing rates and a 12-state FSM for diverse behaviors.