MM-Eval unifies evaluation of multimodal summaries by integrating factual text quality, cross-modal relevance via MLLM judge, and visual diversity via truncated CLIP entropy, then calibrates their combination on human preferences.
Distinctive image features from scale-invariant keypoints
19 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
A confidence-feedback-weighted graph matching network achieves 96.36% F1-score on damage site matching by using matchability confidence to weight edge features and applying geometric consistency and hard-example mining.
Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.
SierpinskiCam adds Sierpinski dome texture cues and negative-RoPE reference video conditioning to geometry-guided video diffusion to improve camera controllability and consistency in video retaking.
A decoupled pipeline with YOLO detection, deterministic prompt encoding, and QLoRA-adapted 1.5B LLM achieves superior structured report generation compared to monolithic VLMs on synthetic maintenance data.
SPeCTrA-Sum uses hierarchical cross-modal fusion via DVP and DPP-distilled image selection via VRP to generate more accurate and visually grounded multimodal summaries.
Spike encoders are reformulated as time-causal bandpass wavelets that preserve sparsity and locality while providing reconstruction error bounds comparable to continuous wavelet transforms on ECG and audio signals.
Sphere clouds neutralize density attacks on private 3D maps for visual localization while depth guidance from ToF sensors restores translation scale for accurate pose estimation.
AV1 motion vectors filtered by cosine consistency yield dense sub-pixel correspondences that support structure-from-motion on short video clips with lower CPU cost and higher match density than sequential SIFT.
An adaptive multi-criteria feature selection policy for 3D scene reconstruction that outperforms random, texture-only, and uniform-grid baselines on synthetic multi-view tests for completeness and RMSE.
STAR-IOD applies scale-decoupled topology alignment and K-Means-based pseudo-label refinement to reduce catastrophic forgetting in remote sensing incremental object detection, reporting 1.7% and 2.1% mAP gains on new DIOR-IOD and DOTA-IOD datasets.
A refraction-aware neural radiance field pipeline reconstructs underwater point clouds from UAV imagery, reducing mean error from ~0.5 m to ~0.06 m on a simulated scene by tracing bent light rays.
Training-inference input alignment outweighs framework choice for longitudinal retinal image prediction, with deterministic regression matching complex models when acquisition variability dominates disease progression.
OpenPRC provides a schema-driven framework with five modules for GPU physics simulation, experimental vision ingestion, reservoir learning, information analysis, and physics-aware optimization to enable consistent PRC evaluation from simulations and real experiments.
CNN separation of He I 10830Å chromospheric signal from photospheric contamination in quiet Sun reveals R ≈ -0.84 anti-correlation with 304Å and magnetic-field-dependent EUV coupling.
Automated pipeline identifies textures in sports videos, builds 3D scene models, places ads in 3D, projects them to 2D, and tracks with homography for seamless augmentation.
Sparse graph topologies (especially minimum spanning trees) match or exceed denser graphs for image classification on Fashion-MNIST when using a fixed three-layer GCN.
The paper experimentally compares SIFT and ORB on GPS-annotated satellite image tiles by measuring inlier ratios after RANSAC homography estimation and analyzing how the number of keypoints affects matching quality.
citing papers explorer
-
Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity
MM-Eval unifies evaluation of multimodal summaries by integrating factual text quality, cross-modal relevance via MLLM judge, and visual diversity via truncated CLIP entropy, then calibrates their combination on human preferences.
-
Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference
A confidence-feedback-weighted graph matching network achieves 96.36% F1-score on damage site matching by using matchability confidence to weight edge features and applying geometric consistency and hard-example mining.
-
Lighting-Consistent Object Transfer Across Radiance Fields
Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.
-
SierpinskiCam: Camera-Controlled Video Retaking with Sierpinski Triangle Pattern Cues
SierpinskiCam adds Sierpinski dome texture cues and negative-RoPE reference video conditioning to geometry-guided video diffusion to improve camera controllability and consistency in video retaking.
-
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
A decoupled pipeline with YOLO detection, deterministic prompt encoding, and QLoRA-adapted 1.5B LLM achieves superior structured report generation compared to monolithic VLMs on synthetic maintenance data.
-
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention
SPeCTrA-Sum uses hierarchical cross-modal fusion via DVP and DPP-distilled image selection via VRP to generate more accurate and visually grounded multimodal summaries.
-
Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets
Spike encoders are reformulated as time-causal bandpass wavelets that preserve sparsity and locality while providing reconstruction error bounds comparable to continuous wavelet transforms on ECG and audio signals.
-
Depth-Guided Privacy-Preserving Visual Localization Using 3D Sphere Clouds
Sphere clouds neutralize density attacks on private 3D maps for visual localization while depth guidance from ToF sensors restores translation scale for accurate pose estimation.
-
Leveraging AV1 motion vectors for Fast and Dense Feature Matching
AV1 motion vectors filtered by cosine consistency yield dense sub-pixel correspondences that support structure-from-motion on short video clips with lower CPU cost and higher match density than sequential SIFT.
-
Feature-Optimized Vision for Adaptive 3D Scene Reconstruction
An adaptive multi-criteria feature selection policy for 3D scene reconstruction that outperforms random, texture-only, and uniform-grid baselines on synthetic multi-view tests for completeness and RMSE.
-
STAR-IOD: Scale-decoupled Topology Alignment with Pseudo-label Refinement for Remote Sensing Incremental Object Detection
STAR-IOD applies scale-decoupled topology alignment and K-Means-based pseudo-label refinement to reduce catastrophic forgetting in remote sensing incremental object detection, reporting 1.7% and 2.1% mAP gains on new DIOR-IOD and DOTA-IOD datasets.
-
BathyFacto: Refraction-Aware Two-Media Neural Radiance Fields for Bathymetry
A refraction-aware neural radiance field pipeline reconstructs underwater point clouds from UAV imagery, reducing mean error from ~0.5 m to ~0.06 m on a simulated scene by tracing bent light rays.
-
Training-inference input alignment outweighs framework choice in longitudinal retinal image prediction
Training-inference input alignment outweighs framework choice for longitudinal retinal image prediction, with deterministic regression matching complex models when acquisition variability dominates disease progression.
-
OpenPRC: A Unified Open-Source Framework for Physics-to-Task Evaluation in Physical Reservoir Computing
OpenPRC provides a schema-driven framework with five modules for GPU physics simulation, experimental vision ingestion, reservoir learning, information analysis, and physics-aware optimization to enable consistent PRC evaluation from simulations and real experiments.
-
Machine Learning-based Separation of the He I 10830{\AA} Chromospheric Signal: Quantitative Analysis of Chromosphere-Corona Intensity in the Quiet Sun
CNN separation of He I 10830Å chromospheric signal from photospheric contamination in quiet Sun reveals R ≈ -0.84 anti-correlation with 304Å and magnetic-field-dependent EUV coupling.
-
Markerless Augmented Advertising for Sports Videos
Automated pipeline identifies textures in sports videos, builds 3D scene models, places ads in 3D, projects them to 2D, and tracks with homography for seamless augmentation.
-
Visual graphs for image classification: does the structure affect performance?
Sparse graph topologies (especially minimum spanning trees) match or exceed denser graphs for image classification on Fashion-MNIST when using a fixed three-layer GCN.
-
Mathematical Analysis of Image Matching Techniques
The paper experimentally compares SIFT and ORB on GPS-annotated satellite image tiles by measuring inlier ratios after RANSAC homography estimation and analyzing how the number of keypoints affects matching quality.
- SIFT-VTON: Geometric Correspondence Supervision on Cross-Attention for Virtual Try-On