Predicting quantized latent residuals (Latent Drift with FSQ) avoids identity collapse and noise interpolation, improving patient-specific 3D MRI neuro-forecasting over diffusion and autoregressive baselines.
In: Proceedings of the IEEE/CVF international confer- ence on computer vision
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 4years
2026 4roles
method 1polarities
use method 1representative citing papers
A multi-view transformer framework integrating rotational micro-ultrasound sweeps with biopsy frames achieves 87.2% patient-level AUROC for prostate cancer detection, outperforming single-frame and video baselines.
Hi-GaTA is a hierarchical gated temporal aggregation adapter that uses short-to-long temporal pyramids and gated fusion to enable surgical video report generation, backed by a new 214-video benchmark and a surgical ViViT pretrained on 40,000 minutes of video.
Logic Gate Networks produce compact Boolean-circuit descriptors for video copy detection that match or exceed prior accuracy at over 11k inferences per second and orders-of-magnitude smaller size.
citing papers explorer
-
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
Predicting quantized latent residuals (Latent Drift with FSQ) avoids identity collapse and noise interpolation, improving patient-specific 3D MRI neuro-forecasting over diffusion and autoregressive baselines.
-
Compass: Prostate Cancer Detection Needs Multi-View Context
A multi-view transformer framework integrating rotational micro-ultrasound sweeps with biopsy frames achieves 87.2% patient-level AUROC for prostate cancer detection, outperforming single-frame and video baselines.
-
Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation
Hi-GaTA is a hierarchical gated temporal aggregation adapter that uses short-to-long temporal pyramids and gated fusion to enable surgical video report generation, backed by a new 214-video benchmark and a surgical ViViT pretrained on 40,000 minutes of video.
-
Efficient Logic Gate Networks for Video Copy Detection
Logic Gate Networks produce compact Boolean-circuit descriptors for video copy detection that match or exceed prior accuracy at over 11k inferences per second and orders-of-magnitude smaller size.