REVIEW 12 cited by
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce Recurrent All-Pairs Field Transforms (RAFT), a new deep network architecture for optical flow. RAFT extracts per-pixel features, builds multi-scale 4D correlation volumes for all pairs of pixels, and iteratively updates a flow field through a recurrent unit that performs lookups on the correlation volumes. RAFT achieves state-of-the-art performance. On KITTI, RAFT achieves an F1-all error of 5.10%, a 16% error reduction from the best published result (6.10%). On Sintel (final pass), RAFT obtains an end-point-error of 2.855 pixels, a 30% error reduction from the best published result (4.098 pixels). In addition, RAFT has strong cross-dataset generalization as well as high efficiency in inference time, training speed, and parameter count. Code is available at https://github.com/princeton-vl/RAFT.
Forward citations
Cited by 12 Pith papers
-
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
WorldRoamBench is a new benchmark for interactive world models that evaluates four stability dimensions with custom metrics and finds no tested model performs reliably across all.
-
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
FactorJEPA splits a video prediction model into layout, agent, and interaction channels with a visibility gate, and a new DENSEWORLD dataset tests it on crowded Indian city scenes.
-
Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps
VLMM is a 3D map representation where each object carries a fused, uncertainty-aware motion attribute (language-based movability prior + observed geometric motion) that makes motion queries such as 'what is moving' an...
-
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
Spatio-temporal contrastive SAEs recover temporal coherence lost by hard TopK, improve action probes by +3.9% and retrieval by up to 2.8× R@1, and expose a monosemanticity metric artifact.
-
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
An audio-conditioned video animation model is pretrained on noisy auto-curated videos and fine-tuned on a few clean examples, achieving top synchronization scores on a new 48-class benchmark with only 1.9% additional ...
-
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
Motion-X++ provides 19.5M 3D whole-body pose annotations across 120.5K sequences with text, audio, video, and motion modalities.
-
GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction
A geometry-aware, training-free inference framework that refines pretrained video diffusion predictions with projected static history content and view-conditioned routing achieves fifth place on AI City Challenge Track 5.
-
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data
A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.
-
Adam SLAM - the last mile of camera calibration with 3DGS
Backpropagating the 3D Gaussian Splatting color loss into camera parameters, with an L2 loss, focal-length training, and a Hessian-based reparameterization, raises average PSNR by 0.4 dB over COLMAP calibration on ref...
-
HKT: A Biologically Inspired Framework for Modular Hereditary Knowledge Transfer in Neural Networks
HKT is a modular feature-level distillation method whose genetic attention residual improves compact vision models on optical flow, classification, and segmentation benchmarks.
-
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
CosmosAlign adapts Cosmos3-Nano with two-stage LoRA, medoid sample selection, and motion-adaptive blending, achieving first place (76.49) on the AI City Challenge 2026 Track 5 traffic video forecasting benchmark.
-
Uncertainty Aware Mapping for Vision-Based Underwater Robots
A vision-based underwater mapping pipeline that colors Voxblox TSDF maps with RAFT-Stereo depth confidence and replaces weight accumulation with a running average.
Discussion (0). Continue with ORCID to comment.