REVIEW 2 cited by
Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Tracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual information from image frames with the high temporal resolution of events, achieving high frame rate and robust point tracking under various challenging conditions. Specifically, we designed an Evolution Fusion module (EvoFusion) to model the image generation process guided by events. This module can effectively integrate valuable information from both modalities operating at different frequencies. To achieve smoother point trajectories, we employed a transformer-based refinement strategy that updates the point's trajectories and features iteratively. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, particularly improving expected feature age by 24$\%$ on EDS datasets. Finally, we qualitatively validated the robustness of our algorithm in real driving scenarios using our custom-designed high-resolution image-event synchronization device. Our source code will be released at https://github.com/ljx1002/FE-TAP.
Forward citations
Cited by 2 Pith papers
-
ETAP: Event-based Tracking of Any Point
ETAP introduces the first event-only tracking-any-point network, trained on a new EventKubric synthetic dataset with a motion-invariance feature-alignment loss, and reports state-of-the-art results on the EDS and EC f...
-
MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking
MATE tracks any point from event cameras alone, using motion vectors extracted from time surfaces to guide matching, and reports higher accuracy and survival than video- and event-based baselines on four benchmarks.
Discussion (0). Continue with ORCID to comment.