REVIEW 3 cited by
Pedestrian Action Anticipation using Contextual Feature Fusion in Stacked RNNs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
One of the major challenges for autonomous vehicles in urban environments is to understand and predict other road users' actions, in particular, pedestrians at the point of crossing. The common approach to solving this problem is to use the motion history of the agents to predict their future trajectories. However, pedestrians exhibit highly variable actions most of which cannot be understood without visual observation of the pedestrians themselves and their surroundings. To this end, we propose a solution for the problem of pedestrian action anticipation at the point of crossing. Our approach uses a novel stacked RNN architecture in which information collected from various sources, both scene dynamics and visual features, is gradually fused into the network at different levels of processing. We show, via extensive empirical evaluations, that the proposed algorithm achieves a higher prediction accuracy compared to alternative recurrent network architectures. We conduct experiments to investigate the impact of the length of observation, time to event and types of features on the performance of the proposed method. Finally, we demonstrate how different data fusion strategies impact prediction accuracy.
Forward citations
Cited by 3 Pith papers
-
TrajFusionNet: Pedestrian Crossing Intention Prediction via Fusion of Sequential and Visual Trajectory Representations
A transformer model that fuses predicted pedestrian trajectories and vehicle speed with scene images achieves competitive accuracy and the lowest inference time on pedestrian crossing intention benchmarks.
-
EEvAct: Early Event-Based Action Recognition with High-Rate Two-Stream Spiking Neural Networks
A high-rate two-stream spiking network with a lightweight gated fusion unit achieves 94.9% on THU EACT-50 and enables early prediction within 100 ms.
-
Pedestrian Intention and Trajectory Prediction in Unstructured Traffic Using IDD-PeD
IDD-PeD is a new unstructured-traffic pedestrian dataset with 685K bounding boxes and 19 behavioral attributes, and eight intention plus four trajectory baselines all degrade on it.
Discussion (0). Sign in to comment.