Pith. sign in

REVIEW 3 cited by

IntFormer: Predicting pedestrian intention with the aid of the Transformer architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.08647 v1 pith:QVY5ZLBY submitted 2021-05-18 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelarchitecturecalledcrossingintformerpedestriantransformerapprox
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Understanding pedestrian crossing behavior is an essential goal in intelligent vehicle development, leading to an improvement in their security and traffic flow. In this paper, we developed a method called IntFormer. It is based on transformer architecture and a novel convolutional video classification model called RubiksNet. Following the evaluation procedure in a recent benchmark, we show that our model reaches state-of-the-art results with good performance ($\approx 40$ seq. per second) and size ($8\times $smaller than the best performing model), making it suitable for real-time usage. We also explore each of the input features, finding that ego-vehicle speed is the most important variable, possibly due to the similarity in crossing cases in PIE dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Reason-enriched annotations and a cross-modal vision-language model improve pedestrian crossing-intent prediction over prior methods.

  2. Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Gemini 2.5 Pro, prompted with 16-frame video clips and ego-vehicle speed, predicts pedestrian crossing intent at 73% accuracy on JAAD-beh without finetuning.

  3. Enhancing Customer Service Chatbots with Context-Aware NLU through Selective Attention and Multi-task Learning

    cs.LG 2025-06 reject novelty 4.0 of 10

    MTL-CNLU-SAWC uses query text plus order-status context and two training labels to boost top-2 intent accuracy by 4.8% over a text-only baseline on Walmart customer care data.

Pith tools