Pith. sign in

REVIEW 22 cited by

Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.08890 v3 pith:UT7QQAJE submitted 2023-02-17 cs.CV

classification cs.CV
keywords cameraseventvisionevent-basedmethodsresearchchangescomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Event cameras are bio-inspired sensors that capture the per-pixel intensity changes asynchronously and produce event streams encoding the time, pixel position, and polarity (sign) of the intensity changes. Event cameras possess a myriad of advantages over canonical frame-based cameras, such as high temporal resolution, high dynamic range, low latency, etc. Being capable of capturing information in challenging visual conditions, event cameras have the potential to overcome the limitations of frame-based cameras in the computer vision and robotics community. In very recent years, deep learning (DL) has been brought to this emerging field and inspired active research endeavors in mining its potential. However, there is still a lack of taxonomies in DL techniques for event-based vision. We first scrutinize the typical event representations with quality enhancement methods as they play a pivotal role as inputs to the DL models. We then provide a comprehensive survey of existing DL-based methods by structurally grouping them into two major categories: 1) image/video reconstruction and restoration; 2) event-based scene understanding and 3D vision. We conduct benchmark experiments for the existing methods in some representative research directions, i.e., image reconstruction, deblurring, and object recognition, to identify some critical insights and problems. Finally, we have discussions regarding the challenges and provide new perspectives for inspiring more research studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Visual Grounding from Event Cameras

    cs.CV 2025-09 conditional novelty 7.0 of 10

    Talk2Event provides 5,567 event-camera driving scenes, 13,458 objects, and 30,690 human-validated referring expressions labeled with appearance, status, relation-to-viewer, and relation-to-others attributes.

  2. Weaving Light and Time: Unified Harmonic-Geometric Representation Learning for Dense RGB-Event Parsing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Evita, a unified RGB-Event backbone with geometric rectification, spectral resonance, and transient routing, plus N-ImageNetV2 pretraining, reports SOTA dense parsing with better accuracy-latency trade-offs.

  3. A Hardware-Aware Open-Source Framework for Design Space Exploration of Mixed-Signal Spiking Neural Networks

    eess.SP 2026-07 conditional novelty 6.0 of 10

    An open-source PyTorch framework embeds calibrated floating-gate and ReRAM synapse non-idealities and mixed-signal neuron models directly into SNN training, enabling cross-layer design space exploration across accurac...

  4. EventTracer: Fast Path Tracing-based Event Stream Rendering

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A path-tracing renderer plus a learned spiking denoiser generates 1000 FPS event streams from 3D scenes and reportedly beats V2E and V2CE on Real2Sim tests.

  5. BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multi-modal semantic segmentation framework that processes RGB and non-RGB sensors separately, matches labels in two stages, and aligns cross-modal queries with a VAE refiner.

  6. Making Every Event Count: Balancing Data Efficiency and Accuracy in Event Camera Subsampling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A causal, density-based event subsampling method preserves classification accuracy better than random, spatial, temporal, event-count, and corner-based baselines in sparse regimes, except when event counts vary widely...

  7. Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.

  8. Neural Ganglion Sensors: Learning Task-specific Event Cameras Inspired by the Neural Circuit of the Human Retina

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Learned spatial event kernels, inspired by retinal ganglion cells, improve the performance-versus-bandwidth tradeoff of event cameras on simulated video interpolation and optical flow.

  9. Dynamic EventNeRF: Reconstructing General Dynamic Scenes from Multi-view RGB and Event Streams

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Dynamic EventNeRF reconstructs dynamic 4D scenes from sparse multi-view event cameras and sparse RGB frames using time-conditioned, multi-segment NeRF models with event-based losses.

  10. EventGPT: Event Stream Understanding with Multimodal Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EventGPT adapts a LLaVA-style MLLM to event camera streams via three-stage training (image-language, event-language, instruction tuning) and outperforms RGB-based MLLMs on its own benchmark.

  11. A Survey of 3D Reconstruction with Event Cameras

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dedicated survey categorizes event-based 3D reconstruction methods by input setup and reconstruction strategy, and catalogs datasets, metrics, and open challenges.

  12. MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection

    cs.CV 2024-12 reject novelty 5.0 of 10

    MAGIC++ trains a semantic segmentation backbone with all available sensors and then uses the plain backbone at test time, reporting strong average results on arbitrary sensor combinations.

  13. Expanding Event Modality Applications through a Robust CLIP-Based Encoder

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A CLIP-based encoder for event cameras, trained with contrastive, consistency, and KL losses, improves zero-shot and few-shot object recognition and extends to video anomaly detection and cross-modal retrieval.

  14. SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation

    cs.RO 2025-08 unverdicted novelty 4.0 of 10

    The paper contributes an open-source v2e-based ROS/Gazebo event camera simulator and reports that event-based transformer policies trained by behavior cloning match or beat RGB-based policies in simulated navigation a...

  15. Edge Intelligence with Spiking Neural Networks

    cs.DC 2025-07 conditional novelty 4.0 of 10

    A comprehensive review of spiking neural networks for edge computing, covering neuron models, learning algorithms, hardware, deployment, security, and evaluation, with a claim to be the first survey on this specific i...

  16. MLLMs are Deeply Affected by Modality Bias

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.

  17. EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    EGFormer dynamically scores and drops the least useful sensor modality at each processing stage, cutting parameters by up to 91 percent and GFLOPs by half while keeping segmentation accuracy competitive.

  18. Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization

    cs.CV 2025-05 reject novelty 4.0 of 10

    A plug-and-play functional-entropy regularizer applied at feature and prediction scales is claimed to reduce unimodal bias in multi-modal semantic segmentation, with large mIoU gains on MUSES, DELIVER, and MCubeS with...

  19. Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection

    cs.CV 2025-05 conditional novelty 4.0 of 10

    IEF-VAD fuses CLIP image and synthetic-event features via learned inverse-variance weighting with Kalman-style updates and iterative refinement, reporting state-of-the-art AUC/AP on UCF-Crime, XD-Violence, ShanghaiTec...

  20. From Events to Enhancement: A Survey on Event-Based Imaging Technologies

    cs.CV 2025-04 conditional novelty 4.0 of 10

    A survey that organizes event-based imaging into enhancement tasks and advanced light-recovery tasks, using a plenoptic light-ray model as the common foundation.

  21. Segment Any RGB-Thermal Model with Language-aided Distillation

    cs.CV 2025-05 conditional novelty 3.0 of 10

    SARTM fine-tunes SAM2 with LoRA and distills CLIP text knowledge to improve RGB-thermal semantic segmentation, reporting top mIoU on PST900, MFNet, and FMB.

  22. Event Camera Guided Visual Media Restoration & 3D Reconstruction: A Survey

    cs.CV 2025-09 conditional novelty 1.0 of 10

    A structured survey of event-camera-guided video restoration and 3D reconstruction, organized by temporal, spatial, and 3D tasks.

Pith tools