{"id":"1ca205a8-9268-4f1f-8832-7b0ce9cd1546","arxiv_id":"2505.05488","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes event-based imaging into enhancement tasks and advanced light-recovery tasks, using a plenoptic light-ray model as the common foundation.","lead":"This paper surveys how event cameras, which record pixel brightness changes instead of full frames, are used in imaging tasks such as video interpolation, deblurring, HDR, light fields, and 3D view synthesis. It proposes a taxonomy based on how much light information each task recovers and lists open problems for the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The physical model does not actually drive the taxonomy: tasks in Sec. 3 collapse to the same plenoptic subset, so the claimed unified foundation is under-specified.","rationale":"The reader's verdict was CONDITIONAL, with the weakest assumption being that the plenoptic function supports a meaningful scoring of imaging tasks. My stress-test pass confirms that this assumption is the load-bearing point, but sharpens it: the problem is not only whether Eq. 1 is a complete light representation; the survey never actually assigns each task a distinct recovered subset of Eq. 1. Several Sec. 3 tasks collapse to the same recovered intensity field, while photometric stereo falls outside the plenoptic-function axis entirely. This makes the claimed physical foundation under-specified rather than simply idealized. The event model of Eq. 6 is standard and sufficiently accurate for the survey's purposes, so I do not treat real-sensor deviations as the main risk. The concern reinforces the existing CONDITIONAL verdict: the survey remains useful as a task map, but the authors need to either demonstrate the mapping from Eq. 1 to their categories or soften the 'unified foundation' claim. Since the reader already conditioned on a related issue, no verdict change is needed.","tokens_in":15896,"tokens_out":7794,"duration_ms":91942,"concrete_test":"Construct a mapping table from each reviewed task (VFI, deblurring, RSC, VSR, LLE, HDR, NeRF, 3DGS, light field, photometric stereo) to the subset of (x,y,z,θ,φ,λ,τ) that its output recovers, using the papers' own formulations. If VFI, deblurring, and RSC occupy the same cell, and photometric stereo has no cell, the model is not actually organizing the field; the authors should either introduce a second axis (e.g., input degradation type) or revise the claim that the taxonomy is grounded in the physical model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the Sec. 2 physical model (plenoptic function, Eq. 1; frame integral, Eq. 5; event threshold, Eq. 6) supplies a foundation that systematically organizes event-based imaging. For this to hold, each task category should correspond to a distinct portion of the light signal that is recovered, and the model should separate enhancement from advanced imaging. The text never provides that mapping. VFI, deblurring, and rolling-shutter correction all recover the same intensity subset I(x,y,t) of Eq. 1; the categories differ only in the degradation of the input, not in the plenoptic content recovered. Photometric stereo (Sec. 4.3) outputs surface normals and lighting estimates, which are not a slice of L(x,y,z,θ,φ,λ,τ) as defined, so it falls outside the model's 'richer light information' axis. Multi-view generation and light-field reconstruction are both labeled richer, but for free-space radiance the (x,y,z) and (θ,φ) arguments of Eq. 1 are redundant along a ray, so 'richer' is not an unambiguous ordering. Consequently, the taxonomy is a reasonable post hoc task map, but the strong claim that it is derived from a unified physical model—and the Sec. 3.4 suggestion that this motivates unified multi-task frameworks—is not supported by the text.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of event-based imaging technologies. It proposes a physical imaging model based on the plenoptic function L(x,y,z,θ,φ,λ,τ) and sensor equations for frame and event outputs (Eqs. 1–6), then organizes existing methods into two categories: image/video enhancement (Sec. 3) and advanced imaging (Sec. 4), followed by a discussion of challenges and open questions (Sec. 5). The survey covers representative works in event-to-video reconstruction, video frame interpolation, deblurring, rolling shutter correction, video super-resolution, low-light enhancement, HDR imaging, multi-task frameworks, NeRF/3DGS-based multi-view generation, light field reconstruction, and photometric imaging.","tokens_in":16137,"tokens_out":6196,"duration_ms":62360,"significance":"If the organizational claim were fully supported, the survey would provide a useful unifying perspective on a rapidly growing field. The paper has clear strengths: it gives a concise and mostly accurate account of the standard event generation model; it covers a broad range of tasks and recent work; and it identifies concrete open problems such as RAW-domain processing and hybrid sensor co-design. The continuously updated resource link is also a practical benefit. However, the central claim that the taxonomy is derived from the physical model is under-supported, and the survey's completeness is asserted rather than demonstrated. With the taxonomy claim revised or properly substantiated, this could be a valuable resource; in its current form, the contribution is primarily a selective and somewhat imbalanced task map.","major_comments":[{"comment":"The taxonomy is not derived from the physical model as claimed. In Sec. 3.2, VFI, deblurring, RSC, and VSR all recover the same intensity-valued slice of Eq. (1); they differ only in the assumed input degradation, so the model does not separate these tasks. In Sec. 4.3, photometric stereo outputs surface normals and illumination estimates, which are not a subset of L(x,y,z,θ,φ,λ,τ) as defined, and in free space the (x,y,z) and (θ,φ) arguments are redundant along a ray, making 'richer light information' ill-defined. The authors should either define explicit projection operators from Eq. (1) for each task and show that they form a partial order, or revise the abstract and Sec. 1 to say the model motivates rather than determines the taxonomy.","section":"Sec. 2, Eqs. (1) and (6); Secs. 3.2 and 4.3"},{"comment":"The paper asserts comprehensiveness ('a comprehensive study', 'over two thousand research papers published annually') without providing a literature-selection methodology, inclusion criteria, or coverage statistics. With roughly seventy references spanning several large subfields, the reader cannot verify the completeness claim, and the open-questions section in Sec. 5 cannot be distinguished from the authors' chosen emphasis. Please add a methodology paragraph covering databases, time window, and filtering criteria, or explicitly reframe the contribution as a selective survey.","section":"Abstract and Sec. 1"},{"comment":"The discussion of multi-task frameworks and future directions relies heavily on the authors' own papers (e.g., Lu et al. 2023a, 2023b, 2024a, 2024b; Liang et al. 2024) without situating them relative to independent work in the same space. This introduces selection bias into the survey's central recommendation that the field should move toward unified multi-task frameworks. The authors should add a balanced comparison that includes works from other groups and should disclose the relationship of their own papers to the claims being made.","section":"Secs. 3.4 and 5"}],"minor_comments":[{"comment":"The phrase 'extend extend' in the second paragraph is a duplicated word and should be corrected.","section":"Sec. 1"},{"comment":"The abstract contains grammatical issues: 'recently advances' should be 'recent advances', and 'photometric' as a standalone noun should be 'photometric imaging'.","section":"Abstract"},{"comment":"The sentence 'the latter preserving spatial resolution in dynamic captures. without sacrificing spatial resolution.' contains a duplicated and grammatically incomplete phrase; it should be rewritten.","section":"Sec. 4.2"},{"comment":"The entries Lu et al. (2024a) and Lu et al. (2025) appear to cite the same UniINR paper with different years; please verify which entry is correct.","section":"References"},{"comment":"The threshold convention in Eq. (6) is stated only verbally; please specify whether θth is a positive constant and whether the comparisons use absolute values.","section":"Sec. 2, Eq. (6)"},{"comment":"The claim that evaluation on 'simulated or limited datasets raises concerns' would benefit from at least one concrete example or citation to support the concern.","section":"Sec. 3.2, Rolling Shutter Correction"}],"recommendation":"major_revision","confidential_remarks":"The core survey content is useful and the paper can be revised. The main risk is overclaiming a unified physical foundation; the authors should either supply a precise mapping from Eq. (1) to each task or soften the claim. The concentration of self-citations in Secs. 3.4 and 5 should also be addressed for balance. The manuscript is within the scope of the journal if the survey framing is made accurate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serviceable survey with a nice organization, but the paper's central claim—that its taxonomy follows from a physical imaging model—doesn't hold up under scrutiny. The Sec. 2 equations are standard and correctly stated, but they don't do the work the paper asks of them. VFI, deblurring, and rolling-shutter correction all recover the same intensity slice of the plenoptic function; they differ in input degradation, not in recovered light information. Photometric stereo outputs normals and lighting, which are not a slice of L(x,y,z,θ,φ,λ,τ). And 'richer' isn't an unambiguous ordering once you notice that along a free-space ray, (x,y,z) and (θ,φ) are redundant. So the taxonomy is a reasonable post-hoc task map, not a derivation. The stress-test note has this right.\n\nWhat's genuinely good: the paper gathers a broad set of recent works, organizes them into enhancement versus advanced imaging in a way that will be convenient for newcomers, and includes a fairly complete reference list plus a maintained GitHub. The representative method descriptions are generally accurate, and the authors do cite prior surveys rather than pretending to be first. That's real value.\n\nSoft spots, in order of size. First, the taxonomy's foundation is overstated; the authors should either make the mapping from plenoptic sub-spaces to task categories explicit or soften the claim to 'organized around a physical model of the sensor output.' Second, selection bias: the paper leans on the authors' own papers in multi-task sections (several Lu et al. entries) and in open questions, and it would benefit from independent coverage. Third, some promotional language—'paradigm shift,' 'disruptive technologies,' 'transformative approach'—should be toned down. Fourth, there's no stated search/completeness methodology, so 'comprehensive' is doing a lot of work; a short inclusion criteria paragraph would help. None of these are load-bearing mathematical errors, because there is no new derivation to check. The physics in Sec. 2 is fine within its simplified scope; the question is what the survey uses it for.\n\nWho this is for: people entering event-based imaging, or researchers in adjacent areas who want a quick map of enhancement vs. advanced tasks. It will be a useful reference even if the unifying narrative doesn't quite work.\n\nRecommendation: send it to peer review, but flag the taxonomy claim and the self-citation imbalance for revision. A serious referee can fix those without changing the paper's bones.","headline":"A useful but overclaimed survey of event-based imaging; the physical-model framing doesn't actually drive the taxonomy, but the map is worth having.","tokens_in":16635,"tokens_out":2255,"would_cite":true,"duration_ms":22654,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68-02","68T45","68U10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Event-based imaging has a single unifying physical model","keywords":["event cameras","event-based imaging","plenoptic function","image enhancement","video enhancement","high dynamic range","light field reconstruction","photometric imaging"],"falsifier":"Measure the full plenoptic function of a controlled scene with an independent instrument, then run a leading event-based video-interpolation method on footage of the same scene. If the interpolated frames contain angular or spectral variation that the standard-video slice of the plenoptic function should not include, or if the method's error does not track the amount of plenoptic information the model says it recovers, the organizing axis fails.","tokens_in":15691,"feed_emoji":"🎥","tokens_out":8272,"duration_ms":69564,"temperature":0.7,"pith_summary":"This survey sets out to show that event-based imaging is not a patchwork of separate tricks but a field with a single physical foundation. It builds a model in which light is described by a plenoptic function—the intensity of every ray in a scene—and in which conventional cameras record an integral of that function over time while event cameras record differential changes (threshold crossings of log-intensity). On this basis, every imaging task can be ranked by how much of the plenoptic function it reconstructs, from standard enhancement tasks (events to video, interpolation, deblurring, HDR) to advanced ones (light-field reconstruction, multi-view generation, photometric imaging). A sympathetic reader would care because the model supplies a shared language for comparing methods, exposing which tasks are related, and pointing toward unified multi-task systems as the natural next step.","feed_headline":"Event-based imaging has a single unifying physical model","feed_subtitle":"A survey ranks every event-camera task by how much of the scene's light field it recovers, from video interpolation to 3D.","key_machinery":"The central object is the physical imaging model built on the plenoptic function, defined in one phrase as the intensity of every light ray in a scene. The load-bearing pair of equations is the frame output $I_0 = f_i(\\int_{t_0}^{t_0+\\Delta t} A(t)\\,dt)$ (Eq. 5) and the event output $E(i,j,t_0)$ as the sign of threshold-crossing log-intensity change (Eq. 6). The model does the taxonomic work: it expresses the two sensor modalities as integral and differential views of the same analog signal, which is why hybrid sensors that capture both are treated as the natural hardware, and why each imaging task can be located by which slice of the plenoptic function it recovers.","core_discovery":"The central claim is that a physical imaging model organizes the entire event-based imaging literature. In this model, the complete light signal is the plenoptic function $L(x,y,z,\\theta,\\phi,\\lambda,\\tau)$—the intensity of every ray at every position, direction, wavelength, and time (Eq. 1). A sensor at the focal plane maps pixels to ray directions, a response function weights wavelengths, and photoelectric conversion introduces Gaussian and Poisson noise, yielding an analog signal $A$ (Eqs. 2–4). Frame cameras output the integral of $A$ over an exposure time (Eq. 5), while event cameras output differential signals: $+1$ or $-1$ when the log-intensity change crosses a threshold (Eq. 6). Because both outputs come from the same analog signal, the survey argues that every enhancement task is a reconstruction of the integral from the differential, and every advanced task is an attempt to recover a higher-dimensional slice of the plenoptic function. The survey maps the field along this axis and draws the conclusion that unified multi-task frameworks are the natural culmination.","pith_inferences":["An immediate extension would turn the taxonomy into a metric: score any imaging system by how many dimensions of the plenoptic function it recovers (spatial, angular, spectral, temporal). The survey does not propose such a metric, but the model makes it well-defined.","The model implies that real sensor deviations from Eq. 6—refractory periods, per-pixel contrast thresholds, readout noise—should translate into measurable errors in task performance. A systematic study linking sensor non-idealities to task accuracy would test whether the shared-foundation claim holds outside ideal simulations.","The same integral/differential decomposition might organize other neuromorphic or computational imaging modalities, such as multispectral event sensors, where the taxonomy's three sensor classes (event-only, beam-splitter, hybrid) would have direct analogues."],"forward_implications":["Enhancement tasks—events to video, frame interpolation, deblurring, rolling-shutter correction, super-resolution, low-light enhancement, HDR—form one family because they all reconstruct the same integral-from-differential relationship.","Advanced tasks—multi-view generation with NeRF and 3D Gaussian splatting, light-field reconstruction, photometric stereo—inherit the video-enhancement techniques because they recover larger slices of the same plenoptic function.","Hybrid sensors that output both frames and events are the hardware consequence of the model, since they capture the integral and differential halves of the same analog signal in spatial and temporal alignment.","The shared foundation makes unified multi-task frameworks (for example, joint deblurring, interpolation, and rolling-shutter correction) a natural goal rather than an ad hoc combination.","The model exposes open research problems as places where the plenoptic-function view has not yet been applied: RAW-domain processing, hybrid sensor/algorithm co-design, multi-camera setups, IMU fusion, and foundation models with events."],"supporting_citations":[{"why":"Supplies the plenoptic function, the model's complete description of light signals.","marker":"(Chan, 2014)"},{"why":"Establishes the event camera model and the standard survey baseline the paper builds on.","marker":"(Gallego et al., 2020)"},{"why":"Provides sensor characteristics (high dynamic range, low latency) and an early hybrid EVS+APS sensor.","marker":"(Finateu et al., 2020)"},{"why":"Describes a modern RGB hybrid event-vision sensor, grounding the hardware classification.","marker":"(Kodama et al., 2023)"},{"why":"Demonstrates real-world advantages of event cameras in low-latency vision, motivating the imaging model.","marker":"(Gehrig & Scaramuzza, 2024)"},{"why":"Foundational learning-based event-to-video reconstruction (E2VID) that enhancement tasks build on.","marker":"(Rebecq et al., 2019)"},{"why":"Introduces TimeLens, the initiating work for event-based video frame interpolation.","marker":"(Tulyakov et al., 2021)"},{"why":"EventNeRF, a representative advanced task showing event-based multi-view generation.","marker":"(Rudnev et al., 2023)"}],"fun_headline_variants":["Event camera imaging: one physical model to unify","All event-based imaging tasks share one physical model","Survey: one physical model unifies event-based imaging","Event imaging survey: a single physical model for all tasks","Event-based imaging's unifying physical model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every piece of light information a camera could ever want can be written down in one mathematical function (the plenoptic function), and that event cameras are faithfully described as threshold crossings of log-intensity change; if either fails for real scenes or real sensors, the unified taxonomy loses its foundation.","fun_headline_variants_meta":{"raw":{"variants":["Event camera imaging: one physical model to unify","All event-based imaging tasks share one physical model","Survey: one physical model unifies event-based imaging","Event imaging survey: a single physical model for all tasks","Event-based imaging's unifying physical model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000507,"raw_usage":{"total_tokens":2456,"prompt_tokens":911,"completion_tokens":1545,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":1472}},"tokens_in":527,"tokens_out":1545,"duration_ms":10773,"temperature":1.0,"reasoning_tokens":1472,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:08:54.476378+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the full plenoptic function of a controlled scene with an independent instrument, then run a leading event-based video-interpolation method on footage of the same scene. If the interpolated frames contain angular or spectral variation that the standard-video slice of the plenoptic function should not include, or if the method's error does not track the amount of plenoptic information the model says it recovers, the organizing axis fails.","supporting_citations":[{"cited_title":"Plenoptic Function, chapter 8, pp.\\ 618--623","cited_arxiv_id":null,"evidence_quote":"Supplies the plenoptic function, the model's complete description of light signals."},{"cited_title":"u ck, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J Davison, J \\","cited_arxiv_id":null,"evidence_quote":"Establishes the event camera model and the standard survey baseline the paper builds on."},{"cited_title":"1.22 m 35.6 mpixel rgb hybrid event-based vision sensor with 4.88 m-pitch event pixels and up to 10k event frame rate by adaptive control on event sparsity","cited_arxiv_id":null,"evidence_quote":"Describes a modern RGB hybrid event-vision sensor, grounding the hardware classification."},{"cited_title":"Low-latency automotive vision with event cameras","cited_arxiv_id":null,"evidence_quote":"Demonstrates real-world advantages of event cameras in low-latency vision, motivating the imaging model."},{"cited_title":"Time lens: Event-based video frame interpolation","cited_arxiv_id":null,"evidence_quote":"Introduces TimeLens, the initiating work for event-based video frame interpolation."},{"cited_title":"Eventnerf: Neural radiance fields from a single colour event camera","cited_arxiv_id":null,"evidence_quote":"EventNeRF, a representative advanced task showing event-based multi-view generation."}],"review_version":1}