{"id":"a257b253-bcf0-4a3a-9766-3889b0ababc3","arxiv_id":"2412.06191","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Events from scanning mirrors or kaleidoscopes can be used to reconstruct light fields at high speed, enabling post-capture refocusing and depth estimation.","lead":"Event cameras usually capture only changes in brightness over time. This paper adds optical attachments so the same cameras can capture light fields, the complete set of light rays from different angles, at high speed. The result is a new way to make high-speed video that can be refocused after capture and used for instant depth measurement.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reconstructed frames from off-distribution networks carry all downstream claims; no quantitative ground-truth validation exists for the central 'light field capture' claim.","rationale":"The reader's weakest assumption identifies exactly the load-bearing step: downstream results are computed from reconstructed intensity frames, not from direct light-field measurements. My reading of the paper confirms this. The galvanometer design's events are angular derivatives for static content and a mixture of angular and temporal derivatives for dynamic content; E2VID is a pretrained, time-sequential reconstruction network. The kaleidoscope design similarly relies on TimeLens, which is designed for event-guided RGB frame interpolation, not for light-field view recovery. The manuscript provides no quantitative comparison of reconstructed views against any ground truth, despite having a physics-based simulator capable of producing such ground truth. The paper's own Sec. 7 admission that training on event-field data would yield a significant performance boost further weakens the empirical support. This is not an internal inconsistency or a violation of consensus; it is an unvalidated generalization of learned models to a new input distribution. Because the central capture claim is plausible but not quantitatively established, the conditional verdict is appropriate. The proposed concrete test would settle whether the concern actually lands by directly measuring reconstruction fidelity in the exact operating regime claimed.","tokens_in":13971,"tokens_out":6086,"duration_ms":67146,"concrete_test":"Run the provided Blender simulator on a static scene with circular galvo scan; feed the generated events through the exact Sec. 6.4 E2VID pipeline and the supplementary template-matching calibration; compute PSNR/SSIM between the 40 reconstructed views and the ground-truth light field views at the corresponding scan positions. Then repeat with a moving-object scene and compare against ground-truth frames at matching times. If PSNR falls below ~20 dB or systematic viewpoint-dependent errors appear, the claim that event fields capture light fields is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Event Fields capture light fields at speed, resolution, and dynamic range—rests on intensity frames produced by pretrained event-to-video networks (E2VID for the galvanometer design, TimeLens for the kaleidoscope). These frames are then used for refocusing, depth, and HDR. But the galvanometer events are not ordinary temporal derivatives of a fixed viewpoint: for static content Eq. (5) makes them angular derivatives ∂B/∂ω scaled by the scan velocity C'(t), and for dynamic content they are an inseparable sum of angular and temporal derivatives. E2VID is trained on time-sequential event streams (fixed viewpoint, scene motion), so it is being applied off-distribution. The paper's only validation is visual: no experiment compares the reconstructed 40 views against ground-truth light field frames, even though the Blender simulator could produce such ground truth. Sec. 7 concedes that training on event-field data 'would yield a significant performance boost.' If the reconstructed frames are biased or hallucinated, the refocused images, depth maps, and 'fastest' headline are unsupported. This is load-bearing because the light field is never measured directly; it appears only as the output of a network not designed for this input type.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'Event Fields', a framework for capturing light fields with event cameras by multiplexing angular information into either the spatial or temporal domain. Two hardware designs are presented: a kaleidoscope-based spatial multiplexer and a galvanometer-based temporal multiplexer. The authors provide a mathematical model (Sec. 4), a Blender-based simulator (Sec. 5.1), hardware prototypes (Secs. 5.2, 5.3), and demonstrations of HDR light fields, slow-motion refocusing, post-capture refocusing, and depth estimation (Sec. 6). The central claim is that event fields enable light-field capture at high speed, resolution, and dynamic range, with the galvanometer design achieving 250 fps and 40 views at megapixel resolution.","tokens_in":14266,"tokens_out":3785,"duration_ms":38390,"significance":"The concept of event fields is novel and timely, and the mathematical derivation in Sec. 4 is sound for the ideal noiseless event camera. The two complementary multiplexing designs are clearly articulated, and the authors have built a simulator and real prototypes, with code and data promised open-source. The potential applications—high-speed refocusing, HDR light-field capture, and instant depth estimation—are compelling. However, the significance is currently limited by the qualitative nature of the evaluation and the reliance of all downstream results on off-the-shelf event-to-video networks that are not designed for event-field inputs. If the authors add quantitative validation and address the reconstruction-fidelity concern, this could be a strong contribution to computational imaging.","major_comments":[{"comment":"The evaluation in Sec. 6 is almost entirely qualitative. Fig. 5 compares kaleidoscope vs. galvanometer on a simulated scene, but the comparison rests on visual inspection of insets; no PSNR/SSIM or other quantitative metrics are reported for the reconstructed views, refocused images, or depth maps. Since the Blender simulator can generate ground-truth light fields, the authors should provide quantitative numbers for view reconstruction accuracy, refocusing fidelity, and depth error. This is load-bearing because the paper's central claim is that event fields 'capture light fields at high speed, resolution, and dynamic range,' and that claim cannot be assessed without quantitative evidence.","section":"Sec. 6.1, 6.4"},{"comment":"The use of pre-trained E2VID and TimeLens networks is a major correctness risk. Equation (5) shows that, for static content, the galvanometer events encode angular derivatives ∂B/∂ω scaled by the scan velocity, not temporal derivatives; for dynamic content they encode an inseparable sum of angular and temporal derivatives. These input distributions differ substantially from the time-sequential event streams for which E2VID and TimeLens were trained. All downstream results—the reconstructed light field views, refocused images, depth maps, and HDR output—are computed from these network outputs, not from direct measurements. The paper itself concedes in Sec. 7 that 'training models specifically for our event fields would yield a significant performance boost.' Without a ground-truth comparison of the reconstructed views against known light fields, or an analysis of reconstruction error on off-distribution inputs, the central demonstrations could be artifacts of the network rather than genuine captures. Please add a validation experiment using the simulator where ground truth is available, or fine-tune the networks on event-field data and re-evaluate.","section":"Sec. 4.2, Sec. 6.4, Sec. 7"},{"comment":"The 'instant true depth estimation' claim is overstated. The method in Sec. 6.5 relies on a depth-from-focus algorithm to obtain disparity, which is then converted to physical depth via a linear calibration fit to seven LED positions between 15 and 100 inches (Supp. Fig. 3). This is a fitted, scene-dependent calibration, not a first-principles depth measurement. The paper does not report depth accuracy or error bars, and the linear model's validity outside the calibrated range is unknown. Please provide quantitative depth error against known depths, and qualify the term 'true depth' accordingly.","section":"Sec. 6.5, Supp. Fig. 3"},{"comment":"The claim that the reconstructed light field at '250 fps and 40 views' is 'the fastest recorded so far at a megapixel resolution' is unsupported: no comparison baseline, citation, or quantitative measure of 'fastest' is given. Additionally, Fig. 11 shows that the event camera loses events due to readout bandwidth limits at scanning frequencies beyond 250 Hz, so the actual light-field quality is affected by event loss. Please provide a clear definition and evidence for this claim, and discuss how bandwidth limitations affect the reported frame rate and view count.","section":"Sec. 6.4"}],"minor_comments":[{"comment":"The notation in the mapping function 'xs = x mod (r/n)' is unclear. Please define x, r, and n explicitly (e.g., pixel index, sensor width, number of angular views) and specify whether this is a 1D or 2D mapping.","section":"Sec. 4.1, Eq. (3)"},{"comment":"The derivation uses an equality where an approximation is implied: ∂B/∂t = pkC/∆tk is a finite-difference approximation. Consider using '≈' consistently and mention that this holds in the limit of small ∆t.","section":"Sec. 3.2, Eq. (2)"},{"comment":"The text states the kaleidoscope captures a '3 × 3 views' light field, but Fig. 3 does not clearly show the view grid. Please annotate the figure or add a separate illustration of the 3×3 angular views.","section":"Sec. 5.2, Fig. 3"},{"comment":"The calibration in Supp. Fig. 1 uses template matching on E2VID-reconstructed frames to determine the scanning curve. Please state how potential reconstruction artifacts (e.g., blur or hallucination) might bias this calibration, and whether the calibration was verified against independent measurements.","section":"Supp. Fig. 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript presents an interesting idea with sound mathematical foundations and real hardware prototypes, but the evidence currently falls short of the strong claims in the title and abstract. The major revisions should focus on quantitative evaluation against ground truth and on validating or adapting the reconstruction networks. If the authors can provide these, the paper could be a strong candidate for acceptance. I would also like the editor to verify that the open-source promise and double-blind anonymity are handled correctly, since the anonymous repository link is already referenced."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real idea with two working prototypes and a clean mathematical skeleton, but the central claim currently rests on intensity frames produced by pre-trained event-to-video networks applied to event streams they were not designed for. The conditional verdict is about right, and the stress-test concern lands. What is new and good: the galvanometer-based temporal multiplexing reads angular derivatives directly by scanning the view, which is a legitimately different way to get light field information from an event camera. The event field abstraction cleanly unifies spatial and temporal multiplexing, and the math in Sec. 4 is correct for the ideal camera. The kaleidoscope variant is honestly described as functionally same as the MLA-based EventLFM, so the fresh engineering is the galvo design plus the Blender simulator and open data. The paper is also admirably honest in Sec. 7, conceding that E2VID and TimeLens are optimized for time-sequential event streams and that training on event field data would yield a significant performance boost. That concession pre-validates the main weakness. Soft spots: first, the galvanometer demonstrations (refocusing, depth) are built from E2VID frames. For static scene content the events are angular derivatives, not ordinary temporal derivatives, so E2VID is off-distribution. There is no comparison of reconstructed views against ground truth, even though the Blender simulator would make that easy. So the 40 views at 250 fps claim is really about frames from a network that saw something else. Second, the instant true depth numbers come from a fitted depth-disparity calibration, not first principles. Third, the fastest recorded headline lacks a baseline comparison. The HDR demonstration is one qualitative figure without a dynamic range measurement. The kaleidoscope/SloMoRF results are comparatively safer because TimeLens is closer to its training distribution, though the evaluation is still largely qualitative. Bottom line: this deserves a serious referee. The core idea is worth publishing if the reconstructions get validated, for example with simulator ground-truth comparisons or a fine-tuned reconstruction model. As is, I would want revision before acceptance. It is a solid computational-imaging paper, not a field-shattering one.","headline":"A genuinely new galvanometer-based light field sensing idea with a clean theoretical core, but the headline results ride on off-distribution neural reconstructions and lack quantitative validation; worthy of peer review with substantial revision.","tokens_in":782,"tokens_out":1464,"would_cite":true,"duration_ms":28781,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Event cameras, fitted with a kaleidoscope or a galvanometer, can capture light fields — the full set of ray directions — at high speed, high resolution, and high dynamic range, enabling post-capture refocusing and instant depth estimation…","keywords":["event camera","light field","temporal multiplexing","spatial multiplexing","galvanometer","kaleidoscope","high dynamic range","depth estimation"],"falsifier":"Render a scene with a known light field and depth map using the paper's own simulator, synthesize events from the galvanometer's temporal-multiplexing model, run the full reconstruction-plus-refocusing pipeline, and compare the output against ground-truth refocused images and depth; large error specifically along the angular-derivative direction, or depth bias that changes with surface texture, would falsify the claim that the galvanometer design captures a usable structured light field.","tokens_in":13798,"feed_emoji":"⚡","tokens_out":9382,"duration_ms":80306,"temperature":0.7,"pith_summary":"Event cameras report per-pixel brightness changes asynchronously, which makes them fast and high-dynamic-range but blind to angular information — ordinary cameras integrate over all ray directions. The paper proposes 'event fields,' a way to multiplex angular information into either space or time so the same event sensor records a light field. A kaleidoscope design spatially splits the view into multiple angles and captures temporal derivatives per view; a galvanometer design scans the view along a Lissajous curve and directly captures angular derivatives. With the galvanometer, the authors report a structured light field at 250 frames per second with 40 views at megapixel resolution, which they state is the fastest recorded at that resolution. If this holds, light-field features like refocusing and depth estimation become available for fast, high-dynamic-range scenes.","feed_headline":"Fastest megapixel light field: 250 fps via event camera","feed_subtitle":"A scanning mirror unlocks refocusing and depth on high-speed, high-dynamic-range scenes.","key_machinery":"The load-bearing mechanism is the derivative identity produced by optical multiplexing. A galvanometer steering the rays along a periodic curve $C(t)$ makes the event camera's thresholded brightness change proportional to the angular derivative of the light field times the steering velocity; for static scenes this reduces to $\\partial L/\\partial \\omega$ along the scan. The scan curve is a Lissajous curve, and because the steering signal is known, every event carries a known angular view, giving a structured light field with no loss of spatial resolution. The kaleidoscope instead uses a spatial mapping $x_s = x \\bmod (r/n)$ to assign each pixel to one of $n$ views, trading spatial resolution for angular resolution while keeping events as temporal derivatives. Reconstruction from these derivative streams uses pre-trained event-to-video networks, and refocusing integrates the recovered views over the angular dimension.","core_discovery":"The paper establishes that an event camera can act as a light-field sensor if the angular dimension is multiplexed before the sensor. Spatial multiplexing places several angular views side by side, so each pixel responds to the temporal derivative of the brightness of its view. Temporal multiplexing steers the rays with a fast mirror, so the brightness change at a pixel is the angular derivative of the light field along the scan curve times the steering speed, i.e., $\\partial B/\\partial t \\approx (\\partial B/\\partial \\omega)(\\partial C/\\partial t)$. Because light fields are typically smooth, both temporal and angular derivatives are sparse, so the event stream stays within the camera's bandwidth. The galvanometer design yields a structured event field — the view at every event time is known from the calibrated scan — from which the paper reconstructs intensity frames, refocuses them over depth, and estimates depth from focus on a moving person at 100 Hz. The kaleidoscope design, combined with a co-located RGB camera and an event-to-video interpolation network, produces slow-motion color light-field video with refocusing on dynamic scenes.","pith_inferences":["We would push the bandwidth analysis further: since the 250 Hz ceiling comes from event readout saturation, adaptive scan rates that slow for static regions and speed up for moving ones—flagged in the paper as future work—could raise the dynamic light-field rate beyond 250 fps without new hardware.","The reliance on networks trained for time-sequential events suggests the reported reconstructions likely understate the hardware's potential; generating event-field training data, as the paper suggests, should measurably improve refocused image quality, a claim that could be tested by fine-tuning on synthetic event fields.","The kaleidoscope's spatial-resolution loss and blindness to static scenes make the two designs complementary; a hybrid that uses the galvanometer for full-resolution static context and the kaleidoscope for high-speed dynamic views is a natural engineering extension.","Depth-from-focus assumes disparity is proportional to depth across the Lissajous scan; we would test the method's behavior near scan-curve self-crossings and on non-Lambertian surfaces, where this proportionality is expected to break down."],"forward_implications":["Post-capture refocusing becomes possible on fast dynamic scenes—the paper demonstrates refocusing on a fan spinning at 480 RPM—something frame-based light-field cameras cannot do at these speeds.","Instant true depth becomes available at the scan rate (e.g., 100 Hz) via depth-from-focus, without SLAM or multi-camera arrays.","The event camera's wide dynamic range carries over to the light field, so HDR light fields can be captured in a single shot without bracketing.","The galvanometer design preserves full spatial resolution and gives a structured, pre-calibrated field, so per-frame calibration is not needed.","Because angular derivatives are sparse, the bandwidth cost of adding the angular dimension is modest; in the paper's experiments, 250 Hz scanning stays within the event camera's readout limits."],"supporting_citations":[{"why":"Provides the 4D light field parameterization and the refocusing equation used for all refocused images and depth-from-focus.","marker":"[25]"},{"why":"The kaleidoscope add-on design that the spatial multiplexing prototype is adapted from.","marker":"[30]"},{"why":"The event-to-video reconstruction network used to turn galvanometer event streams into intensity frames.","marker":"[44]"},{"why":"The event-based video interpolation network used to create the slow-motion color light field (SloMoRF).","marker":"[48]"},{"why":"Prior microlens-array event camera for dynamic light fields; serves as the main comparison baseline for the galvanometer design.","marker":"[14]"},{"why":"Supports the claim that light fields are smooth and their angular derivatives are sparse, which motivates the bandwidth argument for temporal multiplexing.","marker":"[21]"}],"fun_headline_variants":["Event fields: light fields at event camera speeds","Event fields: high-speed, high-dynamic-range light fields","Multiplexed event cameras capture light fields for refocusing","Refocusing and depth from event-based light fields at 100 Hz"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recovered frames, and therefore the refocusing and depth results, are produced by pre-trained event-to-video networks that were optimized for ordinary time-sequential event streams, not for the angular-derivative streams these optics generate; if those networks do not faithfully invert angular derivatives, the demonstrated light-field capabilities are not actually delivered by the hardware.","fun_headline_variants_meta":{"raw":{"variants":["Event fields: light fields at event camera speeds","Event fields: high-speed, high-dynamic-range light fields","Multiplexed event cameras capture light fields for refocusing","Refocusing and depth from event-based light fields at 100 Hz"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000522,"raw_usage":{"total_tokens":2567,"prompt_tokens":1029,"completion_tokens":1538,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":1467}},"tokens_in":645,"tokens_out":1538,"duration_ms":13057,"temperature":1.0,"reasoning_tokens":1467,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:54:32.766044+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Render a scene with a known light field and depth map using the paper's own simulator, synthesize events from the galvanometer's temporal-multiplexing model, run the full reconstruction-plus-refocusing pipeline, and compare the output against ground-truth refocused images and depth; large error specifically along the angular-derivative direction, or depth bias that changes with surface texture, would falsify the claim that the galvanometer design captures a usable structured light field.","supporting_citations":[{"cited_title":"Light field rendering","cited_arxiv_id":null,"evidence_quote":"Provides the 4D light field parameterization and the refocusing equation used for all refocused images and depth-from-focus."},{"cited_title":"A reconfigurable camera add-on for high dynamic range, multispectral, polarization, and light-field imaging","cited_arxiv_id":null,"evidence_quote":"The kaleidoscope add-on design that the spatial multiplexing prototype is adapted from."},{"cited_title":"Eventlfm: Event camera integrated fourier light field microscopy for ultrafast 3d imaging","cited_arxiv_id":null,"evidence_quote":"Prior microlens-array event camera for dynamic light fields; serves as the main comparison baseline for the galvanometer design."},{"cited_title":"Learn- ing light field angular super-resolution via a geometry-aware network","cited_arxiv_id":null,"evidence_quote":"Supports the claim that light fields are smooth and their angular derivatives are sparse, which motivates the bandwidth argument for temporal multiplexing."}],"review_version":1}