{"id":"7adb750e-10fb-477b-9674-2e031e58c41e","arxiv_id":"2411.18597","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured light system using acousto-optic line scanning and an event camera demonstrates full-frame depth scanning at 1000 fps, four times faster than prior event-based structured light.","lead":"Researchers built a 3D scanner that projects a million light planes per second through water-based acousto-optic lenses and pairs them with an event camera, reaching 1000 full-frame depth scans per second. This could make fast-moving objects, like rotating fans or robot arms, easier to capture in 3D than existing structured-light systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full-frame 1 kfps claim is contradicted by the paper's own dense-scene data: at 1 kfps the cat F1 score drops to 0.253 with missing columns, and the authors must accumulate events over multiple scans to recover dense depth, so single-pass full-frame 1000 fps is demonstrated only for sparse scenes.","rationale":"I read the paper in good faith and credit the hardware contribution: the photodetector traces in Fig. 5(d–f) independently confirm line scanning at 1 kfps, 10 kfps, and 100 kfps, and the arbitrary line-placement demonstrations give direct support to the acousto-optic scanning mechanism and the cost reduction relative to Pediredla et al. Those parts are not the weak link. The weak link is the transition from 'the AO device can scan light planes at MHz rates' to 'the system enables full-frame 3D scanning at 1000 fps.' The paper's own Section 5.1 shows that, for dense static scenes, the event camera drops events at 1 kfps, producing missing columns and F1 = 0.253, and the authors must accumulate multiple frames to recover a complete depth map. Accumulation over several scans means the effective temporal resolution is lower than 1 ms, so the reconstructed depth is not a single-pass full-frame 1 kfps frame. The dynamic scenes are demonstrated at 500 fps or lower, and the 1000 fps fan result in Fig. 1 is a sparse scene. This is exactly the assumption the reader flagged: one usable event per sensor row per light-plane sweep. The paper's own data violate it for dense scenes. I also considered whether the 1 µs event-camera timestamp resolution is a separate load-bearing issue, since at 1 kfps the light-plane period is about 1.39 µs; this could make plane-index assignment fragile, but the reader's concern is already sufficient to justify a conditional verdict and is directly supported by the paper's table. Equation (6) also appears dimensionally ambiguous as typeset, but it is a minor presentational issue compared with the event-loss evidence. Therefore I agree with the reader's verdict: CONDITIONAL, requiring clarification of single-pass versus accumulated reconstruction and a qualified statement of when 1 kfps full-frame depth is actually achieved.","tokens_in":17423,"tokens_out":8206,"duration_ms":106099,"concrete_test":"Reprocess the raw 1 kfps cat and dolphin event streams using only events from a single sweep of 720 light planes (no accumulation), and compute depth coverage, chamfer distance, and F1. If single-sweep F1 remains near 0.253 with large missing-column gaps, while the accumulated result is 0.641, then the 1000 fps full-frame claim is unsupported for dense scenes and the paper should be revised to claim sparse/ROI 1 kfps. Additionally, count the average number of events generated per light-plane sweep at 1 kfps; if this exceeds roughly 1.4e3 events per plane (1 GEvents/s divided by 720 planes per frame and 1000 frames/s), the observed dropouts are a necessary consequence of the sensor bandwidth ceiling rather than an incidental artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim, 'full-frame 3D scanning at speeds of 1000 fps', is not established by the reported experiments for dense scenes. At 1 kfps on the cat figurine, Table 2 reports F1 = 0.253, chamfer distance 4.56 mm, and missing columns are visible; the authors attribute this to the event camera randomly losing events when the number of events exceeds its readout bandwidth. The '1000 (acc.)' row, which improves F1 to 0.641, explicitly accumulates events over multiple scans to 'emulate a higher-bandwidth camera,' so it is not a single-pass 1 ms depth frame. Dynamic full-frame results are shown at 500 fps for the fan and at 180/200 fps for non-periodic scenes, not at 1 kfps; the 1 kfps fan in Fig. 1 is a sparse scene. Thus the load-bearing assumption behind the headline—that each sensor row produces one usable event per light-plane sweep, yielding one complete depth frame per sweep—is violated by the paper's own Table 2 and Fig. 6 in the dense case. The adaptive '10x beyond the theoretical limit' claim is likewise limited to sparse ROIs (two knobs in Fig. 10), not full-frame scanning. The hardware validation via photodetector traces is credible, but it validates line scanning speed, not full-frame depth reconstruction at 1 kfps.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a structured light system that combines a custom acousto-optic line scanner with an event camera. The acousto-optic device uses an ultrasonic transducer in water to create traveling refractive-index lenses that focus a pulsed laser into moving light planes. The authors validate line scanning at 1, 10, and 100 kfps using a photodetector, demonstrate arbitrary line placement, and reconstruct static and dynamic scenes. The central claim is full-frame 3D scanning at 1000 fps, four times faster than prior event-based structured light systems, plus adaptive ROI scanning at 10 kfps. The paper reports quantitative metrics (chamfer distance, F1 score) for static scenes at several rates and qualitative dynamic results for rotating and hand-held objects.","tokens_in":17722,"tokens_out":5161,"duration_ms":44607,"significance":"The hardware contribution is significant: a low-cost acousto-optic line scanner with MHz-rate steering capability, which is validated by independent photodetector oscilloscope traces at 1, 10, and 100 kfps. If the full-frame speed claims held with appropriate qualifiers, this would be the fastest event-based structured light scanner demonstrated to date. The paper is transparent about the quality degradation at high rates and attributes it to event-camera bandwidth limits. However, the headline claims in the abstract and Table 1 are broader than the experimental evidence: dense-scene full-frame reconstruction at 1 kfps is not achieved without event accumulation, dynamic full-frame results are shown at 500 fps or lower, and the 10x adaptive speedup is limited to region-of-interest scanning. These issues do not negate the value of the hardware, but they require a revision of the claims.","major_comments":[{"comment":"The abstract, Table 1, and Conclusion state that the system enables 'full-frame 3D scanning at speeds of 1000 fps', but the dense-scene data in Table 2 and Fig. 6 do not support this without qualification. At 1 kfps on the cat scene, the F1 score is 0.253 and chamfer distance is 4.56 mm, with missing columns visible in the depth map. The '1000 (acc.)' row, which improves F1 to 0.641, explicitly accumulates events over multiple scans to 'emulate a higher-bandwidth camera'. The text in Section 5.1 confirms that at 1 kfps the number of missing columns becomes significant and that only sparse scenes (such as the fan in Fig. 1) avoid this degradation. Thus single-pass full-frame depth at 1 kfps is demonstrated only for sparse scenes, not as a general full-frame capability. Please qualify the headline claim or provide dense-scene single-pass 1 kfps results.","section":"5.1, Table 2, Fig. 6"},{"comment":"Table 1 lists dynamic performance as 1 kfps, but the dynamic full-frame experiments show 500 fps for the fan (Fig. 8) and 180/200 fps for non-periodic scenes (Fig. 9). The 1 kfps example in Fig. 1 is the sparse fan scene, not a dense dynamic scene. Consequently, the claim of 'full-frame 3D scanning at 1000 fps' for dynamic scenes is not substantiated by the reported experiments. The authors should either provide a dense dynamic scene scanned at 1 kfps or revise the abstract and Table 1 to reflect the demonstrated dynamic rates.","section":"5.2, Figs. 7-9, Table 1"},{"comment":"The abstract claims 'achieving effective scanning speeds an order of magnitude beyond the camera's theoretical limit', but this result is demonstrated only for adaptive scanning of two regions of interest (the two orange knobs in Fig. 10). This is not full-frame scanning; the 10x speedup applies only when illuminating sparsely selected ROIs. The abstract and Section 5.3 should explicitly state that this speed gain is for ROI-only adaptive scanning, not full-frame. Additionally, the demonstrated 10 kfps ROI rate is limited by laser power rather than by the AO device itself, as the text notes; this practical constraint should be stated alongside the claim.","section":"5.3, Fig. 10, Abstract"}],"minor_comments":[{"comment":"The photodetector validation demonstrates line scanning at 1, 10, and 100 kfps, but the abstract and Section 3 claim a capability of 'two million light planes per second'. The paper does not directly validate the 2 MHz rate; it appears to be a theoretical maximum derived from the beat-frequency model of Eq. (5). Please clarify whether this rate is a theoretical capability or a measured one, and state the highest validated scanning rate in the abstract.","section":"4, Fig. 5"},{"comment":"The dynamic full-frame results use retroreflective tape on the targets (the servo blades in Fig. 7 and the characters in Fig. 9). This is a significant experimental condition that is not mentioned in the abstract or the experimental setup overview. Please state this limitation clearly, as it affects the generality of the dynamic scanning claims.","section":"5.2"},{"comment":"The claim that 'our system produces depth results nearly identical to those from the Galvo-based system at low scanning rates' is not quantified. A direct quantitative comparison between the AO system and the Galvo baseline at 10 fps (e.g., chamfer distance or F1) would strengthen this assertion.","section":"5.1"},{"comment":"The Discussion lists several prototype limitations (small aperture, low laser power, small field of view of 38 mm, large form factor) that are not mentioned in the abstract. The abstract should include a sentence on the current prototype's field-of-view restriction and the need for retroreflective targets in dynamic scenes, to avoid overgeneralizing the demonstrated results.","section":"6"}],"recommendation":"major_revision","confidential_remarks":"The paper's hardware contribution is genuine, and the photodetector-based validation of line scanning is a strong piece of evidence. The central issue is the overstatement of the full-frame 1 kfps claim: the authors' own dense-scene data in Table 2 and Fig. 6 show that single-pass full-frame depth at 1 kfps is only achievable for sparse scenes, and dynamic results cap at 500 fps. The adaptive-speed claim is also ROI-limited. These are load-bearing problems because the abstract and Table 1 present the 1 kfps full-frame capability as the main result. The claims can be fixed within the manuscript's scope by adding explicit qualifiers and revising the abstract; the hardware and method novelty remain publishable. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, read this for the hardware, not for the headline. The acousto-optic line scanner is a genuine advance: pulsing the laser at a frequency offset from the ultrasound, instead of phase-modulating it as in Pediredla et al., is a different and cheaper mechanism, and the photodetector oscilloscope traces at 1, 10, and 100 kfps independently confirm the scanner really moves lines at those rates. Arbitrary line placement is also demonstrated with clean waveforms. The cost reduction (a $250 narrowband amplifier instead of a $23,000 broadband one) is real, and the system-level combination with an event camera is new for structured light. The adaptive ROI scanning at 10 kfps, though only two knobs, is a nice demonstration of the projector's redistributive capability.\n\nThe soft spot is exactly what the stress test says. The abstract's 'full-frame 3D scanning at speeds of 1000 fps' is not backed for dense scenes. The paper's own Table 2 shows the cat at 1 kfps has F1 0.253 and precision 0.146, with missing columns; the improvement to 0.641 comes only from accumulating events over multiple scans, which is explicitly emulating a higher-bandwidth camera. The dynamic scenes are shown at 500 fps for the fan and 180–200 fps for non-periodic motion. The 1 kfps fan in Fig. 1 is a sparse scene and the authors themselves say sparse scenes don't suffer the readout bottleneck. So the honest statement is: the system can scan light planes at 1 kfps and faster, but full-frame dense depth at 1 kfps is not delivered. That's still a useful result, but the headline is over-drawn.\n\nOther, smaller issues: Eq. (6) as printed looks dimensionally off (x_l + l lambda_us over l c_us?) and should be checked. No code or data is released for the event processing/triangulation pipeline, which makes the reconstruction results hard to reproduce. The 10-pixel linewidth and 1 mW laser are acknowledged limitations, and they are the reasons the dense-scene results degrade.\n\nWho is this for? People building fast structured light or scanning hardware will get real ideas from the AO design and the event-camera synchronization details. The paper deserves a serious referee: the hardware contribution is novel and validated, and a revision that recalibrates the claims, fixes the equation, and releases data would be a solid contribution. I would not cite this as a 1 kfps full-frame system without checking whether a revision narrows the claim.","headline":"Real hardware win in AO line scanning, but the 'full-frame 1000 fps' headline only holds for sparse scenes and the paper's own dense-scene data say so.","tokens_in":18286,"tokens_out":3158,"would_cite":true,"duration_ms":29213,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a structured light 3D scanner whose custom acousto-optic projector sweeps two million light planes per second, enabling full-frame depth capture at 1000 fps with an event camera, four times faster than prior…","keywords":["structured light","event camera","acousto-optic scanning","3D scanning","adaptive depth scanning","swept-plane triangulation","ultrasonic GRIN lens"],"falsifier":"Count events per row during a single 1 kfps light-plane sweep over a dense high-contrast object: if many rows produce zero or multiple events, one sweep cannot give a complete depth frame, and the full-frame rate claim would require accumulation to hold.","tokens_in":17217,"feed_emoji":"🔦","tokens_out":10847,"duration_ms":93571,"temperature":0.7,"pith_summary":"The paper introduces a structured light 3D scanner that projects up to two million light planes per second using a custom acousto-optic device, and pairs it with an event camera so that each sweep of a light plane yields a depth map. It reports full-frame depth capture at 1000 frames per second, about four times faster than the previous fastest event-based structured light system. The paper argues that the bottleneck for speed shifts from illumination steering to the event camera's readout bandwidth, and that adaptive scanning of regions of interest reaches effective rates of 10,000 fps, about ten times the camera's theoretical full-frame limit. If the claims hold, this would be the fastest full-frame structured light scanner demonstrated, with a scanning head that costs roughly an order of magnitude less than the prior acousto-optic approach.","feed_headline":"Two million light planes per second push 3D scanning to 1000 fps","feed_subtitle":"Custom acousto-optic projector outruns event-camera readout; adaptive scanning reaches 10 kfps for regions of interest.","key_machinery":"The carrying mechanism is a traveling cylindrical gradient-index (GRIN) lens sculpted by ultrasound. Sinusoidal ultrasound in water creates a refractive index profile $n(x,t)=n_0+n_{us}\\cos(2\\pi x/\\lambda_{us}-2\\pi f_{us}t)$, whose convex lobes act as cylindrical lenses that focus a collimated laser beam into a line moving at the speed of sound. The paper controls the line by pulsing the laser at a frequency ratio $\\alpha$ to the ultrasound, so the focused line steps by $(1-\\alpha)\\lambda_{us}$ per pulse and the sweep rate equals the beat frequency. Programmed pulse timing places lines at arbitrary positions, making the device a redistributive line projector that puts all laser power on selected lines and enables adaptive scanning.","core_discovery":"The central claim is that light-plane steering can be made so fast that the imaging sensor, not the projector, becomes the limiting component in structured light. The paper's acousto-optic projector sculpts traveling cylindrical gradient-index lenses in water with a 2 MHz ultrasonic transducer; each lens focuses a pulsed laser into a line, and the line moves because the laser pulse frequency is offset from the ultrasound frequency, creating a beat that determines the sweep rate. At the event camera, each light plane ideally triggers one event per sensor row, so one sweep is enough to triangulate a full-frame depth map by intersecting backprojected rays with the known plane. The paper demonstrates this at 1000 fps, and shows that adaptively placing lines only where depth changes gives effective 10 kHz scanning, limited in the prototype by laser power rather than by the acousto-optic line rate.","pith_inferences":["The same arbitrary line-placement control could project non-serial patterns such as adaptive grids or coded sequences, trading some event sparsity for robustness or resolution; the paper demonstrates serial sweeps and two-line adaptive scanning only.","The degradation on the dense cat scene implies a natural benchmark for this class of systems: report events per sweep and missing-column counts alongside depth metrics, since dropped columns are a sensor-bandwidth effect rather than a triangulation error.","A clean test of the claimed bottleneck shift is to couple the same projector to an event sensor with higher readout bandwidth; if full-frame depth rate rises past 1 kfps, the acousto-optic line rate is not the practical limit."],"forward_implications":["Full-frame 1000 fps depth is achievable in a single sweep for sparse scenes; for dense scenes, accumulating several sweeps restores complete depth at a lower effective rate.","Adaptive scanning reaches an effective rate of 10 kfps for regions of interest, with the practical limit set by laser power rather than by the acousto-optic steering rate.","The projector can place light lines at arbitrary positions without mechanical motion, enabling redistributive illumination where all laser power goes to the scanned lines.","Further speed gains depend on sensor readout bandwidth rather than illumination steering, with single-photon avalanche diodes identified as a candidate for closing the remaining gap."],"supporting_citations":[{"why":"Supplies the acousto-optic virtual-waveguide light-steering method at MHz rates that this system adapts from point scanning to line scanning.","marker":"[1]"},{"why":"Introduces the adaptive region-of-interest scanning strategy that the AO projector implements at high speed.","marker":"[2]"},{"why":"Establishes the event-based structured light pipeline and the full-frame bandwidth argument that sets the camera-side speed limit.","marker":"[35]"},{"why":"Provides the prior state-of-the-art event-based scanning rate of 250 fps that this work claims to exceed by about four times.","marker":"[36]"},{"why":"Shows the basic combination of an event camera with a laser line scanner for structured light, the foundation for swept-plane depth.","marker":"[65]"},{"why":"Supplies the event-camera parameter settings (bias-hpf, bias-refractory) that let the sensor respond at high sweep rates.","marker":"[74]"},{"why":"Provides the concept of a redistributive projector that places all light on selected lines, used to justify the adaptive advantage.","marker":"[9]"},{"why":"Supplies the geometric calibration and triangulation procedure used to turn line-plane events into depth.","marker":"[72]"}],"fun_headline_variants":["2M light planes/s push 3D scanning to 1000 fps","Acousto-optic projector: 2M planes/s, 3D at 1k fps","2M/s light planes, adaptive scanning: 10k fps effective","Event camera + 2M planes/s: 3D at 1000 fps, 10k adaptive","Speed record: 2M light planes/s, 3D scanning at 1k fps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The full-frame 1000 fps depth claim assumes each projected light plane triggers no more than one usable event per image row, so a single sweep yields a complete frame; the paper's own dense-scene measurement at 1 kfps shows dropped columns, making the claim clean only for sparse scenes.","fun_headline_variants_meta":{"raw":{"variants":["2M light planes/s push 3D scanning to 1000 fps","Acousto-optic projector: 2M planes/s, 3D at 1k fps","2M/s light planes, adaptive scanning: 10k fps effective","Event camera + 2M planes/s: 3D at 1000 fps, 10k adaptive","Speed record: 2M light planes/s, 3D scanning at 1k fps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001335,"raw_usage":{"total_tokens":5389,"prompt_tokens":866,"completion_tokens":4523,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":4406}},"tokens_in":482,"tokens_out":4523,"duration_ms":29882,"temperature":1.0,"reasoning_tokens":4406,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:01:47.829387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count events per row during a single 1 kfps light-plane sweep over a dense high-contrast object: if many rows produce zero or multiple events, one sweep cannot give a complete depth frame, and the full-frame rate claim would require accumulation to hold.","supporting_citations":[{"cited_title":"Megahertz Light Steering Without Mov- ing Parts,","cited_arxiv_id":null,"evidence_quote":"Supplies the acousto-optic virtual-waveguide light-steering method at MHz rates that this system adapts from point scanning to line scanning."},{"cited_title":"Event guided depth sensing,","cited_arxiv_id":null,"evidence_quote":"Introduces the adaptive region-of-interest scanning strategy that the AO projector implements at high speed."},{"cited_title":"ESL: Event-based structured light,","cited_arxiv_id":null,"evidence_quote":"Establishes the event-based structured light pipeline and the full-frame bandwidth argument that sets the camera-side speed limit."},{"cited_title":"Event-based Motion-Robust Accurate Shape Estimation for Mixed Reflectance Scenes","cited_arxiv_id":"2311.09652","evidence_quote":"Provides the prior state-of-the-art event-based scanning rate of 250 fps that this work claims to exceed by about four times."},{"cited_title":"Adaptive pulsed laser line extraction for terrain reconstruction using a dynamic vision sensor,","cited_arxiv_id":null,"evidence_quote":"Shows the basic combination of an event camera with a laser line scanner for structured light, the foundation for swept-plane depth."},{"cited_title":"Event-based vision: A survey,","cited_arxiv_id":null,"evidence_quote":"Supplies the event-camera parameter settings (bias-hpf, bias-refractory) that let the sensor respond at high sweep rates."},{"cited_title":"Homogeneous codes for energy-efficient illumination and imaging,","cited_arxiv_id":null,"evidence_quote":"Provides the concept of a redistributive projector that places all light on selected lines, used to justify the adaptive advantage."},{"cited_title":"Build your own 3d scanner: 3d photography for beginners,","cited_arxiv_id":null,"evidence_quote":"Supplies the geometric calibration and triangulation procedure used to turn line-plane events into depth."}],"review_version":1}