{"id":"fedbe437-e75c-4c44-9525-a212d98187db","arxiv_id":"2506.22426","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Optically shuffling the scene before GRR-shutter capture yields per-pixel random exposures, letting a simple TV-prior inverse problem recover single-shot HDR up to 73 dB from an 8-bit sensor.","lead":"Researchers combined a sensor's global reset release (GRR) shutter mode with a random fiber bundle that shuffles the image, so each scene point lands on a pixel with a different exposure time. The result is a single-shot HDR camera system that can recover up to 73 dB of dynamic range from an 8-bit sensor (48 dB native), shown in simulation and in a physical prototype.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The GRR shutter exposure model S[u,v]=T0+t_r(u-1) is assumed from specifications and never experimentally validated; if the real GRR row timing deviates from linearity, the 73 dB hardware claim is not trustworthy.","rationale":"The reader identified the fidelity of P and the ideal S model as the weakest assumption, but treated both as mitigated by calibration. However, the calibration procedure in Sec. 5.1 applies only to P; S is taken directly from the sensor timing equation without experimental verification. A flat-field GRR capture would settle whether Eq. (5) is accurate. If S is wrong, the 73 dB result is not reproducible and the central quantitative claim would require revision or a calibrated S. This is a specific, testable gap rather than a general concern about hardware nonidealities. The paper is otherwise strong: the forward modeling is correct, the simulation studies are thorough, and the prototype demonstration is meaningful. The requested test is straightforward to perform with equipment already used in the paper.","tokens_in":24925,"tokens_out":10569,"duration_ms":125774,"concrete_test":"Capture a flat-field image in GRR mode with the exact T0 and t_r settings used in Sec. 5.3 (T0=189 us, t_r=51 us), illuminating the fiber bundle with a spatially uniform source (e.g., the OLED TV displaying a constant gray level). Average each sensor row and plot the mean value versus row index; fit to S[u,v]=T0+t_r(u-1). If the fitted line deviates by more than 1-2% from the ideal ramp, or if the max/min exposure ratio differs from the expected 820:1 by more than a few percent, re-run the ND-filter reconstruction with the measured S. If the recovered dynamic range changes by more than 3 dB, the 73 dB claim depends on an incorrect shutter model and should be revised or the paper should report a calibrated S.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (73 dB dynamic range from a single 8-bit measurement, Sec. 5.3) rests on the forward model A = SP in Eq. (4), where S is the GRR exposure matrix from Eq. (5), S[u,v]=T0+t_r(u-1). The calibration procedure in Sec. 5.1 measures only P (in rolling-shutter mode with constant exposure). S is never measured or validated: no flat-field GRR capture is shown that confirms the linear row-wise exposure ramp, the assumed T0=189 us and t_r=51 us values, or the absence of row-timing nonlinearities (e.g., reset delays, dead time, clock quantization). If the true GRR exposure map deviates from Eq. (5), the data-fidelity term ||E A x - E b||^2 in Eq. (8) is evaluated with a wrong S, systematically biasing the recovered radiance. The 73 dB figure is extracted from a single reconstructed intensity profile of a smooth ND-filter ramp; a bias in S would directly contaminate the max/min ratio. Sec. 6.4 admits model mismatch (stray light, crosstalk, broken fibers) but does not separate S errors from P errors. Because P is calibrated and S is assumed, the unvalidated shutter model is the most load-bearing gap in the hardware claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a single-shot HDR imaging architecture that combines the global reset release (GRR) shutter mode of an off-the-shelf CMOS sensor with optical randomization provided by a random fiber bundle. The forward model is b = Q{SPx + eta}, where S is a diagonal row-wise exposure ramp and P is a (calibrated, approximately permuting) optical transport matrix. The authors show analytically that for an ideal permutation P, the de-shuffled measurement is Q{Lambda x} with Lambda a diagonal matrix of randomly reordered exposures, so the system implements spatially random exposure before quantization. HDR recovery is posed as an erasure-regularized inverse problem with total variation. Extensive simulations on the 180-image SI-HDR dataset compare the method against HDR-CNN, Deep Optics HDR, and a 2x2 ND filter array, showing competitive or better performance, especially at high saturation rates. A hardware prototype using a random fiber bundle and an Allied Vision 1800 U-1240 camera is calibrated by measuring the system PSF, and a quantitative ND-filter experiment reports a recovered dynamic range of approximately 73 dB from a single 8-bit measurement. Limitations, including model mismatch, motion blur, and depth-dependent defocus, are discussed in Section 6.","tokens_in":25214,"tokens_out":4476,"duration_ms":52468,"significance":"If the hardware claims hold, this is a practically valuable contribution: it extends the dynamic range of an inexpensive, unmodified sensor using off-the-shelf optics and a lightweight TV prior, avoiding the cost and complexity of custom sensors. The core analytical observation that GRR exposure combined with a permutation matrix yields random exposures is correct and is a clean, parameter-free design insight. The simulation study is unusually thorough, covering six quality metrics, ablations, saturation sweeps, and a 180-scene dataset. The hardware validation is also a strength: the authors calibrate a full system PSF, validate against external ground truth (multi-shot HDR and a 14-bit camera), test two orientations of the ND filter, and provide a control simulation showing that fiber transmittance variation alone cannot explain the result. The main weakness is that the GRR shutter model S in Eq. (5) is assumed from specifications and never directly validated, which directly affects the trustworthiness of the headline 73 dB hardware number.","major_comments":[{"comment":"The GRR exposure model S[u,v] = T0 + t_r (u-1) is used to form the forward model A = SP in Eq. (4), but it is never experimentally validated. No flat-field capture in GRR mode is shown to confirm the assumed linear row-wise exposure ramp, the values T0 = 189 us and t_r = 51 us, or the absence of row-timing nonlinearities (e.g., reset delays, clock quantization, dead time). Because the data-fidelity term in Eq. (8) uses this S, a systematic error in the exposure map would bias the recovered radiance and directly contaminate the max/min ratio used to compute the 73 dB figure. I recommend adding a direct calibration of the effective exposure map (e.g., flat-field GRR captures at several controlled light levels, or a comparison of the measured flat-field gradient against Eq. (5)) and/or a sensitivity analysis showing that the reconstructed dynamic range is robust to plausible deviations from linear row timing.","section":"Section 5.3, Eq. (5)"},{"comment":"The headline claim of 73 dB dynamic range rests on a single measured intensity profile from a single ND-filter sample. The maximum-to-minimum ratio of approximately 5000 is extracted from one line across the reconstruction, and although Appendix C.2 repeats the measurement with two orientations, this is still a single scene and a single profile. The paper should report the uncertainty in this estimate (e.g., variation across multiple profile positions, repeat captures, or the noise floor of the reconstruction) or present additional quantitative samples. As written, the reader cannot distinguish a robust system-level property from a favorable single realization.","section":"Section 5.3, Fig. 9(c)"},{"comment":"The paper acknowledges model mismatch from clumping, crosstalk, broken fibers, stray light, and nonuniform transmittance, but it never quantifies how far the calibrated P deviates from a permutation or how sensitive the recovered HDR is to calibration errors. The 73 dB hardware result is obtained by inverting this P via Eq. (8), so the accuracy of P is load-bearing. I recommend reporting calibration residual statistics (e.g., the fraction of energy in the expected fiber location, the condition number or coherence of P, or the correlation between measured and predicted PSFs) and, if feasible, a simulation in which the calibrated P is perturbed by the observed mismatch to bound the effect on reconstructed dynamic range.","section":"Section 6.4 and Section 5.1"}],"minor_comments":[{"comment":"The sentence \"we reconstruct images from two different LED lamp orientations to eliminate ensure that we did not achieve high dynamic range\" contains a typo ('eliminate ensure' should be 'eliminate the possibility' or 'ensure').","section":"Section 5.3"},{"comment":"The sentence \"Event sensors with learning-based methods can reconstruct high-speed HDR video. [Rebecq et al. 2019; Zou et al. 2021]\" is missing connecting text and has a period before the citation; it should be merged into the preceding discussion.","section":"Section 2"},{"comment":"The paper uses both \"up to 73 dB\" (abstract, introduction) and \"around 73 dB\" (Section 5.3). Please standardize the phrasing so the claim is unambiguous.","section":"Section 5.3"},{"comment":"The control simulation in Fig. 19 is described only briefly; it would help to state explicitly what metric or visual criterion shows that the fiber transmittance variation 'fails completely' for the test scene.","section":"Appendix C.4"},{"comment":"The phrase \"approximately <5%\" is awkward; consider writing 'approximately 5% or fewer' or 'less than 5%'.","section":"Section 6.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well suited to a computational imaging venue. The theoretical contribution and simulation study are strong, and the hardware prototype is a genuine step toward practical single-shot HDR. The main risk is that the headline 73 dB claim depends on an unvalidated shutter model; I would like to see that addressed before publication, but I do not see a fundamental flaw in the approach. The manuscript is likely to be acceptable after the validation/sensitivity analysis is added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine single-shot HDR architecture that works on a real bench, and the math behind it is simple enough to check in an afternoon. The main thing I'd want before believing the headline 73 dB number is an independent check of the GRR shutter's row-exposure ramp; the paper assumes it from specs.\n\nWhat's new: pairing the global reset release shutter mode with a pseudorandom spatial permutation. The observation in Eq. (7) — that with a permutation matrix P, P^T S P is just the exposure matrix with entries randomly reordered — is clean and correct. That turns the GRR exposure gradient into a spatially random exposure map, which makes the inpainting problem much better posed than raw GRR or global shutter with TV. The simulations are thorough: 180 SI-HDR scenes, several metrics, ablations showing the combination beats either component alone, and fair comparison to HDR-CNN, Deep Optics HDR, and ND filter arrays. The hardware demonstration is real: a random fiber bundle, a calibrated PSF matrix, and an ND-filter ramp whose reconstructed profile tracks a multi-shot ground truth over about 5000:1. They also show the fiber-to-fiber transmittance variation can't explain the result. That's solid evidence the idea works.\n\nSoft spots: the stress-test note is right that S is never independently measured. The paper states T0=189 us and t_r=51 us and assumes the linear row-wise model Eq. (5). A flat-field GRR capture would have settled whether the ramp is really linear, including reset delays or clock quantization. If the true exposure map is nonlinear, the forward model used in inversion is wrong, and the 73 dB figure is contaminated. The reconstructed profile tracking the ND filter suggests the true S is close to the model, but \"close\" isn't a measured validation. Second, the quantitative hardware result is a single intensity profile from one scene, no error bars, and the qualitative results look decent but clearly not clean. Those are documented limitations, not hidden flaws. The calibration procedure is heavy (0.7 s per column, 940 MB matrix), so this is a proof-of-concept, not a product. No code or data released either.\n\nBottom line: this deserves a serious referee. The core idea is correct, the simulation evidence is strong, and the prototype is a genuine existence proof. I'd want the authors to add a flat-field GRR validation and repeated measurements before I'd fully trust the 73 dB, but that's a revision, not a reason to reject. This is the kind of paper I'd cite and bring to reading group.","headline":"A clever, well-supported single-shot HDR architecture with real hardware evidence; the main unvalidated assumption is the GRR shutter timing model, which should be checked before trusting the 73 dB claim.","tokens_in":25769,"tokens_out":2250,"would_cite":true,"duration_ms":23717,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An off-the-shelf sensor and a random fiber bundle can capture a 73 dB HDR image in a single 8-bit frame.","keywords":["high dynamic range imaging","single-shot HDR","global reset release shutter","random fiber bundle","optical randomization","computational imaging","total variation regularization","spatially varying exposure"],"falsifier":"Repeat the ND-filter quantitative validation with the prototype, then re-measure the system matrix $P$ from scratch and compare the two reconstructions of the same static scene; if the second reconstruction loses more than a few dB of dynamic range or shows visible model-mismatch artifacts, the claim that the calibrated $P$ faithfully supports the 73 dB result is falsified.","tokens_in":24727,"feed_emoji":"📸","tokens_out":5649,"duration_ms":58679,"temperature":0.7,"pith_summary":"This paper proposes a single-shot HDR imaging architecture built entirely from off-the-shelf parts: a conventional sensor run in global reset release (GRR) shutter mode, plus an optical element that randomly permutes the image before it lands on the sensor. Because GRR mode gives each row a slightly different exposure time, the permutation turns the smooth row-wise exposure gradient into a spatially random exposure map across the scene. The authors show that this randomization makes the HDR recovery problem well posed under a simple total variation prior, and they validate the idea in simulation and with a physical prototype using a random fiber bundle. Their central quantitative claim is that the prototype reconstructs a dynamic range of up to 73 dB from a single 8-bit measurement whose sensor is specified at 48 dB. If correct, this means a camera built from commodity components could capture scene contrast of roughly 5000:1 in one frame, with no multi-exposure bracketing and no per-scene tuning.","feed_headline":"One 8-bit shot captures 73 dB of dynamic range","feed_subtitle":"Random optical shuffling plus a stock sensor's global reset release makes single-shot HDR work.","key_machinery":"The load-bearing object is the permutation matrix $P$ and the diagonalization identity $\\Lambda = P^T S P$. Here $S$ is the diagonal exposure matrix of the GRR shutter, whose entry for row $u$ is $S[u,v] = T_0 + t_r(u-1)$ — the initial exposure plus the line-by-line readout time — and $P$ is an idealized optical randomizer that maps each scene point to a pseudorandom sensor pixel. The identity shows that after de-shuffling, the measurement equals a clipped, randomly re-exposed version of the scene, so the GRR gradient is unfolded from a smooth ramp into pseudorandom pixel-wise exposures. The recovery then treats saturated pixels as erasures via a diagonal mask $E$ and solves a TV-regularized least-squares problem, so the machinery's work is to turn an ill-posed inpainting problem over large clipped regions into a well-posed inverse problem over scattered missing pixels.","core_discovery":"The paper's central discovery is that pairing a random permutation of the scene with the GRR shutter's linear exposure gradient is mathematically equivalent to applying a spatially random exposure to the scene before quantization. Concretely, if $P$ is a permutation matrix, the de-shuffled measurement takes the form $\\hat{\\mathbf{b}} = Q\\{\\Lambda \\mathbf{x}\\}$ with $\\Lambda = P^T S P$ a diagonal matrix whose entries are the GRR exposure times randomly reordered; the proof is that $P^T$ commutes with the pointwise nonlinearity $Q$ and permutes the diagonal of $S$. The consequence is that every local patch of the scene contains a mixture of long and short effective exposures, so saturated pixels are scattered rather than contiguous and can be inpainted from well-exposed neighbours using only total variation. The paper demonstrates this both in simulation, where it outperforms prior single-shot methods at 1% and 10% saturation, and in hardware, where a random fiber bundle coupled to a commercial sensor recovers an ND-filter test scene spanning a measured attenuation ratio of about 5000:1, or 73 dB.","pith_inferences":["The same principle could be ported to video or endoscopy: any sensor with a row-wise exposure gradient, combined with a fixed random mapping, yields per-pixel exposure diversity in a single frame, at the cost of motion blur sensitivity.","The calibration bottleneck (about 0.7 seconds per column of $P$) suggests that moving from a fiber bundle to a bonded, near-permutation mapping, or to learned blind reconstruction, is the natural path to practicality.","If the ideal-permutation analysis carries over to real bundles, the method's saturation-robustness argument predicts that performance should improve as the fiber mapping approaches a true permutation; a direct test would compare reconstruction quality against measured clumping or crosstalk statistics.","The patch-wise dynamic range analysis implies a testable prediction: scenes with isolated highlights (for example, sunlight through leaves) should reconstruct worse than scenes with clustered highlights, consistent with the correlation the paper reports."],"forward_implications":["Single-shot HDR can be achieved with stock sensors and ordinary optics, eliminating the need for custom dual-exposure pixels or costly sensor redesigns.","Large, contiguous highlight regions such as bright sky or lamps are recoverable because the random exposure map scatters saturation into small clusters that a local prior can inpaint.","The operating dynamic range is tunable through two sensor parameters, the base exposure $T_0$ and the row readout time $t_r$, so the same hardware can be reconfigured for different scene brightness ranges.","A real prototype with an off-the-shelf random fiber bundle achieves about 73 dB from an 8-bit measurement, approximately 25 dB beyond the sensor's native range.","Because the forward model is linear and the prior is simple TV, the method needs no per-scene training or exposure optimization."],"supporting_citations":[{"why":"Establishes spatially varying pixel exposures as the mechanism for extending dynamic range, the foundation this paper re-implements with random exposures.","marker":"Nayar and Mitsunaga 2000"},{"why":"Provides the multi-shot HDR fusion method used to generate ground truth and defines the baseline the single-shot approach aims to replace.","marker":"Debevec and Malik 1997"},{"why":"Introduces the total variation prior that the paper uses to inpaint saturated and underexposed pixels.","marker":"Rudin et al. 1992"},{"why":"Supplies FISTA, the optimization algorithm used to solve the TV-regularized inverse problem.","marker":"Beck and Teboulle 2009"},{"why":"Provides the parallel proximal approximation to TV denoising used inside the reconstruction.","marker":"Kamilov 2016"},{"why":"HDR-CNN is a leading learned single-shot baseline whose hallucination behavior the paper compares against.","marker":"Eilertsen et al. 2017"},{"why":"Deep Optics HDR is a PSF-engineering baseline that the paper's method is shown to beat at high saturation rates.","marker":"Metzler et al. 2020"},{"why":"The SI-HDR dataset provides the 180 test scenes used for the simulation comparisons.","marker":"Hanji et al. 2022"},{"why":"HDR-VDP-3 is the perceptual metric used to evaluate reconstruction quality together with PSNR and SSIM.","marker":"Mantiuk et al. 2023"}],"fun_headline_variants":["One shot, 73 dB: shuffle pixels to beat sensor limits","Random pixel shuffle gives single-shot HDR from 8-bit sensor","Stock sensor plus random optics reaches 73 dB in one exposure","Single exposure to 73 dB via randomized GRR shutter capture"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole scheme stands on the premise that the optical system is accurately captured by a known, near-permutation matrix $P$ and that the GRR shutter timing follows the ideal row-wise model $S[u,v] = T_0 + t_r(u-1)$; if the real fiber bundle's mapping drifts after calibration or the shutter timing deviates, the unshuffled measurement and the claimed 73 dB result would not hold.","fun_headline_variants_meta":{"raw":{"variants":["One shot, 73 dB: shuffle pixels to beat sensor limits","Random pixel shuffle gives single-shot HDR from 8-bit sensor","Stock sensor plus random optics reaches 73 dB in one exposure","Single exposure to 73 dB via randomized GRR shutter capture"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3194,"prompt_tokens":1030,"completion_tokens":2164,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":2091}},"tokens_in":646,"tokens_out":2164,"duration_ms":16098,"temperature":1.0,"reasoning_tokens":2091,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:04:29.477127+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the ND-filter quantitative validation with the prototype, then re-measure the system matrix $P$ from scratch and compare the two reconstructions of the same static scene; if the second reconstruction loses more than a few dB of dynamic range or shows visible model-mismatch artifacts, the claim that the calibrated $P$ faithfully supports the 73 dB result is falsified.","supporting_citations":[],"review_version":1}