{"id":"712be4f5-7380-4b66-a496-6f74cc79ac91","arxiv_id":"2511.06126","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"By factorizing a dynamic diffraction sequence into shared spatial and compact temporal neural features, the authors recover ptychographic videos at 30 fps with 308-nm linewidth resolution over a centimeter-scale field.","lead":"This paper demonstrates a lensless, video-rate ptychography system: a moving coded sensor and a space-time neural field reconstruct videos with 308-nm resolution over a ~40 mm² field at 30 frames per second. A generalist should care because it points toward label-free, single-sensor microscopy that watches cells, bacteria, crystals, and dissolving drug microneedles in real time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'gigapixel SBP' claim is not supported by any SBP accounting; the stated sensor and resolution imply ~0.4 GP over the active area, so the central quantitative claim is unverified.","rationale":"The reader's verdict is CONDITIONAL and flags missing SBP computation and the low-rank factorization assumption. I agree that the factorization's validity domain is a genuine concern, but I identify the missing SBP accounting as the single most load-bearing issue because it directly determines whether the central quantitative claim ('gigapixel SBP') is even well-defined. The paper's own sensor numbers make a 1-GP SBP nontrivial: ~20 MP raw frames, ~40 mm² active area, and a 308-nm resolution target yield ~0.4 GP of resolved modes over that area. Therefore 'gigapixel' can only be true if the scanning/translation enlarges the FOV or if the output grid is oversampled beyond physical resolution; neither is specified. This is not an accusation of overclaiming—it may be a definitional ambiguity between output samples and resolved modes—but it is the number the whole abstract rests on. I do not see a reason to change the CONDITIONAL verdict: the issue is addressable by supplying the missing computation and reconstruction grid specification, but until that is done the central claim should not be accepted as stated. The concrete test I propose is deliberately simple: run the released code/config, print the FOV and output array shape, and compare the SBP to 10^9. This settles the concern without requiring new experiments.","tokens_in":14540,"tokens_out":14480,"duration_ms":154241,"concrete_test":"Obtain the reconstruction configuration/code for Fig. 2c and Fig. 5 (the paper promises Zenodo links in Supplementary Note 5). Record the output array size (Nx, Ny) and physical FOV (Lx, Ly). Compute SBP = round(Lx/0.308 μm) × round(Ly/0.308 μm). Also estimate the actual recovered bandwidth from the Fourier support of one reconstructed dynamic frame (threshold at, say, 10× the noise floor). If either number is below 10^9, replace 'gigapixel' in the abstract and title with the actual SBP (≈0.4 GP or the measured value). If both are ≥10^9 and the Fourier support confirms sub-308-nm content across the full FOV, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is quantitative: 'gigapixel space-bandwidth product' with '308-nm linewidths across centimeter-scale fields at 30 fps.' The paper never computes SBP or specifies the reconstruction FOV/grid. The sensor is 5120×3840 with 1.4-μm pixels, i.e. ~7.2×5.4 mm (~40 mm²) active area. At the claimed 308-nm resolution, the number of resolved modes over that area is (7.2 mm/0.308 μm)×(5.4 mm/0.308 μm) ≈ 0.4 GP, below 'gigapixel.' If the coded-sensor translation enlarges the effective FOV, the text must state the reconstructed FOV and output array size; it does not. The statement that 'SBP grows with the efficiency of correlation extraction rather than the number of independent measurements' is a slogan, not a theorem: a prior can produce high-resolution output arrays, but resolved information is still bounded by measurement diversity plus prior validity. This missing SBP accounting is the most load-bearing gap because the headline is precisely the gigapixel number; the low-rank factorization concern is the mechanism by which that number could be achieved, but even a perfect factorization cannot make 0.4 GP of resolved information into 1 GP if the FOV is not enlarged.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a video-rate ptychographic imaging system based on a translating coded sensor and a space-time neural field representation. The complex optical field at each time is factorized as a Hadamard product of multi-resolution hash-encoded spatial features and compact temporal feature vectors, decoded by dual MLPs that output real and imaginary components, and optimized with a gradient-domain loss against a ptychographic forward model. Experiments demonstrate USAF resolution targets claimed to resolve 308-nm linewidths, snowflake melting, Na2CO3 crystallization, stem-cell wound healing and division, E. coli growth in microfluidics, microneedle dissolution, and EUV time-varying probe reconstruction. The central claim is that single-sensor, video-rate gigapixel SBP imaging with centimeter-scale field coverage is achieved at 30 fps.","tokens_in":14884,"tokens_out":3688,"duration_ms":33939,"significance":"If fully supported, the work would be a significant advance: lensless ptychography with mesoscale field of view, sub-micron resolution, and video-rate temporal sampling, with label-free quantitative phase, post-measurement refocusing, and dynamic EUV probe recovery. Strengths include the availability of open-source code and datasets, external validation against USAF linewidths and a known 700-µm microneedle design height, and the use of a literature-based dry-mass conversion. However, the headline quantitative claim of a gigapixel space-bandwidth product currently lacks a direct SBP accounting, and the central low-rank factorization assumption is not characterized. These issues are load-bearing for the paper's main claim.","major_comments":[{"comment":"The headline claim of 'gigapixel SBP' and 'centimeter-scale fields' is never supported by an explicit SBP calculation. The sensor is 5120×3840 with 1.4-µm pixels, giving an active area of about 7.2×5.4 mm. At the claimed 308-nm resolution, the number of resolved modes over that static area is roughly 0.4 gigapixels, below the advertised value. If the coded-sensor translation enlarges the effective FOV, the text must state the reconstructed FOV and output grid size; it does not. The Methods list 'reconstruction grid dimensions and magnification' among unspecified parameters. Without this accounting, the central quantitative claim is unverified.","section":"Abstract; Methods, 'Video-rate gigapixel ptychographic sensing system'"},{"comment":"The reconstruction rests on the assumption that the entire dynamic scene is well approximated by a Hadamard product of shared spatial hash-encoded features and a compact set of T temporal feature vectors (stated to be in R^32, with T not specified). No analysis is provided of the class of dynamics for which this low-rank factorization holds, how T should scale with scene complexity, or what happens for scenes with many independently moving components. This is load-bearing because the conditioning that turns underdetermined single-frame diffraction data into a tractable problem depends directly on the validity of this factorization.","section":"Methods, 'Space-time neural field representations'; Fig. 1a"},{"comment":"The quantitative dry-mass growth curve in Fig. 4h and the microneedle dissolution kinetics in Fig. 5d are presented as quantitative results without error bars, replicate counts, or uncertainty propagation. The text claims 'picogram sensitivity,' but no measurement-noise or sensitivity analysis is given. These omissions weaken the biomedical quantitative claims, which are central to the paper's stated versatility beyond the resolution demonstration.","section":"Fig. 4h; Methods, 'Dry-mass tracking of bacterial growth'"},{"comment":"The statement that 'SBP grows with the efficiency of correlation extraction rather than the number of independent measurements' is presented as the main scaling-law conclusion, but no theorem, bound, or information-theoretic argument supports it. As written it is a slogan rather than a result. Resolved information is bounded by measurement diversity plus validity of the prior/factorization; the paper should either formalize this claim or temper it to match what is actually demonstrated.","section":"Discussion"}],"minor_comments":[{"comment":"Typo: 'EVU' should be 'EUV' in the sentence about extreme ultraviolet ptychography.","section":"Results, first paragraph"},{"comment":"The hash table capacity is given as '2^26 entries' but the text says 'up to 226 entries' (the superscript is lost). Please fix the formatting.","section":"Methods, 'Space-time neural field representations'"},{"comment":"'initial learning rate of 10-3' should read '10^-3'.","section":"Methods, 'Reconstruction via gradient-domain loss'"},{"comment":"The dry-mass formula is referenced as 'In Fig. 4b' but the quantification appears in Fig. 4h; please correct the cross-reference.","section":"Results, paragraph on bacterial growth"},{"comment":"The Methods state 'approximately 100 microneedles per 1×1 cm²' while the Results describe 'hundreds of individual needles.' These numbers should be reconciled.","section":"Fig. 5 and Methods, 'Transdermal microneedle patch'"},{"comment":"The 'gigapixel-scale' claim for the crystal dataset should be accompanied by the actual reconstruction array dimensions and pixel count; the main text never specifies the output grid for any dataset.","section":"Fig. 2 and Supplementary Figs. S3-S4"}],"recommendation":"major_revision","confidential_remarks":"The 'gigapixel SBP' headline will attract significant attention, so the missing FOV/grid accounting must be resolved before publication. The low-rank temporal factorization is the mechanism that makes the method work, and its limitations should be stated more transparently. The paper is otherwise an interesting and potentially high-impact contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a serious hardware-plus-algorithm paper that likely delivers video-rate lensless ptychographic phase imaging over large fields, but the central \"gigapixel\" claim is not backed by the text as written. The sensor is 5120x3840 with 1.4-micron pixels, and at 308-nm resolution over the stated ~40 mm² field, you get roughly 0.4 GP of resolved modes, not 1 GP. If the translating sensor enlarges the effective field of view, the paper never says so—no reconstructed FOV, no output grid size, no SBP computation. That is a load-bearing gap because the title and abstract hang on that number.\n\nWhat is genuinely new: the combination of a translating coded sensor and a space-time neural field reconstruction is a sensible and well-executed idea. The dual real/imaginary MLP parameterization and the gradient-domain loss are pragmatic choices that appear to solve real convergence problems in dynamic phase retrieval. The experimental breadth is impressive—USAF resolution, snowflake melting, stem-cell dynamics, bacterial growth, microneedle dissolution, and EUV probe characterization. They check against external references (USAF linewidths, the 700-micron microneedle design height, literature dry-mass conversion), and the coded-surface calibration uses a static reference rather than the dynamic sample, so the circularity burden is low.\n\nSoft spots, in proportion:\n\n1. The gigapixel accounting is the biggest issue. It is not enough to say \"gigapixel-scale\" when the sensor area alone gives 0.4 GP at the demonstrated resolution. If the synthetic aperture from translation expands the FOV, the paper needs to state the reconstructed field size explicitly. Reviewers should push hard on this.\n\n2. The EUV comparison is unfair as presented: OPR gets 1500 iterations, their method gets 40. They may well have done more in the supplement, but the main text makes the baseline look artificially weak.\n\n3. The low-rank space-time factorization is uncharacterized. The demonstrations are all smooth, spatially correlated dynamics. Many independently moving parts or abrupt non-repeating changes could break the 32-dimensional temporal embedding. The paper should at least state this limitation.\n\n4. The quantitative phase results—dry mass, needle height—have no error bars or replicate counts. Minor, given the breadth of demonstrations, but worth noting.\n\n5. The claim that \"SBP grows with correlation extraction rather than the number of independent measurements\" is a slogan, not a theorem. Information still comes from measurements plus prior validity. The framing overreaches.\n\nWho this is for: computational imaging researchers, especially those in ptychography or neural fields. It deserves a serious referee, but the referee report should ask for the SBP calculation and a fair EUV baseline. If the gigapixel claim cannot be substantiated, the authors should reframe the contribution around video-rate dynamic ptychography, which is already a strong result on its own.","headline":"A real advance in dynamic lensless ptychography, but the headline 'gigapixel' claim is not supported by the numbers as written.","tokens_in":15483,"tokens_out":5496,"would_cite":true,"duration_ms":53680,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper demonstrates that a single coded image sensor, translated beneath a dynamic sample and reconstructed through a space-time neural field factorization, can produce 30-fps phase videos with 308-nm resolution over centimeter-scale fie","keywords":["ptychography","gigapixel imaging","space-time neural fields","phase retrieval","lensless imaging","video-rate imaging","computational imaging","low-rank factorization"],"falsifier":"Use the same hardware and reconstruction approach on a specimen of many independently moving particles with no shared dynamics, such as dense random Brownian motion, and compare recovered frames against ground truth. If the low-rank factorization cannot track independent motions or produces artifacts while a control scene with correlated dynamics reconstructs cleanly, the central scaling claim is falsified. A more quantitative variant: sweep the number of temporal feature vectors on a fixed dataset and show that scenes with N independent moving parts require roughly N temporal vectors for arti","tokens_in":14405,"feed_emoji":"🎥","tokens_out":5230,"duration_ms":48286,"temperature":0.7,"pith_summary":"Ptychography normally requires many overlapping measurements, so capturing a moving sample in real time seems hopeless. This paper argues that the redundancy is not a burden but a resource: if all frames are reconstructed jointly as a low-rank factorization of a space-time field, each new coded-sensor measurement can improve both spatial resolution and temporal tracking. On a compact lensless setup with a 405-nm laser and a coded sensor moving at 30 fps, it claims to resolve 308-nm linewidths across centimeter-scale fields, and demonstrates the approach on melting snowflakes, crystallization, stem cells, bacterial growth, dissolving microneedles, and time-varying EUV probe beams. The reason a reader should care is that, if right, it turns ptychography from a slow quasi-static technique into a general lensless video-rate imaging tool, and it recasts the space-bandwidth bottleneck as a correlation-extraction problem rather than a hardware-scaling problem.","feed_headline":"One coded sensor streams 30-fps gigapixel phase video","feed_subtitle":"Neural space-time factorization turns ptychography's slow scans into real-time lensless imaging with 308-nm resolution.","key_machinery":"The load-bearing object is the space-time neural field factorized as a Hadamard product of hash-encoded spatial features and interpolated temporal feature vectors — a low-rank representation of the complex optical field. Multi-resolution hash encoding compresses the gigapixel spatial coordinates into a shared lookup table (claimed up to roughly 1000-fold compression), and a compact set of temporal vectors (32-dimensional in the implementation) encodes dynamics without explicit motion models. Dual MLPs decode the real and imaginary parts separately to avoid phase-wrap discontinuities, and a gradient-domain loss on spatial derivatives of amplitude supplies the optimization signal. This factori","core_discovery":"At the center of the work is a claim about how to lift the space-bandwidth barrier in ptychography. Instead of reconstructing each frame independently from its own diffraction pattern, the authors factorize the entire space-time volume into multi-resolution hash-encoded spatial features and a small number of learnable temporal feature vectors; every voxel's complex field is formed by the Hadamard product of those features, decoded by two MLPs into real and imaginary parts. A gradient-domain loss on measured versus predicted intensities couples all frames, and the recovered field can be digitally propagated to a chosen focus. The paper reports that this joint optimization converges where fram","pith_inferences":["If the scaling claim is right, then recovered resolution or space-bandwidth product should improve as temporal correlation is exploited; a direct test is to acquire the same dynamic scene with increasing numbers of frames and see whether SBP grows with correlation rather than with measurement count.","The fixed 32-dimensional temporal basis may not scale to scenes with many independent motions; one can test this by imaging several independently moving objects and checking whether artifact-free reconstruction requires a larger temporal dimension or fails entirely.","Because the paper treats probe variation as a temporal function, a natural extension is to model other slowly varying systematic errors, such as sample drift, illumination drift, or stage wobble, as additional temporal features in the same framework.","The success at very low overlap in the EUV demonstration suggests that ptychographic overlap requirements may be governed more by the strength of the temporal prior than by geometric redundancy; this could be tested on simulated data with controlled dynamics and known ground truth."],"forward_implications":["Ptychography can monitor non-repeatable dynamics such as melting, crystallization, cell division, bacterial growth, and drug-device dissolution at video rate, instead of being limited to quasi-static samples.","A single sensor with a moving coded surface suffices for gigapixel-scale video, avoiding multi-camera arrays or long sequential acquisitions.","Post-measurement digital refocusing becomes a natural byproduct, eliminating real-time autofocus requirements for live-cell and incubator imaging.","In EUV, X-ray, and electron ptychography, time-varying illumination or radiation-induced sample evolution can be modeled as temporal features, potentially lowering overlap requirements and radiation dose."],"fun_headline_variants":["Neural fields make gigapixel ptychography video-rate","Single-sensor gigapixel video via space-time neural fields","Ptychography hits video rate with neural space-time factorization","Gigapixel phase video from one sensor using neural correlations","Neural factorization streams gigapixel ptychography at 30 fps"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument rests on the assumption that the full space-time field of any captured scene is well approximated by a low-rank product of shared spatial features and a compact set of temporal feature vectors; if a scene contains many independently moving or rapidly decorrelating regions, this factorization cannot represent it and the reconstruction advantage disappears.","fun_headline_variants_meta":{"raw":{"variants":["Neural fields make gigapixel ptychography video-rate","Single-sensor gigapixel video via space-time neural fields","Ptychography hits video rate with neural space-time factorization","Gigapixel phase video from one sensor using neural correlations","Neural factorization streams gigapixel ptychography at 30 fps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1130,"prompt_tokens":704,"completion_tokens":426,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":339}},"tokens_in":448,"tokens_out":426,"duration_ms":3940,"temperature":1.0,"reasoning_tokens":339,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T23:20:53.458549+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the same hardware and reconstruction approach on a specimen of many independently moving particles with no shared dynamics, such as dense random Brownian motion, and compare recovered frames against ground truth. If the low-rank factorization cannot track independent motions or produces artifacts while a control scene with correlated dynamics reconstructs cleanly, the central scaling claim is falsified. A more quantitative variant: sweep the number of temporal feature vectors on a fixed dataset and show that scenes with N independent moving parts require roughly N temporal vectors for arti","supporting_citations":[],"review_version":1}