REVIEW 4 major objections 6 minor 75 references
Vision without Images: End-to-End Computer Vision from Single Compressive Measurements
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Computer vision can run directly on a single compressive measurement from a sensor with only 8x8 masks, without reconstructing video, and this end-to-end approach outperforms standard full-frame pipelines in ultra-low light.
desk verdict A hardware-motivated 8x8-mask SCI vision system with a genuinely useful multi-task design, but the paper's central 'no reconstruction' claim rests on an undefined 'signal estimate' input and is not actually tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is CompDAE (Compressive Denoising Autoencoder), an asymmetric encoder-decoder built on STFormer, a spatiotemporal Transformer architecture for video snapshot compressive imaging. A 3D-CNN token generator converts the measurement into tokens; stacked STFormer blocks learn space-time dependencies; and a lightweight decoder with $N<M/2$ blocks produces either the denoised spacetime signal during pre-training or a task output after fine-tuning. The mask is constructed as $M_t = J_{n_x/m_x}\otimes A_t$, an $8\times8$ binary sub-mask $A_t$ repeated across the frame by a Kronecker product, which is what makes the pattern physically integrable. The forward model is Poisson-Gaussian, $Y_{\mathrm{noisy}}=\sum_{t=1}^{T}(X_t\odot M_t+\eta_{p,t})+\eta_g$, and the encoder also receives a noise-level map so one model adapts to different APCs. Rate-constrained training adds an Exp-Golomb-motivated term $R(\theta)=\frac{1}{N}\sum_{i=1}^{N}(|\theta_i|+\epsilon)^\nu$ to the task loss, and a half-BackSlash schedule keeps accuracy while improving compressibility.
What would settle it
A decisive test is to run the same trained CompDAE on a real photon-counting sensor with a tiled $8\times8$ mask under 0.1 lux and compare quantitative edge and depth metrics against the simulated values; if real hardware noise pushes edge ODS well below the simulated 0.690 and below a simple denoised full-frame baseline, the central claim fails. A complementary probe is to train the identical model with $4\times4$ masks: a steep drop in edge ODS would show that the method depends on the particular mask size rather than on learned reconstruction-free features.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a spatiotemporal Transformer-based denoising autoencoder can map the raw noisy measurement $Y=\sum_{t=1}^{T} X_t \odot M_t + N$ directly to pixel-level task outputs, and that an $8\times8$ binary mask replicated across the frame is sufficient for that mapping. Quantitatively, CompDAE reaches an edge-detection ODS F1 score of 0.690 at APC=20 with 23.82M parameters, versus 0.474 for DexiNed, and a depth-estimation absolute relative error (AbsRel) of 0.181 on KITTI, lower than Unidepth, Metric3D, MiDaS, and LightedDepth under the same low-light simulation. The same shared encoder, combined with lightweight task-specific decoders, extends to semantic segmentation, and a noise-aware training input lets a single model operate across an APC range from 1 to 60. The central claim is that reconstruction is not a necessary intermediate step: the semantic content of the scene can be read directly from the compressive measurement.
Load-bearing premise
The load-bearing premise is that an $8\times8$ binary mask, repeated across the frame, leaves enough spatial structure in the single summed measurement for pixel-level vision; the paper gives no theoretical bound or resolution analysis for this, and its ablations only vary mask density from 0.3 to 0.5.
Editorial extensions
If this is right
- An SCI camera with a fixed $8\times8$ mask could act as a low-power, low-bandwidth front end that outputs edges, depth, or segmentations directly from the sensor, skipping ISP and reconstruction.
- The shared encoder can be frozen and paired with different small decoders for different tasks, so a single sensing platform supports multiple vision outputs without retraining the expensive part.
- Because the system never produces a full-resolution RGB frame, it naturally cuts communication cost and provides privacy at the sensor level for sensitive scenes.
- The rate-constrained training result implies the model can be aggressively compressed after training, making it deployable on edge devices with limited memory.
- A production system must fix the compression ratio $C_r$ between training and deployment; the ablations show mismatched $C_r$ values degrade accuracy significantly.
Reading between the lines
- The periodicity of the $8\times8$ tiling suggests the same mask pattern could scale to larger sensor formats without redesigning the mask, but the paper only tests $128\times128$ and $256\times256$ simulations, so this transfer remains untested.
- The narrow mask-density sweep (0.3 to 0.5) leaves open whether even simpler masks, such as density 0.2 or deterministic binary patterns, would still support pixel-level tasks; a broader sweep would map the true hardware constraint.
- If the depth and edge results hold under real photon noise, the approach could enable always-on automotive or robotics perception in near darkness, but the simulated Poisson-Gaussian noise may be less temporally correlated than the noise on real photon-counting sensors.
- A further consequence the paper hints at: because the sensor never forms an RGB image, a CompDAE camera could output only edges or depths, leaving no recoverable image at the source and making privacy protection a physical property rather than a software add-on.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CompDAE, an STFormer-based compressive denoising autoencoder that performs edge detection, depth estimation, and semantic segmentation directly from single noisy compressive measurements captured with tiled 8x8 binary masks. The method combines a shared encoder with lightweight task-specific decoders, a unified noise-adaptive training strategy conditioned on a Poisson noise proxy (APC), and a BackSlash-inspired rate-constrained training term. Experiments on simulated low-light SCI measurements from DAVIS, KITTI, and Cityscapes, plus a qualitative real-hardware SPAD result, report large improvements over conventional APS-based and two-stage reconstruction-then-task baselines.
Significance. If the central claim is substantiated, the work would be a meaningful step toward hardware-friendly SCI with small masks and measurement-domain dense vision at low light, potentially reducing bandwidth and latency for edge, depth, and segmentation tasks. The paper is explicitly empirical and does not overclaim theoretical guarantees, which is appropriate. The combination of a multi-task shared encoder, noise-adaptive conditioning, and rate-constrained training is a useful contribution. The main significance is conditional: the central 'without image reconstruction' claim currently rests on an underspecified input representation, and the quantitative comparisons need to be tightened before the reported state-of-the-art margins can be fully credited.
major comments (4)
- [CompDAE encoder and decoder / Unified Noise-Adaptive Training Strategy] The model input is described as 'a multi-channel input tensor comprising a signal estimate, a Gaussian noise map, and a Poisson noise proxy via APC,' but the 'signal estimate' is never defined. The paper also says the estimation module pre-processes the measurement and masks 'as in (Cheng et al. 2021, 2020),' and in those cited works the estimation module computes an initial reconstruction by back-projecting the measurement through the masks. If CompDAE consumes such a back-projected estimate, then it is not operating directly on the raw single measurement, and the abstract's claim of 'directly from noisy compressive raw pixel measurements without image reconstruction' is not supported. This is load-bearing: please define the signal estimate precisely, state whether it is the raw measurement Y, a back-projection, or a full reconstruction, and add an ablation varying the input representation (raw measurement only, back-projection, full reconstruction) to show which representation is necessary for the reported results.
- [Comparison with conventional systems / Tables 1, 2, 5] It is not specified whether the baseline methods (Canny, RCF, BDCN, DexiNed, Unidepth, MiDaS, Metric3D, LightedDepth, SegFormer) receive the same noisy compressive measurement, a clean frame, or a DIC reconstruction. The caption of Table 1 says 'Evaluation on gray-scale video benchmark datasets with our fine-tuned CompDAE' but does not state the input to the baselines. Since the claimed margins (e.g., ODS 0.690 vs. 0.474 for DexiNed in Table 1; AbsRel 0.181 vs. 0.329 for Metric3D in Table 2) could be inflated if baselines are evaluated on different inputs, please specify the exact input and preprocessing for each baseline and report results for all methods on the same measurement. Additionally, none of the tables report error bars or standard deviations; please add them over multiple runs or test splits.
- [Hardware-friendly mask design / Ablation study observations (Table 3c)] The paper provides no resolution or information-theoretic analysis of how an 8x8 mask tiled across the frame preserves enough spatial detail for dense pixel-level tasks, and the only mask-related ablation varies density rho over the narrow range 0.3-0.5. The central hardware claim is that 8x8 masks are sufficient for practical implementation, but all experiments use 128x128 or 256x256 inputs. Please add an analysis or experiments varying mask tile size (e.g., 4x4, 8x8, 16x16, full-frame) and input resolution, and discuss the effect of tiling periodicity on spatial resolution. The real-hardware SPAD result in Figure 6 is qualitative only; quantitative results on photon-limited data would materially support the practical claim.
- [Comparison with Other SCI-based Methods / Table 5] The two-stage baseline DIC is not described with enough detail. It is not stated which DIC variant is used, whether it uses the full adaptive progressive coding system from Liu et al. 2024 or only the reconstruction network, whether the DIC baseline receives the same noisy measurements as CompDAE, or how the downstream task models are configured after reconstruction. Also, 'DIC+DDN' appears in Table 5a but DDN is not defined in the text. Please specify these details so the comparison is genuinely apples-to-apples.
minor comments (6)
- [Equations (2)-(3)] The noise model is not fully consistent: Eq. (3) writes (X_t ⊙ M_t + η_p,t), while the text states α(X_t ⊙ M_t + η_p,t) ∼ P(α(X_t ⊙ M_t)), which implies η_p,t is not additive in the usual Poisson thinning model. Please clarify whether the Poisson noise is applied before or after mask modulation.
- [Implementation Details] The APC values, Gaussian noise sigma, and other noise parameters are described only in prose. Please list the exact APC values used for training and testing and the sigma used in each experiment so the experiments are reproducible.
- [Figure 1] The Kronecker construction M_t = J_{nx/mx} ⊗ A_t is not visually evident from the figure; please add a clearer illustration or annotate the dimensions.
- [Table 4] In the unified model row for 'w/ BackSlash' at APC=20, the OIS entry appears as '.674' instead of '0.674'; this is a typo.
- [Discussions] The Discussion states that 'CompDAE effectively reconstructs multi-frame sequences directly from compressive measurements,' which conflicts with the earlier claim of operating 'without image reconstruction.' Please reconcile this wording.
- [References] The reference for Wu, Zhangt, and Mou 2021 contains a typo: 'Zhangt' should be 'Zhang.'
Circularity Check
No derivational circularity; the central results are empirical and self-contained, with minor verification gaps around the undefined 'signal estimate' input and DexiNed-generated edge labels.
full rationale
The paper's derivation chain is empirical rather than deductive: Eq. (3) defines the Poisson-Gaussian forward model, CompDAE is a learned encoder-decoder trained on simulated noisy measurements, and Tables 1-5 report task metrics against external baselines. No stated equation, fitted constant, or trained parameter is reused as a predicted output by construction. The only self-citations are to BackSlash (Wu, Wen, and Han 2025) and to the STFormer estimation module; these are not load-bearing in a circular sense because the rate-regularizer is written out in the paper as J(θ) = L_task(θ) + λ R(θ) with R(θ) = (1/N)Σ(|θ_i|+ε)^ν, and the ablation in Table 4 shows BackSlash can slightly degrade edge detection, so the main result does not rest on an unverified self-cited theorem. Two limitations are worth flagging but do not amount to circular reduction. First, in the Unified Noise-Adaptive Training Strategy, 'a signal estimate' is never defined; if it is a back-projected frame estimate from the STFormer estimation module, then the abstract's 'without image reconstruction' claim is under-tested, but this is a missing definition and missing raw-measurement ablation, not a derivation that reduces to its own input. Second, edge ground-truth maps are said to be 'generated using DexiNed' while DexiNed is also a comparison baseline; this creates a possible teacher/student evaluation artifact, but the paper states the model is tested on six gray-scale benchmark datasets, so the evaluation is not by construction identical to the training target. Neither issue permits exhibiting a specific equation or fitted parameter that is renamed as a prediction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (6)
- Gaussian noise sigma =
0.01
- Mask density rho =
0.3-0.5 (ablated)
- Compression ratio Cr =
8
- Decoder depth N =
2 for edge, 1 for depth
- BackSlash weight lambda and shape nu =
not specified
- APC training range =
1-60
assumptions (4)
- domain assumption CACTI forward model Y = sum_t X_t * M_t + N with Poisson-Gaussian noise
- standard math i.i.d. binary masks are feasible for signal recovery (Zhao and Jalali 2023)
- domain assumption STFormer blocks learned on DAVIS transfer to low-light SCI measurements
- domain assumption Ground-truth edge maps from DexiNed are valid training targets
Cite this review
Pith. "Pith review of Vision without Images: End-to-End Computer Vision from Single Compressive Measurements." pith.science (2026). https://pith.science/paper/AXNH2Y4J
@misc{pith2026250115122,
author = {Pith},
title = {Pith review of: Vision without Images: End-to-End Computer Vision from Single Compressive Measurements},
year = {2026},
howpublished = {\url{https://pith.science/paper/AXNH2Y4J}},
note = {Machine review of arXiv:2501.15122}
}
abstract
Snapshot Compressed Imaging (SCI) offers high-speed, low-bandwidth, and energy-efficient image acquisition, but remains challenged by low-light and low signal-to-noise ratio (SNR) conditions. Moreover, practical hardware constraints in high-resolution sensors limit the use of large frame-sized masks, necessitating smaller, hardware-friendly designs. In this work, we present a novel SCI-based computer vision framework using pseudo-random binary masks of only 8$\times$8 in size for physically feasible implementations. At its core is CompDAE, a Compressive Denoising Autoencoder built on the STFormer architecture, designed to perform downstream tasks--such as edge detection and depth estimation--directly from noisy compressive raw pixel measurements without image reconstruction. CompDAE incorporates a rate-constrained training strategy inspired by BackSlash to promote compact, compressible models. A shared encoder paired with lightweight task-specific decoders enables a unified multi-task platform. Extensive experiments across multiple datasets demonstrate that CompDAE achieves state-of-the-art performance with significantly lower complexity, especially under ultra-low-light conditions where traditional CMOS and SCI pipelines fail.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Baraniuk, R.; Davenport, M.; DeVore, R.; and Wakin, M. 2008. A Simple Proof of the Restricted Isometry Property for Random Matrices. Constructive Approximation, 28(3): 253--263
work page 2008
-
[2]
Bethi, Y. R. T.; Narayanan, S.; Rangan, V.; Chakraborty, A.; and Thakur, C. S. 2021. Real-time object detection and localization in compressive sensed video. In ICIP, 1489--1493. IEEE
work page 2021
-
[3]
Birkl, R.; Wofk, D.; and M \"u ller, M. 2023. Midas v3. 1--a model zoo for robust monocular relative depth estimation. arXiv preprint arXiv:2307.14460
arXiv 2023
-
[4]
Bochkovskii, A.; Delaunoy, A.; Germain, H.; Santos, M.; Zhou, Y.; Richter, S. R.; and Koltun, V. 2024. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second. arXiv preprint arXiv:2410.02073
arXiv 2024
-
[5]
Boyd, S.; Parikh, N.; Chu, E.; Peleato, B.; Eckstein, J.; et al. 2011. Distributed optimization and statistical learning via the alternating direction method of multipliers. Foundations and Trends in Machine learning, 3(1): 1--122
work page 2011
-
[6]
Candes, E.; and Tao, T. 2005. Decoding by linear programming. IEEE Transactions on Information Theory, 51(12): 4203--4215
work page 2005
-
[7]
Candès, E. J.; Romberg, J. K.; and Tao, T. 2006. Stable signal recovery from incomplete and inaccurate measurements. Communications on Pure and Applied Mathematics, 59(8): 1207–1223
work page 2006
-
[8]
Canny, J. 1986. A computational approach to edge detection. IEEE TPAMI, PAMI-8(6): 679--698
work page 1986
Show all 75 references
-
[9]
Chen, M.; Radford, A.; Child, R.; Wu, J.; Jun, H.; Luan, D.; and Sutskever, I. 2020. Generative pretraining from pixels. In ICML, 1691--1703. PMLR
2020
-
[10]
Cheng, Z.; Chen, B.; Liu, G.; Zhang, H.; Lu, R.; Wang, Z.; and Yuan, X. 2021. Memory-efficient network for large-scale video compressive sensing. In CVPR, 16246--16255
2021
-
[11]
Cheng, Z.; Lu, R.; Wang, Z.; Zhang, H.; Chen, B.; Meng, Z.; and Yuan, X. 2020. BIRNAT: Bidirectional recurrent neural networks with adversarial training for video snapshot compressive imaging. In ECCV, 258--275. Springer
2020
-
[12]
Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; and Schiele, B. 2016. The Cityscapes Dataset for Semantic Urban Scene Understanding. In CVPR
2016
-
[13]
Dadkhah, M.; Jamal Deen, M.; and Shirani, S. 2014. Block-Based CS in a CMOS Image Sensor. IEEE Sensors Journal, 14(8): 2897--2909
2014
-
[14]
Dong, X.; Wang, G.; Pang, Y.; Li, W.; Wen, J.; Meng, W.; and Lu, Y. 2011. Fast efficient algorithm for enhancement of low lighting video. In ICME. IEEE
2011
-
[15]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:2010.11929
2020 arXiv
-
[16]
Foi, A.; Trimeche, M.; Katkovnik, V.; and Egiazarian, K. 2008. Practical Poissonian-Gaussian noise modeling and fitting for single-image raw-data. IEEE TIP, 17(10): 1737--1754
2008
-
[17]
Gao, L.; Liang, J.; Li, C.; and Wang, L. V. 2014. Single-shot compressed ultrafast photography at one hundred billion frames per second. Nature, 516(7529): 74--77
2014
-
[18]
Girshick, R.; Donahue, J.; Darrell, T.; and Malik, J. 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 580--587
2014
-
[19]
Gulve, R.; Sarhangnejad, N.; Dutta, G.; Sakr, M.; Nguyen, D.; Rangel, R.; Chen, W.; Xia, Z.; Wei, M.; Gusev, N.; Lin, E. Y. H.; Sun, X.; Hanxu, L.; Katic, N.; Abdelhadi, A.; Moshovos, A.; Kutulakos, K. N.; and Genov, R. 2022. A 39,000 Subexposures/s CMOS Image Sensor with Dual...
2022
-
[20]
Guo, B.; Han, Y.; and Wen, J. 2019. Agem: Solving linear inverse problems via deep priors and sampling. NeurIPS, 32
2019
-
[21]
He, J.; Zhang, S.; Yang, M.; Shan, Y.; and Huang, T. 2019. Bi-directional cascade network for perceptual edge detection. In CVPR, 3828--3837
2019
-
[22]
He, K.; Chen, X.; Xie, S.; Li, Y.; Doll \'a r, P.; and Girshick, R. 2022. Masked autoencoders are scalable vision learners. In CVPR, 16000--16009
2022
-
[23]
Hitomi, Y.; Gu, J.; Gupta, M.; Mitsunaga, T.; and Nayar, S. K. 2011. Video from a single coded exposure photograph using a learned over-complete dictionary. In ICCV, 287--294
2011
-
[24]
Hu, C.; Huang, H.; Chen, M.; Yang, S.; and Chen, H. 2021 a . Video object detection from one single image through opto-electronic neural network. APL Photonics, 6(4)
2021
-
[25]
Hu, C.; Huang, H.; Chen, M.; Yang, S.; and Chen, H. 2021 b . Video object detection from one single image through opto-electronic neural network. APL Photonics, 6(4): 046104
2021
-
[26]
Hu, M.; Yin, W.; Zhang, C.; Cai, Z.; Long, X.; Chen, H.; Wang, K.; Yu, G.; Shen, C.; and Shen, S. 2024. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE Transactions on Pattern Analysis and Machine Int...
2024
-
[27]
Huang, H.; Hu, C.; Jingwei, L.; Dong, X.; and Chen, H. 2022. CoCoCs: co-optimized compressive imaging driven by high-level vision. Optics Express, 30: 30894--30910
2022
-
[28]
Iliadis, M.; Spinoulas, L.; and Katsaggelos, A. K. 2020. Deepbinarymask: Learning a binary mask for video compressive sensing. Digital Signal Processing, 96: 102591
2020
-
[29]
Inagaki, Y.; Kobayashi, Y.; Takahashi, K.; Fujii, T.; and Nagahara, H. 2018. Learning to Capture Light Fields Through a Coded Aperture Camera. In ECCV, 431--448. ISBN 978-3-030-01234-2
2018
-
[30]
Jiang, X.; Raskutti, G.; and Willett, R. 2015. Minimax optimal rates for Poisson inverse problems with physical constraints. IEEE TIP, 61(8): 4458--4474
2015
-
[31]
Karpathy, A.; Toderici, G.; Shetty, S.; Leung, T.; Sukthankar, R.; and Fei-Fei, L. 2014. Large-scale video classification with convolutional neural networks. In CVPR, 1725--1732
2014
-
[32]
Khademi, W.; Rao, S.; Minnerath, C.; Hagen, G.; and Ventura, J. 2021. Self-supervised poisson-gaussian denoising. In WACV, 2131--2139
2021
-
[33]
B.; et al
Kinga, D.; Adam, J. B.; et al. 2015. A method for stochastic optimization. In ICLR, volume 5, 6. San Diego, California
2015
-
[34]
Kwan, C.; Chou, B.; Yang, J.; Rangamani, A.; Tran, T.; Zhang, J.; and Etienne-Cummings, R. 2019. Target tracking and classification using compressive measurements of MWIR and LWIR coded aperture cameras. Journal of Signal and Information Processing, 10(3): 73--95
2019
-
[35]
Li, Y.; Guo, B.; Wen, J.; Xia, Z.; Liu, S.; and Han, Y. 2021. Learning model-blind temporal denoisers without ground truths. In ICASSP, 2055--2059. IEEE
2021
-
[36]
Liang, Y.; Huang, H.; Jingwei, L.; Dong, X.; Chen, M.; Yang, S.; and Chen, H. 2022. Action recognition based on discrete cosine transform by optical pixel-wise encoding. APL Photonics, 7
2022
-
[37]
Liao, X.; Li, H.; and Carin, L. 2014. Generalized alternating projection for weighted-2,1 minimization with applications to model-based compressive sensing. SIAM Journal on Imaging Sciences, 7(2): 797--823
2014
-
[38]
Liu, X.; Zhu, M.; Zheng, S.; Luo, R.; Wu, H.; and Yuan, X. 2024. Video snapshot compressive imaging using adaptive progressive coding for high-quality reconstruction under different illumination circumstances. Opt. Lett., 49(1): 85--88
2024
-
[39]
Liu, Y.; Cheng, M.-M.; Hu, X.; Wang, K.; and Bai, X. 2017. Richer convolutional features for edge detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 3000--3009
2017
-
[40]
J.; and Dai, Q
Liu, Y.; Yuan, X.; Suo, J.; Brady, D. J.; and Dai, Q. 2018. Rank minimization for snapshot compressive imaging. IEEE TPAMI, 41(12): 2990--3006
2018
-
[41]
Llull, P.; Liao, X.; Yuan, X.; Yang, J.; Kittle, D.; Carin, L.; Sapiro, G.; and Brady, D. J. 2013. Coded aperture compressive temporal imaging. Optics express, 21(9): 10526--10545
2013
-
[42]
Lu, S.; Yuan, X.; and Shi, W. 2020. Edge compression: An integrated framework for compressive imaging processing on cavs. In 2020 IEEE/ACM Symposium on Edge Computing (SEC), 125--138. IEEE
2020
-
[43]
Meng, Z.; Qiao, M.; Ma, J.; Yu, Z.; Xu, K.; and Yuan, X. 2020. Snapshot multispectral endomicroscopy. Opt. Lett., 45(14): 3897--3900
2020
-
[44]
Okawara, T.; Yoshida, M.; Nagahara, H.; and Yagi, Y. 2020. Action recognition from a single coded image. In ICCP, 1--11. IEEE
2020
-
[45]
Ottewill, A.; and Rahman, M. 1997. Correlation Function and Power Spectrum of Non-Stationary Shot Noise. Technical Report LIGO-T970085-00-D, LIGO Project, California Institute of Technology and Massachusetts Institute of Technology. Preliminary version, printed June 12, 1997; ...
1997
-
[46]
Pathak, D.; Krahenbuhl, P.; Donahue, J.; Darrell, T.; and Efros, A. A. 2016. Context encoders: Feature learning by inpainting. In CVPR, 2536--2544
2016
-
[47]
V.; and Yu, F
Piccinelli, L.; Yang, Y.-H.; Sakaridis, C.; Segu, M.; Li, S.; Gool, L. V.; and Yu, F. 2024. UniDepth: Universal Monocular Metric Depth Estimation. In CVPR, 10106--10116
2024
-
[48]
S.; Riba, E.; and Sappa, A
Poma, X. S.; Riba, E.; and Sappa, A. 2020. Dense extreme inception network: Towards a robust cnn model for edge detection. In WACV, 1923--1932
2020
-
[49]
Pont-Tuset, J.; Perazzi, F.; Caelles, S.; Arbel\'aez, P.; Sorkine-Hornung, A.; and Van Gool , L. 2017. The 2017 DAVIS Challenge on Video Object Segmentation. arXiv preprint arXiv:1704.00675
2017 arXiv
-
[50]
Reddy, D.; Veeraraghavan, A.; and Chellappa, R. 2011. P2C2: Programmable pixel compressive camera for high speed imaging. In CVPR, 329--336
2011
-
[51]
Sun, L.; Wu, F.; Ding, W.; Li, X.; Lin, J.; Dong, W.; and Shi, G. 2024. Multi-Scale Spatio-Temporal Memory Network for Lightweight Video Denoising. IEEE Transactions on Image Processing, 33: 5810--5823
2024
-
[52]
Uhrig, J.; Schneider, N.; Schneider, L.; Franke, U.; Brox, T.; and Geiger, A. 2017. Sparsity Invariant CNNs. In International Conference on 3D Vision (3DV)
2017
-
[53]
N.; Kaiser, L.; and Polosukhin, I
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is all you need. In NeurIPS, 6000–6010. Curran Associates Inc
2017
-
[54]
Vincent, P.; Larochelle, H.; Bengio, Y.; and Manzagol, P.-A. 2008. Extracting and composing robust features with denoising autoencoders. In ICML, 1096--1103
2008
-
[55]
Wang, L.; Cao, M.; and Yuan, X. 2023. Efficientsci: Densely connected network with space-time factorization for large-scale video snapshot compressive imaging. In CVPR, 18477--18486
2023
-
[56]
Wang, L.; Cao, M.; Zhong, Y.; and Yuan, X. 2022. Spatial-temporal transformer for video snapshot compressive imaging. IEEE TPAMI, 45(7): 9072--9089
2022
-
[57]
Wang, N.; and Yeung, D.-Y. 2013. Learning a deep compact image representation for visual tracking. NeurIPS, 26
2013
-
[58]
Wang, P.; Wang, L.; and Yuan, X. 2023. Deep Optics for Video Snapshot Compressive Imaging. In ICCV, 10612--10622
2023
-
[59]
Wen, J.; and Villasenor, J. 1999. Structured prefix codes for quantized low-shape-parameter generalized Gaussian sources. IEEE Transactions on Information Theory, 45(4): 1307--1314
1999
-
[60]
Wu, J.; Wen, J.; and Han, Y. 2025. BackSlash: Rate Constrained Optimized Training of Large Language Models
2025
-
[61]
Wu, Z.; Zhangt, J.; and Mou, C. 2021. Dense deep unfolding network with 3D-CNN prior for snapshot compressive imaging. In ICCV. IEEE
2021
-
[62]
Yoshida, M.; Sonoda, T.; Nagahara, H.; Endo, K.; Sugiyama, Y.; and Taniguchi, R.-i. 2020. High-Speed Imaging Using CMOS Image Sensor With Quasi Pixel-Wise Exposure. IEEE Transactions on Computational Imaging, 6: 463--476
2020
-
[63]
Yoshida, M.; Torii, A.; Okutomi, M.; Endo, K.; Sugiyama, Y.; Taniguchi, R.; and Nagahara, H. 2018. Joint optimization for compressive video sensing and reconstruction under hardware constraints. In Hebert, M.; Ferrari, V.; Sminchisescu, C.; and Weiss, Y., eds., ECCV, Lecture N...
2018
-
[64]
Yuan, X. 2016. Generalized alternating projection based total variation minimization for compressive sensing. In ICIP, 2539--2543. IEEE
2016
-
[65]
Yuan, X.; Liu, Y.; Suo, J.; and Dai, Q. 2020. Plug-and-play algorithms for large-scale snapshot compressive imaging. In CVPR, 1447--1457
2020
-
[66]
Yuan, X.; and Pang, S. 2016. Structured illumination temporal compressive microscopy. Biomed. Opt. Express, 7(3): 746--758
2016
-
[67]
Zhang, J.; Xiong, T.; Tran, T.; Chin, S.; and Etienne-Cummings, R. 2016. Compact all-CMOS spatiotemporal compressive sensing video camera with pixel-wise coded exposure. Optics Express, 24: 9013--9024
2016
-
[68]
Zhang, R.; Isola, P.; and Efros, A. A. 2016. Colorful image colorization. In ECCV, 649--666. Springer
2016
-
[69]
J.; and Dai, Q
Zhang, Z.; Zhang, B.; Yuan, X.; Zheng, S.; Su, X.; Suo, J.; Brady, D. J.; and Dai, Q. 2022. From compressive sampling to compressive tasking: retrieving semantics in compressed domain with low bandwidth. PhotoniX, 3(1): 19
2022
-
[70]
Zhao, H.; Shi, J.; Qi, X.; Wang, X.; and Jia, J. 2017. Pyramid Scene Parsing Network. In CVPR, 6230--6239
2017
-
[71]
Zhao, M.; and Jalali, S. 2023. Theoretical analysis of binary masks in snapshot compressive imaging systems. In 2023 59th Annual Allerton Conference on Communication, Control, and Computing (Allerton), 1--8. IEEE
2023
-
[72]
Zheng, S.; Yang, X.; and Yuan, X. 2022. Two-Stage is Enough: A Concise Deep Unfolding Reconstruction Network for Flexible Video Compressive Sensing. arXiv preprint arXiv:2201.05810
2022 arXiv
-
[73]
Zhu, S.; and Liu, X. 2023. LightedDepth: Video Depth Estimation in light of Limited Inference View Angles. In CVPR
2023
-
[74]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[75]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.