REVIEW 1 major objections 7 minor 60 references
AI Makes Plausible Images But Gets the Imaging Physics Wrong
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-09 17:40 UTC pith:IHGE4NRX
load-bearing objection First systematic benchmark showing frontier VLMs lag specialized methods on physics-grounded computational imaging, but the central comparative claim is overgeneralized relative to the baselines actually provided. the 1 major comments →
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Frontier multimodal AI systems can generate visually plausible images but cannot reliably solve computational imaging problems that require understanding and inverting the physics of image formation. The gap is largest for computational sensing tasks where a structured forward operator must be explicitly inverted, and adaptive planning does not meaningfully close it. Visually natural outputs and physically correct reconstructions are not the same thing, and current models produce the former without achieving the latter.
What carries the argument
The benchmark formalizes each imaging task as a forward model x = A(z; m) + n, where z is the latent clean signal, A is the task-dependent forward operator, and n is noise. Three evaluation protocols isolate different competences: Expert tests fixed-prompt inverse execution, Planner tests adaptive per-image planning by a VLM before execution, and Forward tests whether the model can simulate the forward degradation process itself. Performance is measured with reference-based metrics (PSNR, SSIM, LPIPS) and the no-reference metric NIQE, then normalized into a unified score for cross-task comparison.
Load-bearing premise
The benchmark assumes that the three evaluation protocols and their fixed prompt templates adequately probe whether models understand imaging physics, rather than merely testing their sensitivity to how a task is described. The paper itself acknowledges that the protocols do not prove the model causally uses the forward model rather than generic image priors, yet the central conclusion about a gap between semantic and physical competence depends on this assumption holding.
What would settle it
If a frontier model were shown to produce high-fidelity reconstructions on the computational sensing tasks — lensless imaging, event-based reconstruction, time-of-flight, holography — with reference-based metrics matching specialized baselines, the claimed gap between visual plausibility and physical fidelity would collapse for that model class.
If this is right
- If the gap between visual plausibility and physical fidelity is real, then benchmarks that score only perceptual quality will systematically overestimate AI competence on safety-critical imaging tasks in medicine, remote sensing, and scientific imaging.
- The near-zero benefit of adaptive planning suggests that the bottleneck for agentic imaging is not instruction quality but the executor's lack of operator-aware reasoning, meaning progress will likely require architectural or training changes rather than better prompting.
- The Forward protocol's results imply that models struggle not only to invert physics but to simulate it, which would undermine any pipeline that relies on a VLM to generate training data or consistency checks for imaging systems.
- If specialized non-agentic methods remain substantially stronger, there is a practical case for hybrid systems where VLMs handle semantic reasoning and specialized solvers handle physical inversion, rather than end-to-end agentic pipelines.
Where Pith is reading between the lines
- The finding that models produce natural-looking but physically incorrect outputs on sensing tasks suggests they are applying learned image priors rather than reasoning about the measurement process — a failure mode that is invisible to human evaluators and to no-reference quality metrics.
- The fact that planner guidance rarely helps could indicate that the planner models themselves cannot diagnose the degradation type from visual inspection alone, which would mean the observe-plan-execute loop breaks at the observation stage rather than the execution stage.
- A natural next test would be to give the planner explicit access to the forward operator parameters (sampling ratio, noise statistics, PSF) rather than asking it to infer them from the image, which would isolate whether the failure is in diagnosis or in execution.
- The consistent superiority of specialized methods on sensing tasks raises the question of whether the gap would narrow if VLMs were fine-tuned on paired (measurement, reconstruction) data from specific forward operators, or whether the architectural mismatch between autoregressive/diffusion generation and inverse problem structure is more fundamental.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ImagingBench, a benchmark of 20 computational imaging tasks across five categories (ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration). Three evaluation protocols are defined: Expert (fixed expert-guided inverse reconstruction), Planner (planner-guided inverse reconstruction), and Forward (forward-system simulation for consistency checking). The authors evaluate Gemini, GPT, and Qwen image-editing models, finding that agentic models remain consistently weaker than specialized non-agentic baselines, especially on computational sensing tasks, and that planner guidance provides only modest gains over fixed expert prompts. The benchmark is well-motivated: it addresses a genuine gap between semantic vision benchmarks and physics-grounded computational imaging evaluation.
Significance. The paper makes a timely and useful contribution by providing the first unified benchmark spanning physical image formation, inverse reconstruction, calibration, and planner-executor evaluation for agentic AI. The three-protocol design (Expert, Planner, Forward) is a thoughtful framework for disentangling execution ability, per-instance planning, and forward-model consistency. The inclusion of ablation studies (noise sweeps, sampling-ratio sweeps, spherical-aberration sweeps) and cost analysis adds practical value. The finding that visually plausible outputs do not correspond to physically accurate solutions is important for the community. The benchmark infrastructure and task diversity are commendable.
major comments (1)
- The central claim that agentic models are 'consistently weaker than specialized methods' is directly supported by head-to-head comparison on only 9 of 20 tasks; 11 tasks in Table 3 report '—' for the Non-agent SoTA column. Critically, for computational sensing—the category the abstract emphasizes as where agentic models are 'especially' weaker—only lensless imaging (1/5 tasks) has a direct baseline. The other four sensing tasks (lightfield extrapolation, lightfield interpolation, event-based intensity, ToF depth) lack baselines entirely. The abstract specifically names 'event-based reconstruction' and 'time-of-flight imaging' as examples of where agentic models struggle most, yet these are exactly the tasks without specialized comparison. The claim therefore rests on low absolute PSNR values (e.g., ToF at ~5 dB, event-based at ~7-10 dB) rather than direct head-to-head evaluation. While 5
minor comments (7)
- Figure 1 caption refers to 'CIBench' but the benchmark is named 'ImagingBench' throughout the rest of the paper. This inconsistency should be corrected.
- Table 2 lists Comp. Gen. Holography under both 'Ray and Wave Optics' and 'Computational Sensing' categories. The rationale for this dual placement should be clarified.
- The normalized aggregate score (Section 3.3) uses hand-tuned weights (w_psnr=0.3, w_ssim=0.3, w_lpips=0.3, w_niqe=0.1) and normalization ranges (PSNR [15,40], NIQE [3,20]). A sensitivity analysis showing how the leaderboard ranking changes under alternative weight choices would strengthen confidence in the aggregate score's robustness.
- Section 3.4 mentions 'Nano Banana 2' as the Gemini model name but Table 3 and other references use 'Gemini-3.1-Flash-Image'. The naming should be made consistent.
- Table 3 uses color coding for metric values but the color scale is not always legible in print. Adding explicit numerical thresholds for 'good' and 'bad' ranges in the caption would aid interpretation.
- The paper states (Section 3.2) that the protocols 'do not by themselves prove that the model causally uses the provided forward model rather than generic image priors.' This is an important caveat that should be more prominently discussed, as the central conclusion about the gap between semantic and physical competence depends on this assumption.
- Section 6.5 mentions safety filter refusals from hosted models but does not quantify the refusal rate. Reporting the percentage of calls that triggered safety filters would be useful for practitioners.
Circularity Check
No circularity present.
full rationale
ImagingBench is an empirical benchmark paper, not a derivation paper. Its central claims rest on comparing agentic model outputs against external ground-truth data and independent task-specific baselines (e.g., FFDNet, ESRGAN, AutoLens). The forward models used to generate synthetic degradations (Poisson-Gaussian noise, Zernike aberrations, Bayer mosaicking, compressive sensing masks) are standard physics-based operators from the computational imaging literature, not fitted parameters repackaged as predictions. The normalized scoring (Eqs. 2-5) introduces hand-chosen clipping ranges and weights, but these are transparent aggregation choices for cross-task summarization, not fitted-to-data quantities presented as derived results. The paper explicitly states its protocols 'do not by themselves prove that the model causally uses the provided forward model rather than generic image priors,' which is a self-acknowledged limitation rather than a circular claim. No step in the benchmark construction or evaluation reduces to its own inputs by definition. The absence of non-agentic baselines on some tasks is a coverage gap (correctness risk), not a circularity issue.
Axiom & Free-Parameter Ledger
free parameters (4)
- PSNR normalization range [15, 40] =
15 to 40 dB
- NIQE normalization range [3, 20] =
3 to 20
- Metric weights (w_psnr, w_ssim, w_lpips, w_niqe) =
0.3, 0.3, 0.3, 0.1
- Strehl normalization threshold =
0.8
axioms (3)
- domain assumption Vision-language models that perform well on semantic visual tasks should be tested on physics-grounded tasks to assess transfer.
- ad hoc to paper The three protocols (Expert, Planner, Forward) are sufficient to disentangle execution ability, planning ability, and forward-model consistency.
- domain assumption Standardized 1024x1024 input resolution preserves task-defining structure for all subtasks.
read the original abstract
Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates three complementary settings: Expert, fixed expert-guided inverse reconstruction; Planner, planner-guided inverse reconstruction; and Forward, forward-system simulation for consistency checking. We benchmark leading proprietary and open-source image-centric multimodal systems, including Gemini, GPT, and Qwen, and compare them with representative task-specific non-agentic baselines. Across tasks, agentic models remain consistently weaker than specialized methods, especially on computational sensing problems such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography. Planner guidance provides only modest and inconsistent gains over the fixed-prompt Expert baseline. Although the models often generate visually plausible outputs, their reference-based fidelity remains poor, revealing a substantial gap between semantic visual competence and physically grounded imaging performance. ImagingBench provides a unified testbed for measuring this gap and tracking progress in agentic AI for computational imaging.
Figures
Reference graph
Works this paper leans on
-
[1]
Imagenet large scale visual recognition challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015
work page 2015
-
[2]
The PASCAL visual object classes (VOC) challenge,
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zis- serman, “The PASCAL visual object classes (VOC) challenge,” International Journal of Computer Vision, vol. 88, no. 2, pp. 303–338, 2010
work page 2010
-
[3]
Microsoft COCO: Common objects in context,
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P . Perona, D. Ramanan, P . Doll´ar, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” inEuropean Conference on Computer Vision (ECCV), 2014, pp. 740–755
work page 2014
-
[4]
GLUE: A multi-task benchmark and analysis platform for nat- ural language understanding,
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A multi-task benchmark and analysis platform for nat- ural language understanding,” inProceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2018, pp. 353–355
work page 2018
-
[5]
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,
A. Srivastava, A. Rastogi, A. Raoet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,”Transactions on Machine Learning Research, 2023. [Online]. Available: https://openreview.net/forum?id=uyTL5Bvosj
work page 2023
-
[6]
MMBench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin, “MMBench: Is your multi-modal model an all-around player?” inEuropean Conference on Computer Vision (ECCV), 2024, pp. 216–233
work page 2024
-
[7]
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen, “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,” inProceedings of the IEEE/CVF Conference on Computer Vision a...
work page 2024
-
[8]
AgentBench: Evaluating LLMs as Agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang, “AgentBench: Evaluating LLMs as agents,” 2023. [Online]. Available: https://arxiv.org/abs/2308.03688
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[9]
GAIA: A benchmark for general AI assistants,
G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and T. Scialom, “GAIA: A benchmark for general AI assistants,” inInternational Conference on Learning Representations (ICLR), 2024. [Online]. Available: https://openreview.net/forum?id=fibxvahvs3
work page 2024
-
[10]
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
W. Chow, J. Mao, B. Li, D. Seita, V . Guizilini, and Y. Wang, “Physbench: Benchmarking and enhancing vision-language models for physical world understanding,” inICLR, 2025. [Online]. Available: https://arxiv.org/abs/2501.16411
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[11]
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
J. Yao, Y. Hu, Y. Yi, B. Han, S. Feng, G. Yang, B. Wen, R. Krishna, L. L. Wang, Y. Tsvetkov, N. A. Smith, and B. Zhu, “Mmmg: a comprehensive and reliable evaluation suite for multitask multimodal generation,” 2025. [Online]. Available: https://arxiv.org/abs/2505.17613
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[12]
Com- putational imaging and artificial intelligence: The next revolution of mobile vision,
J. Suo, W. Zhang, J. Gong, X. Yuan, D. J. Brady, and Q. Dai, “Com- putational imaging and artificial intelligence: The next revolution of mobile vision,”Proceedings of the IEEE, vol. 111, no. 12, pp. 1607– 1639, 2023
work page 2023
-
[13]
do: A differentiable engine for deep lens design of computational imaging systems,
C. Wang, N. Chen, and W. Heidrich, “do: A differentiable engine for deep lens design of computational imaging systems,”IEEE Transactions on Computational Imaging, vol. 8, pp. 905–916, 2022
work page 2022
-
[14]
Curriculum learning for ab initio deep learned refractive optics,
X. Yang, Q. Fu, and W. Heidrich, “Curriculum learning for ab initio deep learned refractive optics,”Nature communications, vol. 15, no. 1, p. 6572, 2024
work page 2024
-
[15]
Bayesian-based iterative method of image restoration,
W. H. Richardson, “Bayesian-based iterative method of image restoration,”Journal of the Optical Society of America, vol. 62, no. 1, pp. 55–59, 1972
work page 1972
-
[16]
An iterative technique for the rectification of observed distributions,
L. B. Lucy, “An iterative technique for the rectification of observed distributions,”The Astronomical Journal, vol. 79, pp. 745–754, 1974
work page 1974
-
[17]
Nonlinear total variation based noise removal algorithms,
L. I. Rudin, S. Osher, and E. Fatemi, “Nonlinear total variation based noise removal algorithms,”Physica D: Nonlinear Phenomena, vol. 60, no. 1-4, pp. 259–268, 1992
work page 1992
-
[18]
Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising,
K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang, “Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising,”IEEE transactions on image processing, vol. 26, no. 7, pp. 3142–3155, 2017
work page 2017
-
[19]
Learning a single convolutional super-resolution network for multiple degradations,
K. Zhang, W. Zuo, and L. Zhang, “Learning a single convolutional super-resolution network for multiple degradations,” inProceed- ings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3262–3271
work page 2018
-
[20]
VQA: Visual question answering,
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual question answering,” inProceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, pp. 2425–2433
work page 2015
-
[21]
nocaps: Novel object captioning at scale,
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P . Anderson, “nocaps: Novel object captioning at scale,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 8948–8957
work page 2019
-
[22]
MedMNIST v2: A large-scale lightweight benchmark for 2d and 3d biomedical image classification,
J. Yang, R. Shi, D. Wei, Z. Liu, L. Zhao, B. Ke, H. Pfister, and B. Ni, “MedMNIST v2: A large-scale lightweight benchmark for 2d and 3d biomedical image classification,”Scientific Data, vol. 10, no. 1, p. 41, 2023
work page 2023
-
[23]
CheXpert: A large chest radiograph dataset with uncertainty labels and expert comparison,
J. Irvin, P . Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya, J. Seekins, D. A. Mong, S. S. Halabi, J. K. Sandberg, R. Jones, D. B. Larson, C. P . Langlotz, B. N. Patel, M. P . Lungren, and A. Y. Ng, “CheXpert: A large chest radiograph dataset with uncertainty labels and expert comparison,” inProceeding...
work page 2019
-
[24]
A. E. W. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P . Lungren, C.-y. Deng, R. G. Mark, and S. Horng, “MIMIC- CXR, a de-identified publicly available database of chest radio- graphs with free-text reports,”Scientific Data, vol. 6, p. 317, 2019
work page 2019
-
[25]
VinDr-CXR: 14 An open dataset of chest x-rays with radiologist’s annotations,
H. Q. Nguyen, K. Lam, L. T. Le, H. H. Nguyen, H. H. Pham, H. Tong, D. Dinh, D. Nguyen, M. Dao, V . Vuet al., “VinDr-CXR: 14 An open dataset of chest x-rays with radiologist’s annotations,” Scientific Data, vol. 9, no. 1, p. 429, 2022
work page 2022
-
[26]
The mul- timodal brain tumor image segmentation benchmark (BRATS),
B. H. Menze, A. Jakab, S. Bauer, J. Kalpathy-Cramer, K. Farahani, J. Kirby, Y. Burren, N. Porz, J. Slotboom, R. Wiestet al., “The mul- timodal brain tumor image segmentation benchmark (BRATS),” IEEE Transactions on Medical Imaging, vol. 34, no. 10, pp. 1993–2024, 2015
work page 1993
-
[27]
A dataset of clinically generated visual questions and answers about radiology images,
J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,”Scientific Data, vol. 5, p. 180251, 2018
work page 2018
-
[28]
SLAKE: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,
B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu, “SLAKE: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,” in2021 IEEE 18th International Sympo- sium on Biomedical Imaging (ISBI), 2021, pp. 1650–1654
work page 2021
-
[29]
PathVQA: 30000+ Questions for Medical Visual Question Answering
X. He, Y. Zhang, L. Mou, E. Xing, and P . Xie, “PathVQA: 30000+ questions for medical visual question answering,” 2020. [Online]. Available: https://arxiv.org/abs/2003.10286
work page internal anchor Pith review Pith/arXiv arXiv 2020
-
[30]
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
X. Zhang, C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie, “PMC-VQA: Visual instruction tuning for medical visual question answering,” 2023. [Online]. Available: https://arxiv.org/abs/2305.10415
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[31]
GMAI-MMBench: A comprehensive multimodal evaluation benchmark towards general medical ai,
P . Chen, J. Ye, G. Wang, Y. Li, Z. Deng, W. Li, T. Li, H. Duan, Z. Huang, Y. Su, B. Wang, S. Zhang, B. Fu, J. Cai, B. Zhuang, E. J. Seibel, Y. Qiao, and J. He, “GMAI-MMBench: A comprehensive multimodal evaluation benchmark towards general medical ai,” inAdvances in Neural Information Processing Systems, vol. 37, 2024. [Online]. Available: https://proceed...
work page 2024
-
[32]
MMMG: A massive, multidisciplinary, multi-tier generation benchmark for text-to-image reasoning,
Y. Luo, Y. Yuan, J. Chen, H. Cai, Z. Yue, Y. Yang, F. Z. Daha, J. Li, and Z. Lian, “MMMG: A massive, multidisciplinary, multi-tier generation benchmark for text-to-image reasoning,”
-
[33]
Available: https://arxiv.org/abs/2506.10963
[Online]. Available: https://arxiv.org/abs/2506.10963
-
[34]
J. Chang, V . Sitzmann, X. Dun, W. Heidrich, and G. Wetzstein, “Hybrid optical-electronic convolutional neural networks with optimized diffractive optics for image classification,”Scientific reports, vol. 8, no. 1, p. 12324, 2018
work page 2018
-
[35]
Vision-language model guided image restoration,
C. Yang, R. Dong, and K.-M. Lam, “Vision-language model guided image restoration,” 2025. [Online]. Available: https: //arxiv.org/abs/2512.17292
-
[36]
Optiagent: A physics-driven agentic framework for automated optical design,
Y. Geng, L. Sun, Y. Gao, X. Hu, Z. Yi, X. Qian, W. Hu, J. Bai, and K. Wang, “Optiagent: A physics-driven agentic framework for automated optical design,” 2026. [Online]. Available: https://arxiv.org/abs/2602.23761
-
[37]
Towards real- time photorealistic 3d holography with deep neural networks,
L. Shi, B. Li, C. Kim, P . Kellnhofer, and W. Matusik, “Towards real- time photorealistic 3d holography with deep neural networks,” Nature, vol. 591, no. 7849, pp. 234–239, 2021
work page 2021
-
[38]
Uni- fied reconstruction of static and dynamic scenes from events,
Q. Gao, P . Duan, H. Lou, M. Teng, Z. Cai, X. Chen, and B. Shi, “Uni- fied reconstruction of static and dynamic scenes from events,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 27 914–27 923
work page 2025
-
[39]
Depth restoration in under-display time-of-flight imaging,
X. Qiao, C. Ge, P . Deng, H. Wei, M. Poggi, and S. Mattoccia, “Depth restoration in under-display time-of-flight imaging,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 5, pp. 5668–5683, 2022
work page 2022
-
[40]
Learned reconstructions for practical mask-based lens- less imaging,
K. Monakhova, J. Yurtsever, G. Kuo, N. Antipa, K. Yanny, and L. Waller, “Learned reconstructions for practical mask-based lens- less imaging,”Optics express, vol. 27, no. 20, pp. 28 075–28 090, 2019
work page 2019
-
[41]
When color constancy goes wrong: Correcting improperly white-balanced im- ages,
M. Afifi, B. Price, S. Cohen, and M. S. Brown, “When color constancy goes wrong: Correcting improperly white-balanced im- ages,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 1535–1544
work page 2019
-
[42]
S. A Sharif, R. A. Naqvi, and M. Biswas, “Beyond joint demosaick- ing and denoising: An image processing pipeline for a pixel-bin image sensor,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 233–242
work page 2021
-
[43]
Burst photography for high dynamic range and low-light imaging on mobile cameras,
S. W. Hasinoff, D. Sharlet, R. Geiss, A. Adams, J. T. Barron, F. Kainz, J. Chen, and M. Levoy, “Burst photography for high dynamic range and low-light imaging on mobile cameras,”ACM Transactions on Graphics (Proc. SIGGRAPH Asia), vol. 35, no. 6, 2016
work page 2016
-
[44]
Fpa-cs: Focal plane array-based compressive imaging in short-wave infrared,
H. Chen, M. Salman Asif, A. C. Sankaranarayanan, and A. Veer- araghavan, “Fpa-cs: Focal plane array-based compressive imaging in short-wave infrared,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 2358–2366
work page 2015
-
[45]
Resolution-robust large mask inpainting with fourier convolutions,
R. Suvorov, E. Logacheva, A. Mashikhin, A. Remizova, A. Ashukha, A. Silvestrov, N. Kong, H. Goka, K. Park, and V . Lem- pitsky, “Resolution-robust large mask inpainting with fourier convolutions,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, 2022, pp. 2149–2159
work page 2022
-
[46]
Gemini 3.1 flash-image model card,
Google DeepMind, “Gemini 3.1 flash-image model card,” https://deepmind.google/models/model-cards/ gemini-3-1-flash-image/, 2026, accessed: 2026-02-26
work page 2026
- [47]
-
[48]
C. Wu, J. Li, J. Zhou, J. Lin, K. Gao, K. Yan, S. ming Yin, S. Bai, X. Xu, Y. Chen, Y. Chen, Z. Tang, Z. Zhang, Z. Wang, A. Yang, B. Yu, C. Cheng, D. Liu, D. Li, H. Zhang, H. Meng, H. Wei, J. Ni, K. Chen, K. Cao, L. Peng, L. Qu, M. Wu, P . Wang, S. Yu, T. Wen, W. Feng, X. Xu, Y. Wang, Y. Zhang, Y. Zhu, Y. Wu, Y. Cai, and Z. Liu, “Qwen-image technical repo...
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[49]
A high-quality de- noising dataset for smartphone cameras,
A. Abdelhamed, S. Lin, and M. S. Brown, “A high-quality de- noising dataset for smartphone cameras,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
work page 2018
-
[50]
Image demosaicing: A system- atic survey,
X. Li, B. Gunturk, and L. Zhang, “Image demosaicing: A system- atic survey,” inVisual Communications and Image Processing 2008, vol. 6822. SPIE, 2008, pp. 489–503
work page 2008
-
[51]
Deep multi-scale convolutional neural network for dynamic scene deblurring,
S. Nah, T. H. Kim, and K. M. Lee, “Deep multi-scale convolutional neural network for dynamic scene deblurring,” inCVPR, July 2017
work page 2017
-
[52]
The stanford light field archive (2016),
Stanford Computer Graphics Laboratory, “The stanford light field archive (2016),” https://lightfields.stanford.edu/LF2016. html, 2016, accessed: 2026-02-26
work page 2016
-
[53]
Ntire 2017 challenge on single image super-resolution: Dataset and study,
E. Agustsson and R. Timofte, “Ntire 2017 challenge on single image super-resolution: Dataset and study,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, July 2017
work page 2017
-
[54]
Learning photo- graphic global tonal adjustment with a database of input / output image pairs,
V . Bychkovsky, S. Paris, E. Chan, and F. Durand, “Learning photo- graphic global tonal adjustment with a database of input / output image pairs,” inThe Twenty-Fourth IEEE Conference on Computer Vision and Pattern Recognition, 2011
work page 2011
-
[55]
Phasecam3d—learning phase masks for pas- sive single view depth estimation,
Y. Wu, V . Boominathan, H. Chen, A. Sankaranarayanan, and A. Veeraraghavan, “Phasecam3d—learning phase masks for pas- sive single view depth estimation,” in2019 IEEE International Conference on Computational Photography (ICCP). IEEE, 2019, pp. 1–12
work page 2019
-
[56]
A physics-based noise formation model for extreme low-light raw denoising,
K. Wei, Y. Fu, J. Yang, and H. Huang, “A physics-based noise formation model for extreme low-light raw denoising,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2758–2767
work page 2020
-
[57]
M. Afifi and M. S. Brown, “Deep white-balance editing,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020
work page 2020
-
[58]
Ffdnet: Toward a fast and flexible solution for cnn-based image denoising,
K. Zhang, W. Zuo, and L. Zhang, “Ffdnet: Toward a fast and flexible solution for cnn-based image denoising,”IEEE Transactions on Image Processing, vol. 27, no. 9, pp. 4608–4622, 2018
work page 2018
-
[59]
Banet: A blur-aware attention network for dynamic scene deblurring,
F.-J. Tsai, Y.-T. Peng, C.-C. Tsai, Y.-Y. Lin, and C.-W. Lin, “Banet: A blur-aware attention network for dynamic scene deblurring,” IEEE Transactions on Image Processing, vol. 31, pp. 6789–6799, 2022
work page 2022
-
[60]
Esrgan: Enhanced super-resolution generative adversarial networks,
X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, Y. Qiao, and C. Change Loy, “Esrgan: Enhanced super-resolution generative adversarial networks,” inEuropean Conference on Computer Vision (ECCV), 2018, pp. 0–0
work page 2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.