Pith. sign in

REVIEW 3 major objections 5 minor 3 references

Color-Blind Image Sensors: Towards Digital Twin of Human Retina

T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read A camera sensor whose filter array copies the human fovea can reproduce both normal and color-deficient vision, the authors demonstrate.

desk verdict A physically fabricated quasi-random CFA from a real foveal patch, plus a protanopic version; solid engineering, but the spectral fidelity claims need quantification. read the letter →

arxiv 2509.08518 v1 pith:2FE645SN submitted 2025-09-10 physics.optics

classification physics.optics
keywords imagesensorquasi-randomcolorfilterarrayfovealconemosaicvisiondeficiencyprotanopiademosaicingneuralnetworkreconstructionbio-inspired
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a camera sensor can faithfully emulate human color sampling if its color filter array copies the quasi-random spatial distribution of foveal cone cells, instead of using the periodic red-green-blue grid of conventional cameras. The authors fabricate such a quasi-random array from a real patch of human fovea, with long-, medium-, and short-wavelength cones in their natural proportions, bond it to a monochrome CMOS sensor, and train a fully convolutional neural network to turn the raw mosaiced output into full-color sRGB images. They then build a second array that removes the long-wavelength cone band, creating a sensor that mimics red-green color deficiency. Its reconstructed images of standard color-vision test plates show the numerals expected for a protanopic or protanomalous observer, while normal-vision plates reconstruct correctly on the first sensor. If the claim holds, this provides a physical, tunable platform for studying color vision deficiencies and for building personalized digital-twin retinas from an individual's own cone mosaic.

What carries the argument

The quasi-random color filter array (Q-CFA): a photolithographically fabricated mosaic whose spatial statistics are taken directly from a real foveal cone patch (492 L, 236 M, 39 S cones), and whose three transmission bands are designed to approximate L, M, and S cone spectral sensitivities. The array is the load-bearing object because it replaces periodic sampling with retina-like quasi-random sampling; the companion P-CFA removes the L-like band to simulate protanopia. The second piece of machinery is the fully convolutional U-Net neural network, which solves the otherwise ill-posed 'random demosaicing' problem and combines it with raw RGBM-to-sRGB color conversion, effectively learning th

What would settle it

Capture a set of metameric pairs—pairs of distinct light spectra that normal human vision matches as the same color—with the Q-CFA camera under controlled illumination. If the reconstructed images separate the two members of any pair into visibly different colors, the sensor's spectral sensitivities are not cone-equivalent and the eye-like sampling claim is refuted. For the protanopic sensor, a second check is chromaticity collapse: pure protanopia compresses color to a line on the CIE diagram, while the paper's own plots show a triangle; finding a real protanopic observer who sees more color

Watch

Extended reading notes

Core claim

The central discovery is that a physical color filter array patterned after a real foveal cone mosaic—not a periodic grid—can be integrated with a monochrome sensor and a neural-network decoder to produce both normal and color-deficient vision. The quasi-random array (Q-CFA) uses a retinal patch containing 492 L, 236 M, and 39 S cones, repeated across a 20.48-megapixel sensor, with filter spectra whose peaks and overlaps approximate human cone sensitivities. A U-Net, which the authors connect to the concept of receptive fields in retinal circuitry, performs demosaicing together with raw-to-sRGB color conversion. For color blindness, the protanopic array (P-CFA) is designed with an L:M:S cone

Load-bearing premise

The claim rests on the assumption that the fabricated filter spectra are close enough to real human cone sensitivities—especially the deliberately red-shifted red channel—that the sensor's sampling really represents how a retina samples color; the paper does not quantify how much fidelity is lost by that shift.

Editorial extensions

If this is right

  • The same fabrication and training pipeline can produce a sensor from any individual's cone mosaic, enabling personalized simulations of color vision deficiencies rather than population averages.
  • Quasi-random mosaics, once paired with a learned decoder, can be used without sacrificing image quality; the paper shows the network generalizes to random CFA patterns beyond the specific fabricated array.
  • A protanopic variant is the first of a family: changing the band set and cone ratio should yield deutan and tritan emulations on the same sensor hardware.
  • Because the physical filter array, not just software, defines the color deficiency, the camera can serve as a testbed for personalized corrective filters and for studying how high-frequency patterns are perceived in color-deficient vision.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The deliberate red-shift of the Q-CFA's red filter means the 'normal vision' sensor is not spectrally identical to a real retina; a useful next step would be measuring the sensor's color-matching functions against human observer data to quantify this gap.
  • The color-blind validation depends heavily on the trained network: the same P-CFA mosaiced raw data fed through a different decoder would likely produce different 'percepts.' Isolating the mosaic's contribution with a linear least-squares demosaicer would clarify what is hardware and what is learned.
  • The random-sampling advantage for high-frequency patterns suggests a testable extension: measuring how the Q-CFA camera resolves achromatic fine detail compared with a Bayer camera of the same pixel count, particularly for textures and camouflage-like stimuli.
  • If personalized mosaics become practical, the sensor could be used to search for filter patterns that improve an individual's performance on real-world color tasks, effectively using the hardware itself as an optimization target for color-blindness aids.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports the design, fabrication, and integration of a quasi-random color filter array (Q-CFA) on a monochrome CMOS image sensor. The Q-CFA mimics the spatial distribution and spectral responses of foveal L, M, and S cones, with a U-Net neural network performing demosaicing and RGBM-to-sRGB conversion. A second prototype, the protanopic CFA (P-CFA), is fabricated with only S- and M-like bands and used with the same reconstruction network to emulate protanopia/protanomaly. Validation is performed with Macbeth ColorChecker charts and Ishihara plates, including comparisons with Affinity Photo 2 protanopia simulations and CIE chromaticity diagrams. The stated goal is a digital twin of the human retina for studying color vision deficiencies.

Significance. The work is technically ambitious and demonstrates, for the first time to my knowledge, a fabricated image sensor whose color filter geometry is derived from real foveal cone mosaics, and a second sensor designed to mimic a color-deficient retina. The use of a real retinal patch, the hybrid synthetic/hardware training set, and the explicit validation against Ishihara plates are commendable. If the spectral fidelity claim can be quantified, this would be a meaningful contribution to bio-inspired imaging and color-vision research. However, the central claim of 'similar spectral characteristics' is currently supported only qualitatively, and the color-blindness validation is entangled with the neural-network reconstruction. These gaps are load-bearing and need to be addressed before the result can be accepted as stated.

major comments (3)
  1. [Results, 'The quasi-random color filter array (Q-CFA) for eye-inspired image sensor'; Fig. 2C; Discussion] The manuscript states that the red filter peak was deliberately shifted to longer wavelengths to reduce neural-network complexity, and later acknowledges that this limits the ability to 'perfectly emulate the retina.' No quantitative measure of the deviation from the L-cone spectral sensitivity (e.g., centroid shift, spectral similarity coefficient, or colorimetric error) is reported. Since the 'similar spectral characteristics' claim is central, this unquantified shift leaves the core assertion unsupported. Please report a metric comparing the measured filter spectra to the Stockman-Brainard cone fundamentals, and assess how the shift affects downstream color accuracy.
  2. [Results, 'A color-blind image sensor based on a protanopic color filter array (P-CFA)'; Fig. 5F; Table 3] The protanopic behavior is validated only on neural-network-reconstructed images, and the network was trained using data from the same sensor (hybrid synthetic/hardware method). Thus the observed loss of red-green information and the appearance of Ishihara numbers could be partly an artifact of the learned color transform rather than a faithful sensor-level spectral property. The paper should validate the P-CFA at the raw-data level, e.g., by comparing measured band responses of the sensor to predicted cone-excitation ratios, or by simulating the full optical-to-RGBM chain without the trained network. Additionally, Table 3's CIE diagrams are qualitative; provide a quantitative gamut comparison with established protanopia/protanomaly models and report repeatability.
  3. [Results, 'Full-color image recovery using a deep neural network'; Fig. 4C; Discussion] Color reconstruction accuracy is described only qualitatively ('satisfactory color accuracy') for the Macbeth ColorChecker and Ishihara plates. No quantitative color-difference metric (e.g., CIEDE2000 or mean ΔE) is reported, and there are no error bars or repeated-capture statistics. Given the paper's central claim that the sensor faithfully represents human retinal sampling, a quantitative colorimetric evaluation is essential. The Discussion also notes that pixels with severe crosstalk are replaced by neural-network estimates; the impact of this replacement on measured color accuracy should be quantified.
minor comments (5)
  1. [Abstract; Discussion] The abstract claims 'similar distribution, spacing, ratios and spectral characteristics of an actual foveal mosaic,' but the Q-CFA is a repeated 24×24 arcmin patch rather than a full fovea. The Discussion acknowledges this; the abstract should be tempered or the limitation stated more prominently.
  2. [Fig. 3B; Fig. 3C; Methods] Measured filter pixel intensities are shown as mean values without error bars or per-pixel variance, despite the Methods noting 20 sampled points per channel. Adding error bars would strengthen the spectral-characterization claims.
  3. [Supplementary references] The text references Sections 1–8 and Fig. S4–S8 in the supplementary information, but the supplementary material is not included with the arXiv submission. Readers cannot verify the demosaicing performance claims or the crosstalk-mitigation procedure.
  4. [Fig. 1; Section 'Full-color image recovery'] The symbol M is used both for the monochrome (rod-like) channel and for the M cone class. This dual use is potentially confusing; consider using 'W' or 'Mono' for the achromatic channel.
  5. [Table 1] For the P-CFA, the cone population ratio is given as 0:0.9:0.1 (L:M:S). Since the manuscript discusses protanomaly as well as protanopia, state explicitly whether this ratio is intended to model absent L cones or a reduced-cone scenario.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the sensor's color-blind behavior is a physical consequence of the two-band P-CFA and is validated against external Ishihara/Affinity Photo benchmarks; self-citations are background only.

full rationale

The paper's derivation chain is not circular. The Q-CFA design uses a real retinal patch (Sabesan et al. 2015) and measured filter spectra; the claimed similarity to foveal mosaic is a design specification, not a prediction fitted to itself. The P-CFA explicitly sets the cone population ratio L:M:S to 0:0.9:0.1 (Table 1), so the absence of L cones is by construction, but the paper then independently validates protanopic behavior by (a) comparing measured band-2 spectra against L/M cone sensitivities via a 100,000-spectrum correlation analysis, (b) reconstructing Ishihara plates through the trained neural network and comparing to expected protanopic views, and (c) plotting CIE diagrams against Affinity Photo 2 protanopia simulations. The neural network is a demosaicing and color-space-conversion tool trained on sensor data to map raw RGBM to sRGB; it is not trained to output protanopic images, and the Ishihara test images are external. The acknowledged red-filter peak shift is a stated limitation ('it limits the ability of the mosaic to perfectly emulate the retina'), not a circular step. Self-citations (He et al. 2020/2021; Kamalakar et al. 2005) appear only as background or application mentions and are not load-bearing for the central claims. No uniqueness theorem is imported from the authors, and no ansatz is smuggled in via citation. The paper is self-contained against external benchmarks, so the circularity score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the accuracy of reference cone spectra, the representativeness of one participant's foveal patch, and the ISETBio protanopic mosaic model. The red filter peak shift is an ad hoc design choice that weakens spectral fidelity.

free parameters (2)
  • Red filter peak wavelength shift = Shifted towards longer wavelength (exact value not given in main text)
    The red channel peak was deliberately moved from the typical L-cone peak to simplify the neural network, affecting the spectral similarity claim.
  • Cone population ratio for P-CFA = L:M:S = 0:0.9:0.1
    Chosen to model protanopia, but the actual band 2 filter still overlaps with L-cone sensitivity, making the model a compromise.
assumptions (4)
  • domain assumption Cone spectral sensitivities from Bowmaker and Dartnall (1980) are an accurate representation of human cone responses.
    Used as the reference for filter design and comparison in Fig. 2C and Fig. 5F.
  • domain assumption The retinal patch from Sabesan et al. (2015) is representative of a typical human foveal mosaic.
    The Q-CFA is created by repeating one 492L/236M/39S cone patch; individual variation is not addressed.
  • domain assumption ISETBio-generated foveal mosaic accurately models protanopic cone arrangement.
    The P-CFA design relies on ISETBio output for cone positions and sizes.
  • ad hoc to paper The neural network's RMSE on a synthetic validation set is a valid measure of reconstruction quality for real scenes.
    The network is trained and selected on synthetic data from the same sensor model, then qualitatively validated on real Ishihara plates; no quantitative metric on real images is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Color-Blind Image Sensors: Towards Digital Twin of Human Retina." pith.science (2026). https://pith.science/paper/2FE645SN

@misc{pith2026250908518,
  author       = {Pith},
  title        = {Pith review of: Color-Blind Image Sensors: Towards Digital Twin of Human Retina},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2FE645SN}},
  note         = {Machine review of arXiv:2509.08518}
}
read the original abstract

The human retina contains a complex arrangement of photoreceptors that convert light into visual information. Conventional image sensors mimic the trichromacy of the retina using periodic filter mosaics responsive to three primary colors. However, this is, at best, an approximation, as an actual retina exhibits a quasi-random spatial distribution of light-sensitive rod and cone photoreceptors, where the ratio of rods to cones and their concentrations vary across the retina. Hence, the periodic mosaics are limited to accurately simulate the properties of the eye. Here, we present an image sensor with similar distribution, spacing, ratios and spectral characteristics of an actual foveal mosaic for emulating eye-like sampling and mimicking color blindness. To perform image reconstruction, we use a fully convolutional U-Net neural network adopting the concept of receptive fields in the retinal circuitry. Our research will enable the development of digital twin of a retina to further understand color vision deficiencies.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 2 canonical work pages

  1. [69]

    https://doi.org/10.1017/CBO9781107337930.004 Vision, N.R.C. (US) C. on, 1981. COLOR VISION TESTS, in: Procedures for Testing Color Vision: Report of Working Group 41. National Academies Press (US). Webster, M.A., 2015. Individual differences in color vision, in: Elliot, A.J., Franklin, A., Fairchild, M.D. (Eds.), Handbook of Color Psychology, Cambridge Ha...

  2. [238]

    A single sensor based multispectral imaging camera using a narrow spectral band color mosaic integrated on the monochrome CMOS image sensor

    https://doi.org/10.1364/OSAC.413709 He, X., Liu, Y., Ganesan, K., Ahnood, A., Beckett, P., Eftekhari, F., Smith, D., Uddin, M.H., Skafidas, E., Nirmalathas, A., Unnithan, R.R., 2020. A single sensor based multispectral imaging camera using a narrow spectral band color mosaic integrated on the monochrome CMOS image sensor. APL Photonics 5, 046104. https://...

  3. [2023]

    Tho’ she kneel’d in that place where they grew…

    Cuttlefish eye–inspired artificial vision for high-quality imaging under uneven illumination conditions. Sci. Robot. 8, eade4698. https://doi.org/10.1126/scirobotics.ade4698 Kim, M.S., Yeo, J.-E., Choi, H., Chang, S., Kim, D.-H., Song, Y.M., 2023. Evolution of natural eyes and biomimetic imaging devices for effective image acquisition. J. Mater. Chem. C 1...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.