Pith. sign in

REVIEW 4 major objections 6 minor 64 references

A neuromorphic vision system for open-world visual intelligence

T0 review · 4 major / 6 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read A hardware task-traction mechanism distills task-relevant light fields so open-world vision runs in 193 μs with large accuracy gains.

desk verdict Real co-designed polarization+RRAM front end with a three-stage task-traction pipeline; the 193 μs figure is measured on small silicon, but the headline open-world accuracy and ~30× latency numbers ride mostly on VTEAM-simulated RRAM. read the letter →

arxiv 2607.10066 v1 pith:OLXD7S72 submitted 2026-07-11 eess.IV

classification eess.IV
keywords neuromorphicvisiontasktractionRRAMpolarizationimagingopen-worldperceptioninformationbottleneckautonomousdrivingin-sensorcomputing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that open-world visual perception fails when systems process full light-intensity fields with heavy software models or narrow task-specific sensors. Drawing on dragonfly vision and information-bottleneck ideas, it builds a neuromorphic front end that keeps only task-relevant cues. A polarization-sensitive photodiode array feeds a single RRAM array partitioned for feature traction (select intensity vs polarization by gradient entropy), attention traction (modulate conductance to extract ROIs), and prediction traction (anticipate target motion for the next frame). The integrated stack runs tracking, segmentation, and trajectory prediction in 193 μs. Across eight unstructured conditions and vehicle-mounted driving tests, the authors report accuracy gains of roughly 25–38% over strong baselines and about a 30-fold latency cut. The claim is that task-oriented distillation at the sensor, not larger models or more modalities alone, is what makes perception both fast and reliable outdoors.

What carries the argument

Task traction mechanism: three sequential hardware stages on one RRAM array that select the most informative light field (feature traction), extract temporal ROIs (attention traction), and anticipate the next-frame task region (prediction traction), so only distilled cues reach the visual tasks.

What would settle it

Build a larger physical polarization-plus-RRAM camera, re-run the eight open-world and underground-garage driving sequences end-to-end on hardware, and check whether tracking/segmentation/prediction gains and the ~30 imes latency reduction survive device variation and real light fields.

Watch

Extended reading notes

Core claim

A polarization-sensitive imager co-integrated with a functionally partitioned RRAM array can implement a hardware task-traction mechanism—light-field selection, ROI extraction, and short-horizon target anticipation—that executes visual tasks in 193 μs and, across eight open-world scenarios, improves object tracking, segmentation, and trajectory prediction by 25.54%, 37.73%, and 36.10% while cutting latency by about 30.6 imes relative to state-of-the-art software solutions.

Load-bearing premise

The big accuracy and latency numbers measured mostly on a calibrated simulated RRAM for high-resolution scenes will still hold when a scaled physical array is used in real open-world driving.

Editorial extensions

If this is right

  • Perception pipelines can drop full-frame image enhancement and large end-to-end models when the front end already suppresses task-irrelevant light.
  • Autonomous-driving stacks under glare, low light, reflection, or camouflage can replace multi-sensor fusion with a single polarization-plus-RRAM front end for lower latency.
  • Other light-field modalities (infrared, spectral) can plug into the same three-stage traction pipeline without redesigning the RRAM partitioning.
  • The same front-end distillation principle can be reused for multi-target scenes via the lightweight non-task monitoring path that spawns extra ROIs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the simulated-to-hardware gap is closed, real-time edge robots and drones could run closed-loop vision without GPU-class compute.
  • The two principles the authors name—task-guided acquisition and prediction-guided sensing—suggest analogous front-end distillers for audio (selective source tracking) and touch (slip anticipation).
  • Failure modes under extreme multi-target clutter or RRAM drift would show up first as missed monitoring spikes rather than as tracking error inside the predicted ROI.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript presents a neuromorphic vision system that co-integrates a 12×12 polarization-sensitive photodiode array with an 8×8 BEOL 1T1R HfO2 RRAM array and an FPGA to implement a hardware “task traction mechanism” (feature traction via gradient-entropy light-field selection, attention traction via temporal conductance modulation for ROI extraction, and prediction traction via motion-intensity anticipation). Inspired by dragonfly vision and information-bottleneck ideas, the system is claimed to distill task-relevant cues at the sensor front end, execute visual tasks in 193 μs, and, across eight open-world conditions plus vehicle-mounted driving scenes, improve object tracking, segmentation, and trajectory prediction by 25.54%, 37.73%, and 36.10% while reducing latency ~30.6× versus SOTA software baselines. Hardware characterization (I–V, analogue programming, retention >100 ks, 30 ns switching) and a full-stack specular-interference proof-of-concept are reported; larger-scene and driving evaluations use a VTEAM-calibrated simulated RRAM array.

Significance. If the hardware–simulation bridge holds, this is a meaningful systems contribution: it moves task-adaptive light-field selection, ROI extraction, and short-horizon anticipation into a single functionally partitioned RRAM array at the perceptual front end, rather than treating neuromorphic devices only as post-sensor accelerators. The polarization imager (extinction ~9000:1), BEOL 1T1R integration, ablations of the three traction modules, multi-baseline comparisons (base, enhancement, end-to-end, multi-sensor fusion), and vehicle-mounted demos are concrete strengths. The work is of interest to neuromorphic sensing, in-sensor computing, and robust open-world perception. Credit is due for reporting device-level metrics, module ablations (Fig. S26, STable 11), and explicit labeling of simulated-RRAM results in figure captions—though the abstract still packages mixed evidence as a single system result.

major comments (4)
  1. Abstract and opening Results package three headline numbers—(i) 193 μs execution, (ii) +25.54/+37.73/+36.10% accuracy across eight scenarios, (iii) ~30.6× latency cut vs SOTA—as if they come from one physical system. The manuscript itself separates the evidence: 193 μs and large relative gains under specular interference are measured on the integrated 12×12/8×8 stack (Fig. 2i–l; proof-of-concept), whereas the eight-scenario suite, garage driving results, and SOTA percentages are “based on simulated RRAM array” (Fig. 3 caption; Fig. 4d–f; Methods, VTEAM model). This conflation is load-bearing for the central claim. Please restructure abstract, Results, and Discussion so hardware-measured and simulation-only metrics are never co-listed without explicit scope, and state clearly which claims are demonstrated in silicon versus extrapolated under the calibrated model.
  2. Methods and Results (scalability / SNote 7; Figs. 3–4): high-resolution open-world and autonomous-driving evaluations assume that a VTEAM model calibrated to measured pulse responses, plus ideal partitioning of one array into computation/perception/monitoring regions, faithfully represents scaled array behavior. Array-level non-idealities that matter for continuous high-rate operation—device-to-device variation under concurrent multi-region use, sneak paths, write-disturb, retention under sustained pulsing, and spatial scaling from 8×8 to the tiled m×n units (m=204, n=170)—are only lightly addressed (Fig. S27 covers moderate variability). Without either a larger physical array demo or a quantified error budget showing how these non-idealities propagate into IoU/F1/ED and latency, the claim that simulated gains “faithfully represent what a scaled physical system would deliver” remains the
  3. Latency comparisons vs SOTA (Figs. 3–4, S18–S20; STables 3–10): end-to-end times for enhancement models, YOLOv11, Mask-DiFuser, MoETrack, etc., are measured on an RTX 4080 for full-image pipelines, while the proposed system reports ROI-restricted processing after front-end distillation (and 193 μs only for the small physical stack). The ~30× factor is therefore not an apples-to-apples system comparison unless input resolution, output task definition, and what is included in “end-to-end” (sensing, R/W, FPGA post-processing) are matched and stated. Please define a common evaluation protocol (same frames, same task heads where possible, breakdown of sensing vs compute vs I/O) and report both absolute latencies and accuracy–latency Pareto points so the efficiency claim is interpretable.
  4. Feature traction (Eqs. 1–4; Methods): selection between intensity and polarization rests on local/global gradient entropy with Roberts kernels programmed into RRAM. This is a free design choice (axiom that gradient entropy is a sufficient task-relevance criterion). Ablations remove whole modules but do not test alternative selection metrics (e.g., contrast, DoLP variance, learned scores) or failure cases where high-gradient clutter is task-irrelevant (specular edges, water ripples). Given that feature traction removal costs ~37.9% average accuracy (STable 11), a short controlled study on when gradient entropy mis-selects the light field is needed to support the claim of task-adaptive, not merely high-gradient, selection.
minor comments (6)
  1. Several free thresholds and maps are listed without values or ranges in the main text: ROI binarization/area thresholds, motion-to-displacement f(Q) in Eq. (9), monitoring spike/recovery voltages, and programmed conductance targets for gradient kernels. Put numerical defaults and a one-paragraph sensitivity note in Methods or SI.
  2. Figure 1(c–d) and Figure 5 introduce PTR/ROI/task-region terminology that later overlaps (tM, p_t M, R_t P). A small notation table would reduce confusion.
  3. Proof-of-concept claims “214%, 358%, 62%” accuracy improvements vs full-image intensity on GPU (Results). Report absolute IoU/F1/ED for both sides in the main text, not only relative percentages, so effect sizes are readable when baselines are weak.
  4. SNote cross-references (SNote 1–13, STables 1–13) are central to reproducibility but not available in the main PDF package reviewed here; ensure SI is complete and that every main-text percentage maps to a specific table row.
  5. Typographical consistency: “193 μs” vs “193 {\mu}s”, “1T1R” hyphenation, and mixed “open-world” / “open world”. Standardize units and device nomenclature throughout.
  6. Discussion’s generalization to auditory/tactile perception is interesting but speculative; shorten or move to outlook so it does not dilute the vision-systems contribution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: experimental hardware/systems paper evaluated on external metrics, not a derivation that reduces predictions to its inputs.

full rationale

The manuscript is an experimental neuromorphic-systems paper. Its load-bearing claims are measured performance numbers (193 μs end-to-end execution on the integrated stack; IoU/F1/ED gains and latency ratios versus named baselines across eight scenarios and driving tests). Those quantities are external evaluation metrics, not algebraic consequences of the design equations. Feature traction (gradient-entropy selection via programmed Roberts kernels), attention traction (high-pass temporal modulation of RRAM conductance), and prediction traction (motion-intensity gradients mapped by a design function f to next-frame PTR) are engineered modules whose utility is checked by ablation and by comparison to Farneback, enhancement models, YOLOv11, Mask-DiFuser, MoETrack, and multi-sensor modalities. Free thresholds, the mapping f, and the VTEAM model calibrated on measured pulse data are ordinary design/calibration choices; they do not force the reported accuracy or latency figures by construction, nor do they rename a known identity as a prediction. Biological and information-bottleneck citations are external inspiration, not self-citation uniqueness theorems that forbid alternatives. No step reduces a claimed first-principles result to its own fitted inputs. Circularity score is therefore 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 2 invented entities

The central performance claim rests on device physics and engineering design choices more than on free mathematical axioms. Load-bearing pieces are: that gradient entropy on intensity/polarization is a valid task-relevance proxy; that RRAM conductance dynamics can encode temporal cues and implement fixed gradient kernels; that short-horizon center shifts from average motion intensity predict useful next ROIs; and that simulated RRAM dynamics calibrated to pulse data stand in for scaled hardware. Several numerical thresholds and grid sizes are free design parameters that gate ROI formation and monitoring.

free parameters (5)
  • ROI binarization and area thresholds
    Attention traction converts RRAM states to ROIs by a predefined threshold plus morphological open/close and a minimum connected-area cut; values are design choices that directly control reported task regions.
  • Motion-to-displacement map f(Q)
    Prediction traction maps average horizontal/vertical motion intensity to center shifts via a predefined transformation f; this mapping is not derived from first principles in the text.
  • Spatial unit tiling (m=204, n=170) and monitoring grid (6×6)
    High-resolution inputs are blocked into fixed units for RRAM modulation, and non-task monitoring uses a fixed p×q grid; both sizes are hand-chosen and affect temporal cue granularity.
  • Monitoring spike and recovery voltage thresholds
    Lightweight monitoring suppresses or amplifies non-task motion based on thresholds on inter-frame monitoring vectors and recovery voltages; these set false-alarm vs miss tradeoffs.
  • RRAM programmed conductance targets for gradient kernels
    Roberts operators are written into nonvolatile conductance states with write-verify accuracy (~5 μS); kernel choice and programmed levels are implementation parameters of feature/prediction traction.
assumptions (5)
  • domain assumption Information-bottleneck-style compression that preserves task-relevant cues improves efficiency and robustness in unstructured vision.
    Introduction frames the entire system as guided by information bottleneck theory; this is assumed rather than proven for the specific hardware pipeline.
  • ad hoc to paper Local gradient entropy is a sufficient criterion to select between intensity and polarization channels for the current task region.
    Feature traction always picks the max-entropy light field inside the predicted region; success of later modules depends on this proxy being task-aligned.
  • domain assumption RRAM cells can simultaneously provide fast pulse-driven temporal encoding and stable analogue retention for gradient compute in one array.
    Methods and Discussion attribute dual use to HfO2/Ta stack and 1T1R design; the system architecture assumes this dual regime holds under operating conditions.
  • ad hoc to paper VTEAM dynamics calibrated to measured pulse responses adequately model array behavior for high-resolution scene evaluation.
    Complex open-world and driving results are obtained with simulated RRAM; transfer from calibrated model to physical scaled arrays is assumed.
  • domain assumption Standard vision metrics (IoU, F1, Euclidean distance) and chosen software baselines fairly represent SOTA open-world perception latency/accuracy.
    Claimed percentage gains and 30.6× latency reduction are defined relative to Farneback, enhancement models, YOLOv11, and multi-sensor fusion methods on the authors' eight-scenario benchmark.
invented entities (2)
  • task traction mechanism (feature, attention, prediction traction) independent evidence
    purpose: Names the three-stage hardware information-distillation pipeline that selects light fields, extracts ROIs, and anticipates targets at the sensor front end.
    The label packages known operations into a system-level construct; independent support is the hardware/simulation experiments, not an external physical discovery.
  • functionally partitioned single RRAM array (computation / perception / monitoring regions) independent evidence
    purpose: Allows gradient compute, temporal cue encoding, and unexpected-motion monitoring without separate chips.
    Architectural invention of the paper; evidence is device measurements and system demos, but region sizing and roles are design choices, not independently observed natural objects.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A neuromorphic vision system for open-world visual intelligence." pith.science (2026). https://pith.science/paper/OLXD7S72

@misc{pith2026260710066,
  author       = {Pith},
  title        = {Pith review of: A neuromorphic vision system for open-world visual intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OLXD7S72}},
  note         = {Machine review of arXiv:2607.10066}
}
read the original abstract

Time-efficient and robust visual intelligence remains a critical challenge in unstructured open-world environments, yet current approaches often rely on computationally intensive neural architectures or task-specific sensors with limited versatility. Inspired by biological vision and information bottleneck theory, we report a neuromorphic vision system that performs task-oriented visual intelligence through an information distillation strategy (named as task traction mechanism) implemented on hardware. The system integrates a polarization-sensitive imager with a resistive random-access memory (RRAM) array to progressively distill task-relevant information via light field selection, region of interest extraction, and target anticipation. The neuromorphic vision system conducts visual tasks within an execution time of 193 {\mu}s. Evaluation across eight challenging open-world scenarios shows accuracy improvements of 25.54%, 37.73%, and 36.10% for object tracking, object segmentation, and trajectory prediction, respectively, together with an average 30.6-fold reduction in latency relative to state-of-the-art solutions.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 5 canonical work pages

  1. [1]

    Perceiving real-world scenes

    Biederman, I. Perceiving real-world scenes. Science 177, 77–80 (1972)

  2. [2]

    & Melchner, M

    Sperling, G. & Melchner, M. J. The attention operating characteristic: examples from visual search. Science 202, 315–318 (1978)

  3. [3]

    & Julesz, B

    Sagi, D. & Julesz, B. ‘Where’ and ‘what’ in vision. Science 228, 1217–1219 (1985)

  4. [4]

    & Desimone, R

    Moran, J. & Desimone, R. Selective attention gates visual processing in the extrastriate cortex. Science 229, 782–784 (1985)

  5. [5]

    P., Baccus, S

    Ölveczky, B. P., Baccus, S. A. & Meister, M. Segregation of object and background motion in the retina. Nature 423, 401–408 (2003)

  6. [6]

    Dense reinforcement learning for safety validation of autonomous vehicles[J]

    Feng S, Sun H, Yan X, et al. Dense reinforcement learning for safety validation of autonomous vehicles[J]. Nature, 2023, 615(7953): 620-627

  7. [7]

    Cao, Z. et al. Continuous improvement of self-driving cars using dynamic confidence-aware reinforcement learning. Nat. Mach. Intell. 5, 145–158 (2023)

  8. [8]

    Wang, Y., Yue, Y., Yue, Y. et al. Emulating human-like adaptive vision for efficient and flexible machine visual perception. Nat Mach Intell (2025). https://doi.org/10.1038/s42256-025-01130-7

Show all 64 references
  1. [9]

    Kaufmann, E. et al. Champion-level drone racing using deep reinforcement learning. Nature 620, 982–987 (2023)

  2. [10]

    A matched case-control analysis of autonomous vs human-driven vehicle accidents[J]

    Abdel-Aty M, Ding S. A matched case-control analysis of autonomous vs human-driven vehicle accidents[J]. Nature communications, 2024, 15(1): 4931

  3. [11]

    Yolov11: An overview of the key architectural enhancements[J]

    Khanam R, Hussain M. Yolov11: An overview of the key architectural enhancements[J]. arXiv preprint arXiv:2410.17725, 2024

  4. [12]

    Sam 2: Segment anything in images and videos[C]//International Conference on Learning Representations

    Ravi N, Gabeur V, Hu Y T, et al. Sam 2: Segment anything in images and videos[C]//International Conference on Learning Representations. 2025, 2025: 28085-28128

  5. [13]

    Ithaca365: Dataset and driving perception under repeated and challenging weather conditions[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Diaz-Ruiz C A, Xia Y, You Y, et al. Ithaca365: Dataset and driving perception under repeated and challenging weather conditions[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 21383-21392

  6. [14]

    Learning naturalistic driving environment with statistical realism[J]

    Yan X, Zou Z, Feng S, et al. Learning naturalistic driving environment with statistical realism[J]. Nature communications, 2023, 14(1): 2037

  7. [15]

    OpenAI Gpt-4 Technical Report (OpenAI, 2023)

  8. [16]

    Color shift estimation-and-correction for image enhancement[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Li Y, Xu K, Hancke G P, et al. Color shift estimation-and-correction for image enhancement[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 25389-25398

  9. [17]

    Single image reflection separation via component synergy[C]//Proceedings of the IEEE/CVF international conference on computer vision

    Hu Q, Guo X. Single image reflection separation via component synergy[C]//Proceedings of the IEEE/CVF international conference on computer vision. 2023: 13138-13147

  10. [18]

    Flowie: Efficient image enhancement via rectified flow[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Zhu Y, Zhao W, Li A, et al. Flowie: Efficient image enhancement via rectified flow[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 13-22

  11. [19]

    Hvi: A new color space for low-light image enhancement[C]//Proceedings of the Computer Vision and Pattern Recognition Conference

    Yan Q, Feng Y, Zhang C, et al. Hvi: A new color space for low-light image enhancement[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 5678-5687

  12. [20]

    Zero-reference low-light enhancement via physical quadruple priors[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Wang W, Yang H, Fu J, et al. Zero-reference low-light enhancement via physical quadruple priors[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 26057-26066

  13. [21]

    Robust single image reflection removal against adversarial attacks[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Song Z, Zhang Z, Zhang K, et al. Robust single image reflection removal against adversarial attacks[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 24688-24698

  14. [22]

    Underwater ranker: Learn which is better and how to be better[C]//Proceedings of the AAAI conference on artificial intelligence

    Guo C, Wu R, Jin X, et al. Underwater ranker: Learn which is better and how to be better[C]//Proceedings of the AAAI conference on artificial intelligence. 2023, 37(1): 702-709

  15. [23]

    Toward better than pseudo-reference in underwater image enhancement[J]

    Liu Y, Jiang Q, Li X, et al. Toward better than pseudo-reference in underwater image enhancement[J]. IEEE Transactions on Image Processing, 2025

  16. [24]

    Mask-DiFuser: A masked diffusion model for unified unsupervised image fusion[J]

    Tang L, Li C, Ma J. Mask-DiFuser: A masked diffusion model for unified unsupervised image fusion[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  17. [25]

    Revisiting rgbt tracking benchmarks from the perspective of modality validity: A new benchmark, problem, and solution[J]

    Tang Z, Xu T, Wu X J, et al. Revisiting rgbt tracking benchmarks from the perspective of modality validity: A new benchmark, problem, and solution[J]. IEEE Transactions on Image Processing, 2025

  18. [26]

    Harnessing disordered photonics via multi-task learning towards intelligent four-dimensional light field sensors[J]

    Zhu S, Zheng Z, Meng W, et al. Harnessing disordered photonics via multi-task learning towards intelligent four-dimensional light field sensors[J]. PhotoniX, 2023, 4(1): 26

  19. [27]

    High discrimination infrared polarimetry via intrinsic anisotropic photogating in black phosphorus meta-polarization detectors[J]

    Bu Y, Ye T, Wang R, et al. High discrimination infrared polarimetry via intrinsic anisotropic photogating in black phosphorus meta-polarization detectors[J]. Nature Communications, 2025, 16(1): 10137

  20. [28]

    Polarization-selective unidirectional and bidirectional diffractive neural networks for information security and sharing[J]

    Guo Z, Tan Z, Zang X, et al. Polarization-selective unidirectional and bidirectional diffractive neural networks for information security and sharing[J]. Nature Communications, 2025, 16(1): 4492

  21. [29]

    Matrix Fourier optics enables a compact full-Stokes polarization camera[J]

    Rubin N A, D’Aversa G, Chevalier P, et al. Matrix Fourier optics enables a compact full-Stokes polarization camera[J]. Science, 2019, 365(6448): eaax1839

  22. [30]

    J. Wei, Y. Chen, Y. Li, W. Li, J. Xie, J. Xie, C. Lee, K. S. Novoselov, C.-W. Qiu, Geometric filterless photodetectors for mid-infrared spin light. Nat. Photonics 17, 171–178 (2023)

  23. [31]

    Near-infrared and visible light dual-mode organic photodetectors[J]

    Lan Z, Lei Y, Chan W K E, et al. Near-infrared and visible light dual-mode organic photodetectors[J]. Science advances, 2020, 6(5): eaaw8065

  24. [32]

    Intelligent infrared sensing enabled by tunable moiré quantum geometry[J]

    Ma C, Yuan S, Cheung P, et al. Intelligent infrared sensing enabled by tunable moiré quantum geometry[J]. Nature, 2022, 604(7905): 266-272

  25. [33]

    A broadband hyperspectral image sensor with high spatio-temporal resolution[J]

    Bian L, Wang Z, Zhang Y, et al. A broadband hyperspectral image sensor with high spatio-temporal resolution[J]. Nature, 2024, 635(8037): 73-81

  26. [34]

    Yako M , Yamaoka Y , Kiyohara T ,et al.Video-rate hyperspectral camera based on a CMOS-compatible random array of Fabry–Pérot filters[J].Nature Photonics, 2023, 17(3):8.DOI:10.1038/s41566-022-01141-5

  27. [35]

    Hyperspectral confocal imaging for high-throughput readout and analysis of bio-integrated microlasers[J]

    Titze V M, Caixeiro S, Dinh V S, et al. Hyperspectral confocal imaging for high-throughput readout and analysis of bio-integrated microlasers[J]. Nature Protocols, 2024, 19(3): 928-959

  28. [36]

    The information bottleneck method[J]

    Tishby N, Pereira F C, Bialek W. The information bottleneck method[J]. arXiv preprint physics/0004057, 2000

  29. [37]

    Information theory[M]

    Ash R B. Information theory[M]. Courier Corporation, 2012

  30. [38]

    Polarized vision in the eyes of the most effective predators: dragonflies and damselflies (Odonata)[J]

    Cezário R R, Lopez V M, Datto-Liberato F, et al. Polarized vision in the eyes of the most effective predators: dragonflies and damselflies (Odonata)[J]. The Science of Nature, 2025, 112(1): 8

  31. [39]

    Network adaptation improves temporal representation of naturalistic stimuli in the Drosophila eye

    Nikolaev A., Leung B., Odermatt B., Lagnado L. “Network adaptation improves temporal representation of naturalistic stimuli in the Drosophila eye.” PLoS Biology, 2009

  32. [40]

    Exploration of motion inhibition for the suppression of false positives in small target detection

    Melville-Smith A., et al. “Exploration of motion inhibition for the suppression of false positives in small target detection.” Biological Cybernetics, 2022

  33. [41]

    A lattice filter model of the visual pathway

    Gregor K., Chklovskii D.B. “A lattice filter model of the visual pathway.” NeurIPS, 2012

  34. [42]

    Internal models direct dragonfly interception steering[J]

    Mischiati M, Lin H T, Herold P, et al. Internal models direct dragonfly interception steering[J]. Nature, 2015, 517(7534): 333-338

  35. [43]

    Preattentive facilitation of target trajectories in a dragonfly visual neuron[J]

    Lancer B H, Evans B J E, Fabian J M, et al. Preattentive facilitation of target trajectories in a dragonfly visual neuron[J]. Communications biology, 2022, 5(1): 829

  36. [44]

    A predictive focus of gain modulation encodes target trajectories in insect vision[J]

    Wiederman S D, Fabian J M, Dunbier J R, et al. A predictive focus of gain modulation encodes target trajectories in insect vision[J]. Elife, 2017, 6: e26478

  37. [45]

    Capture success and efficiency of dragonflies pursuing different types of prey[J]

    Combes S A, Salcedo M K, Pandit M M, et al. Capture success and efficiency of dragonflies pursuing different types of prey[J]. Integrative and comparative biology, 2013, 53(5): 787-798

  38. [46]

    BEVHeight++: Toward Robust Visual Centric 3D Object Detection,

    L. Yang et al., "BEVHeight++: Toward Robust Visual Centric 3D Object Detection," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 6, pp. 5094-5111, June 2025, doi: 10.1109/TPAMI.2025.3549711

  39. [47]

    Adaptive Feature-Manipulated Vehicle and Pedestrian Detection in Infrared Images,

    G. Chen, P. Zhang, Y. Zhang, Z. He and B. Shi, "Adaptive Feature-Manipulated Vehicle and Pedestrian Detection in Infrared Images," in IEEE Transactions on Intelligent Transportation Systems, vol. 26, no. 4, pp. 4579-4591, April 2025, doi: 10.1109/TITS.2025.3545844

  40. [48]

    Self-Supervised Learning of LiDAR 3D Point Clouds via 2D-3D Neural Calibration,

    Y. Zhang, J. Hou, S. Ren, J. Wu, Y. Yuan and G. Shi, "Self-Supervised Learning of LiDAR 3D Point Clouds via 2D-3D Neural Calibration," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 10, pp. 9201-9216, Oct. 2025, doi: 10.1109/TPAMI.2025.3584625

  41. [49]

    AiDT: Toward Radar-Based Joint Anti-Interference Detection and Tracking for Weak Extended Targets Under Zero-Trust Autonomous Perception Tasks,

    Z. Zhang, Y. Zhang, D. Huang, X. Fang, M. Zhou and Y. Zhang, "AiDT: Toward Radar-Based Joint Anti-Interference Detection and Tracking for Weak Extended Targets Under Zero-Trust Autonomous Perception Tasks," in IEEE Transactions on Robotics, vol. 41, pp. 3368-3384, 2025, doi: 1...

  42. [50]

    Lattice-allocated real-time line segment feature detection and tracking using only an event-based camera[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision

    Ikura M, Glover A, Mizuno M, et al. Lattice-allocated real-time line segment feature detection and tracking using only an event-based camera[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025: 4645-4654

  43. [51]

    Two-frame motion estimation based on polynomial expansion[C]//Scandinavian conference on Image analysis

    Farnebäck G. Two-frame motion estimation based on polynomial expansion[C]//Scandinavian conference on Image analysis. Berlin, Heidelberg: Springer Berlin Heidelberg, 2003: 363-370

  44. [52]

    Localization of autonomous vehicles in tunnels based on roadside multi-sensor fusion[J]

    Zhu K, Chen S, Shi J, et al. Localization of autonomous vehicles in tunnels based on roadside multi-sensor fusion[J]. IEEE Transactions on Intelligent Vehicles, 2024

  45. [53]

    Autonomous Driving in Underground Mines via Parallel Driving Operation Systems: Challenges, Frameworks and Cases Study[J]

    Tian B, Yang J, Zhang C, et al. Autonomous Driving in Underground Mines via Parallel Driving Operation Systems: Challenges, Frameworks and Cases Study[J]. IEEE Transactions on Intelligent Vehicles, 2024

  46. [54]

    Multipath ghost recognition for indoor MIMO radar[J]

    Feng R, De Greef E, Rykunov M, et al. Multipath ghost recognition for indoor MIMO radar[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 1-10

  47. [55]

    Virtual point removal for large-scale 3d point clouds with multiple glass planes[J]

    Yun J S, Sim J Y. Virtual point removal for large-scale 3d point clouds with multiple glass planes[J]. IEEE transactions on pattern analysis and machine intelligence, 2019, 43(2): 729-744

  48. [56]

    RGBT salient object detection: A large-scale dataset and benchmark[J]

    Tu Z, Ma Y, Li Z, et al. RGBT salient object detection: A large-scale dataset and benchmark[J]. IEEE Transactions on Multimedia, 2022, 25: 4163-4176

  49. [57]

    Low-latency automotive vision with event cameras[J]

    Gehrig D, Scaramuzza D. Low-latency automotive vision with event cameras[J]. Nature, 2024, 629(8014): 1034-1040

  50. [58]

    Precise and scalable analogue matrix equation solving using resistive random-access memory chips[J]

    Zuo P, Wang Q, Luo Y, et al. Precise and scalable analogue matrix equation solving using resistive random-access memory chips[J]. Nature Electronics, 2025: 1-12

  51. [59]

    In situ spectral reconstruction based on a memristor chip for energy-efficient computational spectrometry[J]

    Zhao H, Wang L, Zhou Y, et al. In situ spectral reconstruction based on a memristor chip for energy-efficient computational spectrometry[J]. Nature Electronics, 2026: 1-13

  52. [60]

    A full-stack memristor-based computation-in-memory system with software-hardware co-development[J]

    Yu R, Wang Z, Liu Q, et al. A full-stack memristor-based computation-in-memory system with software-hardware co-development[J]. Nature Communications, 2025, 16(1): 2123

  53. [61]

    Fully hardware-implemented memristor convolutional neural network[J]

    Yao P, Wu H, Gao B, et al. Fully hardware-implemented memristor convolutional neural network[J]. Nature, 2020, 577(7792): 641-646

  54. [62]

    Full hardware implementation of neuromorphic visual system based on multimodal optoelectronic resistive memory arrays for versatile image processing[J]

    Zhou G, Li J, Song Q, et al. Full hardware implementation of neuromorphic visual system based on multimodal optoelectronic resistive memory arrays for versatile image processing[J]. Nature communications, 2023, 14(1): 8489

  55. [63]

    Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip[J]

    Yao M, Richter O, Zhao G, et al. Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip[J]. Nature communications, 2024, 15(1): 4464

  56. [64]

    VTEAM: A general model for voltage-controlled memristors[J]

    Kvatinsky S, Ramadan M, Friedman E G, et al. VTEAM: A general model for voltage-controlled memristors[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2015, 62(8): 786-790

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.