REVIEW 6 major objections 6 minor 2 cited by
AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models
T0 review · 6 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A hat-mounted Raspberry Pi camera with a vision-language model gives visually impaired users real-time object naming, personalized face recognition, and spoken scene descriptions; the authors report 90% accuracy in good light and an 85…
desk verdict A plausible low-cost assistive wearable prototype undermined by a thin, internally inconsistent evaluation and a reference list that looks fabricated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrier of the argument is a three-mode processing pipeline built around the Raspberry Pi 4. In the default mode the system continuously runs MobileNet SSD object detection and FaceNet face recognition, with the ultrasonic sensor triggering a buzzer when an object is closer than 20 cm. Pressing the green button records a voice-labelled face or object into the SQLite embedding database; pressing the blue button sends the captured image and a prompt containing detected labels, spatial relationships, and environmental cues to GPT-4o-mini through a cloud API, then speaks the returned description through bone-conduction headphones. TensorFlow Lite and model quantization are what make this pipeline fast enough for wearable use on the Pi's limited CPU/GPU.
What would settle it
Run the same pipeline on a fixed, publicly documented set of several hundred labelled images in good and low lighting, using the paper's stated 50% confidence threshold, and compare the resulting confusion matrix against the reported 90%/80% accuracy, 0.88/0.92 precision/recall, and 0.76/0.74 values. If the measured values fall materially below those numbers, or if the paper's own confusion matrix in Figure 5 does not satisfy the standard formulas for the reported precision, recall, F1, and accuracy, the central accuracy claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that a complete wearable assistance loop can be built from low-cost, off-the-shelf parts: MobileNet SSD detects objects in real time, FaceNet produces embeddings matched by cosine similarity to a personalised SQLite database, and a GPT-4o-mini vision-language model converts the current frame into a contextual spoken description when the user presses the blue button. The green button gives users a one-click way to add new people or objects, with GAN-generated synthetic data and real-time fine-tuning claimed to improve recognition over time. The authors report that this combination yields 90% accuracy in good light and 80% in low light, precision/recall/F1 of 0.88/0.92/0.90 and 0.76/0.74/0.75 respectively, a 20 cm ultrasonic collision warning, and a comparative advantage over commercial and research systems that lack either contextual understanding or personalisation.
Load-bearing premise
The reported accuracy and usability figures rest on the assumption that the images and scenarios used in Section 5 are representative of real daily use and are correctly labelled; the paper does not state the test set size, its contents, or the labelling protocol, so if that assumption fails the 90% accuracy and SUS 85 score will not transfer to actual use.
Editorial extensions
If this is right
- At the reported 1.5-second recognition time, users can receive near-immediate audio naming of objects while walking, with collision warnings for anything closer than 20 cm.
- A blue-button description lets a user request a full spoken account of a scene, which the paper demonstrates in market and grocery-store settings, making unassisted shopping a plausible use case.
- Because the green button stores personal face embeddings with voice labels, the system becomes more accurate for the specific people and objects in a user's life, not just generic classes.
- The reported 90% good-light and 80% low-light accuracy place the system in the same performance range as commercial assistive glasses while relying on much cheaper hardware.
- The low-light dip points to a clear improvement path: adding infrared or thermal sensing would directly target the system's weakest operating condition.
Reading between the lines
- The architecture implies a privacy trade-off the paper states only in passing: images are sent to the cloud only when the blue button is pressed, so offline or privacy-sensitive settings lose contextual descriptions entirely; a local open-weight vision-language model would restore that capability at the cost of slower responses.
- The one-click enrolment procedure could be generalised from faces to arbitrary object categories, allowing users to build personal inventories of medication bottles, food packages, or tools, something the evaluation does not directly test.
- A direct ablation study separating the contribution of the LVLM from the local object detector would clarify how much of the reported contextual understanding comes from the cloud model versus the on-device recognition.
- The two-button interaction model suggests a natural extension to voice-triggered capture, which would remove the need for a sighted person to locate the hat buttons.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a wearable assistive system for visually impaired users: a hat-mounted camera connected to a Raspberry Pi 4, using MobileNet SSD for object detection, FaceNet for face recognition, an ultrasonic distance sensor and buzzer for collision warnings, and an API call to GPT-4o-mini for contextual scene descriptions. The authors report recognition metrics under good and low lighting, a 50-person usability study, SUS and satisfaction scores, response times, and a comparative feature table against other assistive devices. The central claim is that the system is a significant advancement over traditional assistive technologies because of its combination of real-time object recognition, personalized face/object databases, and LVLM-based contextual understanding.
Significance. If substantiated, the proposed system would be a useful, low-cost, customizable assistive device that combines obstacle avoidance with contextual scene description and personalization, potentially improving independence for blind and low-vision users. The engineering design is sensible and uses openly available components. However, the paper's significance rests almost entirely on Section 5's evaluation, and that evaluation is not reported in sufficient detail to support the abstract's claim of a 'significant advancement.' The lack of a described test set, the internal inconsistency between the text and Table 2, and the absence of any quantitative or controlled comparison with existing aids mean the claimed benefits are not established by the evidence presented. The paper does not provide reproducible code, data, or analysis scripts, so the headline numbers cannot be independently checked.
major comments (6)
- [§5.1, Table 2, Figure 5] The object-recognition evaluation is not reproducible: the paper never states how many test images or scenes were used, how ground-truth labels were produced, how lighting was controlled, or which object classes were included. The confusion matrix is presented only as a figure with no tabulated counts, so the reported Accuracy 0.90/0.80, Precision 0.88/0.76, Recall 0.92/0.74, and F1 0.90/0.75 cannot be verified. This is load-bearing because these numbers are the primary quantitative evidence for the paper's central claim.
- [§5.1, paragraph 1] The text misstates the error decomposition: it says that with 90% accuracy, '10% were false positives,' but accuracy complement is the total error rate, not the false-positive rate. Given the reported Precision of 0.88 and Recall of 0.92, both false positives and false negatives are present, and their relative magnitudes are unknown without the confusion matrix. The sentence should either present TP/FP/FN/TN counts or be removed.
- [§5.1, low-light sentence] The text says that in lower light 'both precision and recall were at 80%,' but Table 2 reports Precision 0.76 and Recall 0.74. This is a direct internal inconsistency. One of the two statements is wrong, and the discrepancy matters because the low-light robustness is one of the system's claimed advantages over alternatives.
- [§5.3, §4.2] The usability evaluation is reported only as summary scores. The SUS average of 85 is given without item-level SUS scores, individual participant results, task completion rates, or any statistical measure of variability. The scenario ratings in Table 3 (e.g., 'Navigation 4.5/5') are not SUS items, and Table 4 contains internally questionable values (e.g., the '9' for Real-Time Object Detection usage frequency appears to be a typo for 90). Without per-task outcomes and a description of how the 50 participants were recruited and how tasks were standardized, the claim of 'exceptional user acceptance' is not supported.
- [§5.5, §6.1] The comparative analysis against traditional support techniques and other devices is not a controlled evaluation. Section 5.5 reports participants' 'impressions' and Table 5 is a feature checklist, not a measurement of task performance, satisfaction, or safety relative to white canes, guide dogs, or existing electronic travel aids. The abstract's statement that 'comparative analysis shows ... a significant advancement' is therefore not supported by the evidence in Section 5.
- [§3.8, §6.3] The green-button personalization workflow is described as involving GAN-based synthetic data generation and real-time fine-tuning of the recognition model, but no implementation details, dataset sizes, training procedure, or evaluation of the personalized database are provided. Since the personalized database is one of the two advertised improvements over prior systems, its effectiveness needs at least a basic evaluation (e.g., accuracy on newly added faces/objects over time).
minor comments (6)
- [§2.6, §6.2] The paper claims 'for the first time in research' contextual understanding via LLMs in the conclusion, but the literature review does not support such a strong novelty claim; this phrasing should be softened or justified with a targeted survey.
- [Table 4] The usage-frequency value of 9 for Real-Time Object Detection appears to be a typo; it is probably 90, which would be consistent with the other rows.
- [References [54] and [56]] References [54] and [56] both list arXiv:2405.07606, which appears to be a placeholder or an error; the references should be checked and corrected.
- [§4.2] The usability study does not report whether ethical approval or informed-consent procedures were followed for the 50 visually impaired participants; this information should be included in the evaluation section.
- [Throughout] There are numerous typographical and formatting issues, including the repeated word 'participants participants,' 'system system,' inconsistent capitalization, and equations that are not numbered consistently; a careful copyedit is needed before resubmission.
- [§3.4] The phrase 'OpenAI's GPT-4o-mini via Azure' should clarify the exact API/model version and the network latency assumptions, since response-time claims in §5.4 depend on this.
Circularity Check
No significant circularity: the system reports empirical measurements from pre-trained models and an external API, with no parameter fitted to the evaluation outcome.
full rationale
This paper contains no formal derivation chain whose outputs are equivalent to its inputs. The reported performance numbers (Section 5.1, Table 2: Precision 0.88/0.76, Recall 0.92/0.74, F1 0.90/0.75, Accuracy 0.90/0.80) are empirical measurements of a pipeline built from MobileNet SSD, FaceNet, and OpenAI GPT-4o-mini via Azure, all external pre-trained components. No model parameter is fitted to the evaluation set, and the claimed accuracy is not derived from the same data used to define or train the system. The usability scores (SUS 85, satisfaction ratings in Tables 3 and 4) are self-reported survey results, which may be methodologically weak but are not circular in the sense of a prediction reducing to a fitted input. The green-button personalization workflow (Section 3.8) is described as fine-tuning with newly labeled data, but the paper reports no accuracy claim specifically for that workflow, so there is no fitted-input-called-prediction step to flag. The Section 5.1 statement that '10% were false positives' is statistically imprecise and inconsistent with the reported recall/precision decomposition, but that is a reporting error, not a circularity. No load-bearing self-citation is present: the references to prior systems in the comparative Table 5 are external and are not used as unverified premises to justify the central claim. The paper's central claim of being 'a significant advancement' is an interpretation of its own measurements rather than a consequence derived from its own definitions. Accordingly, no circular step meeting the required evidence standard can be exhibited, and the appropriate finding is no significant circularity (score 0).
Assumptions & free parameters
free parameters (2)
- object detection confidence threshold =
0.50
- collision warning distance threshold =
20 cm
assumptions (4)
- domain assumption Quantized MobileNet SSD and FaceNet models preserve accuracy as in their original publications when run on Raspberry Pi via TensorFlow Lite.
- domain assumption The 50 participants' self-reported usability scores and qualitative feedback are truthful and the simulated scenarios reflect real-life conditions.
- domain assumption The GPT-4o-mini API returns accurate, relevant descriptions over the network with latency acceptable for the user.
- domain assumption TensorFlow Lite integer quantization does not materially degrade model accuracy.
Cite this review
Pith. "Pith review of AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models." pith.science (2026). https://pith.science/paper/MMCGGCAQ
@misc{pith2026241220059,
author = {Pith},
title = {Pith review of: AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/MMCGGCAQ}},
note = {Machine review of arXiv:2412.20059}
}
read the original abstract
Visual impairment affects the ability of people to live a life like normal people. Such people face challenges in performing activities of daily living, such as reading, writing, traveling and participating in social gatherings. Many traditional approaches are available to help visually impaired people; however, these are limited in obtaining contextually rich environmental information necessary for independent living. In order to overcome this limitation, this paper introduces a novel wearable vision assistance system that has a hat-mounted camera connected to a Raspberry Pi 4 Model B (8GB RAM) with artificial intelligence (AI) technology to deliver real-time feedback to a user through a sound beep mechanism. The key features of this system include a user-friendly procedure for the recognition of new people or objects through a one-click process that allows users to add data on new individuals and objects for later detection, enhancing the accuracy of the recognition over time. The system provides detailed descriptions of objects in the user's environment using a large vision language model (LVLM). In addition, it incorporates a distance sensor that activates a beeping sound using a buzzer as soon as the user is about to collide with an object, helping to ensure safety while navigating their environment. A comprehensive evaluation is carried out to evaluate the proposed AI-based solution against traditional support techniques. Comparative analysis shows that the proposed solution with its innovative combination of hardware and AI (including LVLMs with IoT), is a significant advancement in assistive technology that aims to solve the major issues faced by the community of visually impaired people
Forward citations
Cited by 2 Pith papers
-
Scalable phonon-laser arrays with self-organized synchronization
Local driving of an Ising-like spin–mechanical chain yields scalable, site-addressable phonon lasers with resonance conditions, on-demand lasing, and self-organized synchronization.
-
Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices
Physical scene text can inject prompts into wearable VLMs, hijacking decisions and content with high success rates across six threat scenarios, partially mitigated by OCR masking and token-drift defenses.
Reference graph
Works this paper leans on
- [56]
-
[1]
S. R. Flaxman, R. R. Bourne, S. Resnikoff, P. Ackland, T. Braithwaite, M. V. Cicinelli, A. Das, J. B. Jonas, J. Keeffe, J. H. Kempen, et al., Global causes of blindness and distance vision impairment 1990–2020: a systematic review and meta-analysis, The Lancet Global Health 5 (12) (2017) e1221–e1234
work page 2017
-
[2]
World Health Organization, World Report on Vision, World Health Organization, 2021
work page 2021
-
[3]
H. R. Marston, S. T. Smith, Mobile phone use by people with impairments, in: Accessibility Conference 2008, University of York, 2008, pp. 1–7
work page 2008
-
[4]
J. M. Gold, I. L. Bailey, A. B. Sekuler, R. Sekuler, Visual span and reading speed in macular degeneration, Vision Research 63 (2012) 40–47
work page 2012
-
[5]
G. Reyes-Cruz, J. E. Fischer, S. Reeves, Reframing disability as competency: Unpacking everyday technology practices of people with visual impairments, in: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 2020, pp. 1–13
work page 2020
-
[6]
D. Dakopoulos, N. G. Bourbakis, Wearable obstacle avoidance electronic travel aids for blind: a survey, IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 40 (1) (2010) 25–35
work page 2010
-
[7]
S. Momayez, J. -M. Robert, M. Rioux, J. De Guise, Design and evaluation of a visuo -haptic augmented reality system for education and training in health sciences, Virtual Reality 17 (1) (2013) 15–29. 17
work page 2013
Show all 56 references
-
[8]
D. G. DEMIRAL, Emerging assistive technologies and challenges encountered, Current Studies in Technology, Innovation and Entrepreneurship (2023) 1
2023
-
[9]
Ahmed, A
N. Ahmed, A. A. Qasem, Challenges and opportunities in developing assistive technology solutions: A literature review, International Journal of Computer Applications 134 (3) (2016) 9–14
2016
-
[10]
M. A. Hersh, Evaluating assistive technology systems for blind people: Methodologies and applications, Neurore- habilitation 27 (3) (2010) 201–217
2010
-
[11]
Kuriakose, R
B. Kuriakose, R. Shrestha, F. E. Sandnes, Tools and technologies for blind and visually impaired navigation support: a review, IETE Technical Review 39 (1) (2022) 3–18
2022
-
[12]
M. H. Abidi, A. N. Siddiquee, H. Alkhalefah, V. Srivastava, A comprehensive review of navigation systems for visually impaired individuals, Heliyon (2024)
2024
-
[13]
M. A. Williams, A. M. Anto´n, A. Nau, Y. Khalifa, J. D. Dorn, M. A. Fahim, Assistive technology for visually impaired and blind people: A narrative review, Journal of Ophthalmology 2019 (2019) 1–15
2019
-
[14]
Bauer, L.-J
I. Bauer, L.-J. Elsaesser, The challenges and opportunities of assistive technology procurement in ghana—a narra- tive literature review, Global Health Action 12 (1) (2019) 1699343
2019
-
[15]
Gogatishvili, A
G. Gogatishvili, A. Malianov, A. Belyaev, A. Karpov, A. Ronzhin, Assistive technologies for people with visual impairments: A literature review, Symmetry 12 (10) (2020) 1698
2020
-
[16]
X. Li, J. Zhou, C. Li, C. Ming, Eye gaze tracking system for people with visual impairment, in: 2018 IEEE International Conference on Consumer Electronics-Taiwan (ICCE-TW), IEEE, 2018, pp. 1–2
2018
-
[17]
Kedia, R
P. Kedia, R. Vishwakarma, An assistive device for the visually impaired that uses deep learning in a raspberry pi embedded system, International Journal of Advanced Computer Science and Applications 9 (8) (2018) 424–428
2018
-
[18]
Prabhushankar, C
J. Prabhushankar, C. Chandru, P. Kannan, S. Udayakumar, Assistive technology device with color identification and text recognition for visually impaired using cnn, International Journal of Recent Technology and Engineering 8 (2S3) (2019) 449–453
2019
-
[19]
U. R. Roentgen, G. J. Gelderblom, M. Soede, L. P. De Witte, Overview of the use of electronic mobility aids by persons who are visually impaired, Journal of Visual Impairment & Blindness 102 (11) (2008) 702–724
2008
-
[20]
I. Khan, S. Khusro, I. Ullah, Technology -assisted white cane: evaluation and future directions, PeerJ 6 (2018) e6058
2018
-
[21]
Attia, D
I. Attia, D. Asamoah, The white cane, its e ffectiveness, challenges and suggestions for e ffective use: A case of akropong school for the blind, Journal of Education, Society and Behavioural Science 33 (3) (2020) 47–55
2020
-
[22]
A. D. P. dos Santos, F. O. Medola, M. J. Cinelli, A. R. Garcia Ramirez, F. E. Sandnes, Are electronic white canes better than traditional canes? a comparative study with blind and blindfolded participants, Universal Access in the Information Society 20 (1) (2021) 93–103
2021
-
[23]
I. Khan, I. U. Shah Khusro, S. Khusro, Technology-assisted white cane: Evaluation and future
-
[24]
Monteiro, J
J. Monteiro, J. P. Aires, R. Granada, R. C. Barros, F. Meneguzzi, Virtual guide dog: An application to support visually-impaired people through deep convolutional neural networks, in: 2017 International Joint Conference on Neural Networks (IJCNN), IEEE, 2017, pp. 2267–2274
2017
-
[25]
C. R. Sanders, The impact of guide dogs on the identity of people with visual impairments, Anthrozoo¨s 13 (3) (2000) 131–139
2000
-
[26]
R. H. Sanders, Guiding the blind with robot dogs, Communications of the ACM 57 (6) (2014) 24–26
2014
-
[27]
Y. Wei, M. Lee, A guide-dog robot system research for the visually impaired, in: 2014 IEEE International Confer- ence on Industrial Technology (ICIT), IEEE, 2014, pp. 800–805
2014
-
[28]
D. Mai, T. Howell, P. Benton, P. C. Bennett, Raising an assistance dog puppy—stakeholder perspectives on what helps and what hinders, Animals 10 (1) (2020) 128
2020
-
[29]
Chuang, N.-C
T.-K. Chuang, N.-C. Lin, J.-S. Chen, C.-H. Hung, Y.-W. Huang, C. Teng, H. Huang, L. -F. Yu, L. Giarre´, H.-C. Wang, Deep trail-following robotic guide dog in pedestrian environments for people who are blind and visually impaired-learning from virtual and real worlds, in: 2018 ...
2018
-
[30]
B. Hong, Z. Lin, X. Chen, J. Hou, S. Lv, Z. Gao, Development and application of key technologies for guide dog robot: A systematic literature review, Robotics and Autonomous Systems 154 (2022) 104104
2022
-
[31]
B. L. Due, Guide dog versus robot dog: Assembling visually impaired people with non-human agents and achieving assisted mobility through distributed co-constructed perception, Mobilities 18 (1) (2023) 148–166
2023
-
[32]
S. Cai, A. Ram, Z. Gou, M. A. W. Shaikh, Y.-A. Chen, Y. Wan, K. Hara, S. Zhao, D. Hsu, Navigating real- world challenges: A quadruped robot guiding system for visually impaired people in diverse environments, in: Proceedings of the CHI Conference on Human Factors in Computing ...
2024
-
[33]
Thiyagarajan, S
K. Thiyagarajan, S. Kodagoda, M. Luu, T. Duggan -Harper, D. Ritchie, K. Prentice, J. Martin, Intelligent guide robots for people who are blind or have low vision: A review, Vision Rehabilitation International 13 (1) (2022) 1–15
2022
-
[34]
Maidenbaum, S
S. Maidenbaum, S. Levy-Tzedek, D.-R. Chebat, A. Amedi, The “eyecane”, a new electronic travel aid for the blind: Technology, behavior & swift learning, Restorative Neurology and Neuroscience 32 (6) (2014) 813–824. 18
2014
-
[35]
OrCam Technologies Ltd, Orcam myeye, https://www.orcam.com/en/myeye/, accessed: 2021-06-01 (2021)
2021
-
[36]
L. Boya, Y. Xu, Design and development of assistive smart glasses for visually impaired people, in: 2020 Interna- tional Conference on Computer Engineering and Application (ICCEA), IEEE, 2020, pp. 336–341
2020
-
[37]
N. G. Bourbakis, Assistive technology for visually impaired people, Springer, 2008, pp. 1141–1156
2008
-
[38]
R. A. Aghaee, H. Al Najada, C. Paul, W. Burleson, J. Gilbert, Developing mobile applications for people with disabilities, in: Proceedings of the 18th Annual Conference on Information Technology Education, ACM, 2017, pp. 97–102
2017
-
[39]
Walle, C
H. Walle, C. De Runz, B. Serres, G. Venturini, A survey on recent advances in ai and vision-based methods for helping and guiding visually impaired people, Applied Sciences 12 (5) (2022). doi:10.3390/app12052308. URL https://www.mdpi.com/2076-3417/12/5/2308
2022 doi
-
[40]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778
2016
-
[41]
Redmon, S
J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: Unified, real-time object detection, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 779–788
2016
-
[42]
B. B. Morrison, R. E. Ladner, S. Doyle, Accessible computing–weaving accessibility throughout the undergraduate cs curriculum, ACM Inroads 8 (4) (2017) 45–48
2017
-
[43]
S. Han, J. Pool, J. Tran, W. Dally, Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, arXiv preprint arXiv:1510.00149 (2015)
2015 arXiv
-
[44]
J. Zhu, Y. Chen, M. Zhang, Q. Chen, Y. Guo, H. Min, Z. Chen, An edge computing platform of guide-dog robot for visually impaired, in: 2019 IEEE 14th International Symposium on Autonomous Decentralized System (ISADS), IEEE, 2019, pp. 1–7
2019
-
[45]
Ghafoor, N
A. Ghafoor, N. Nahar, Assistive technology for university students with visual impairment: a pilot study, Interna- tional Journal of Computer Applications 178 (48) (2019) 1–6
2019
-
[46]
Parsaee, F
M. Parsaee, F. Momeni, G. Jafari, The use of bone conduction headphone to improve auditory perception for visually impaired people, Journal of Biomedical Physics & Engineering 5 (4) (2015) 165
2015
-
[47]
Bradski, The opencv library, Dr
G. Bradski, The opencv library, Dr. Dobb’s Journal of Software Tools 25 (11) (2000) 120–123
2000
-
[48]
Sandler, A
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.-C. Chen, Mobilenetv2: Inverted residuals and linear bottle- necks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4510–4520
2018
-
[49]
Schroff, D
F. Schroff, D. Kalenichenko, J. Philbin, Facenet: A unified embedding for face recognition and clustering, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 815–823
2015
-
[50]
TensorFlow, Tensorflow lite, https://www.tensorflow.org/lite, accessed: 2021-06-01 (2017)
2017
-
[51]
Jacob, S
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, D. Kalenichenko, Quantization and training of neural networks for e fficient integer-arithmetic-only inference, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 2704–2713
2018
-
[52]
Taylor, A
P. Taylor, A. W. Black, R. Caley, The festival speech synthesis system: System documentation, Tech. rep., Human Communication Research Centre, University of Edinburgh (1997)
1997
-
[53]
Viola, M
P. Viola, M. Jones, Rapid object detection using a boosted cascade of simple features, in: Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2001), Vol. 1, IEEE, 2001, pp. I–I
2001
-
[55]
M. Lee, S. Kim, Magiceye: Cnn-based vision assistance for object recognition, arXiv preprint arXiv:2303.13863 (2023). URL https://arxiv.org/abs/2303.13863
2023 arXiv
-
[57]
K. S. Bobba, K. Kartheeban, V. K. S. Boddu, V. M. S. Bolla, D. Bugga, Newvision: Application for helping blind people using deep learning, arXiv preprint arXiv:2311.03395 (2023). URL https://arxiv.org/abs/2311.03395
2023 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.