Pith. sign in

REVIEW 7 cited by

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03118 v2 pith:3E7OP54E submitted 2024-12-04 cs.HC cs.CV

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

classification cs.HC cs.CV
keywords objectobjectfinderassistiveblindsearchtargetinteractiveopen-vocabulary
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos

    cs.CV 2026-05 unverdicted novelty 7.0

    EgoExoMem is the first benchmark for cross-view memory reasoning on synchronized egocentric-exocentric videos, where E2-Select raises MLLM accuracy from 55.3% to 58.2% over baselines.

  2. GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance

    cs.CV 2025-03 unverdicted novelty 7.0

    GuideDog supplies 22K egocentric image-description pairs from 46 countries and an 818-sample QA benchmark showing that current multimodal models still struggle with depth perception and BLV-specific guidance rules.

  3. What if? Emulative Simulation with World Models for Situated Reasoning

    cs.CV 2026-03 conditional novelty 6.5

    WanderDream supplies 15.8K panoramic mental-exploration videos and 158K QA pairs showing that world-model imagination measurably improves situated spatial reasoning without active exploration.

  4. EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos

    cs.CV 2026-05 unverdicted novelty 6.0

    EgoExoMem introduces the first cross-view ego–exo video memory benchmark (2.6K MCQs, eight QA types) and E²-Select, a training-free dual-view frame selector scoring 58.2%.

  5. DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection

    cs.CV 2026-05 unverdicted novelty 5.0

    DSAA improves fine-grained open-vocabulary object detection by injecting attribute priors via APA in text embeddings, modulating K/V vectors in BERT, and using an attribute-aware contrastive loss, with gains shown on ...

  6. VisionAssist: An Open-Source Smartphone Assistant for AI-Based Visual Accessibility

    cs.CV 2026-07 conditional novelty 4.0

    An open-source voice-driven app wraps cloud Gemini AI with on-device contacts and calendar to help low-vision users find, read, and manage everyday tasks.

  7. DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection

    cs.CV 2026-05 unverdicted novelty 4.0

    DSAA boosts fine-grained OVD by injecting attribute priors via APA in text embeddings, modulating K/V in BERT, and using attribute-aware contrastive loss, with gains reported on FG-OVD benchmark.