{"id":"ba3b53c6-caa7-4068-a390-46a6854b9ec7","arxiv_id":"2607.03213","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A sensing-computing split (ESP32 glasses + nearby consumer GPU) delivers ~993 ms median local MLLM visual assistance with safety abstention and open artifacts.","lead":"OpenGlass is an open-source glasses-plus-laptop system that answers spoken visual questions for blind users with local AI, without uploading camera frames to the cloud. It reports roughly one-second median spoken answers over real Wi-Fi and releases hardware, code, prompts, and logs.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Latency claim is solid; the soft spot is whether the 120-frame author rubric can support 'practical assistive quality' without independent scoring or BLV users.","rationale":"I agree with the Reader: the central engineering claim is the real Wi-Fi user\to audio latency under the sensing-computing split, and that claim is supported by Table 4, failure attribution, resolution ablations, and a full open release. The deployment assumption (nearby consumer GPU host) is explicit and the non-navigation disclaimer is clear, so this is not a hidden overclaim about certified mobility. The weakest link for accepting the paper as more than a latency/privacy systems demo is the assistive quality/safety evaluation: small custom set, author rubric, no IAA, no BLV users, and high HCE on sign/QR tasks. That is exactly the Reader's weakest_assumption. It justifies CONDITIONAL rather than ACCEPT or REJECT: accept the measured sub-2 s local pipeline result; condition broader assistive-quality language on independent scoring and human evaluation. No stronger internal inconsistency in the latency argument presents itself, so the verdict should remain CONDITIONAL / UNCHANGED.","tokens_in":13671,"tokens_out":674,"duration_ms":6284,"concrete_test":"Have 2–3 independent raters (ideally including at least one BLV-experienced annotator) re-score the released 120 responses under the paper's own 0/1/2 rubric and HCE definition, blinded to model identity; report IAA (e.g., Cohen/Fleiss κ) and recompute Table 2 Quality/Success/HCE. If κ < 0.6 or HCE on T3A/T3B shifts by >10 absolute points, the quality/safety support for the assistive claim weakens while the latency claim can stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest quantitative claim (Table 4: 993 ms median user\to audio, 97.5% <2 s under real ESP32 Wi-Fi with resize) is well-instrumented and open-sourced. The load-bearing soft spot for the broader 'practical BLV visual assistance' framing is the quality/safety evidence that is packaged with that latency result. Quality, Success, Abstain, and HCE (Tables 1–2) rest on a 120-instance author-collected ESP32 set scored with an author-defined 0/1/2 safety-first rubric, with no inter-annotator agreement, no independent raters, and no BLV participant study (Limitations; §3.3, §3.10). Task-level numbers already show the risk: T3A has 37.5% HCE and T3B 66.7% HCE (Table 2), so 'usable' rates depend heavily on counting safe abstention as success. If those rubric labels are optimistic or non-reproducible, the system can still be a strong latency/privacy reference platform while the assistive-quality half of the claim is overstated. The reader's weakest_assumption correctly flags this; it is not a flaw in the latency measurement itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"OpenGlass presents an open-source, local-first sensing–computing split for multimodal visual assistance aimed primarily at blind and low-vision users: an ESP32-S3/OV5640 glasses unit captures on-demand JPEG frames over Wi-Fi, while a nearby consumer host runs local MiniCPM-V/o 4.5 (llama.cpp, INT4), modular Whisper ASR, and local TTS with sentence-level streaming. The paper’s core quantitative claim is end-to-end user→audio latency under real ESP32 capture—993 ms median (97.5% <2 s) with resized 896×504 payloads and 1625 ms (93.3% <2 s) with raw 1280×720—plus lower and more predictable latency than Gemini 2.5 Flash and Qwen-VL-Max on the same capture path (Table 4). Supporting material includes replay vs. Wi-Fi separation (Tables 1–4), resize/streaming/TTS ablations (Table 3), resolution–slice trade-offs (Table 5), barge-in tests (Table 6), a 120-instance task suite (T1/T2/T3/H1) with a safety-first 0/1/2 rubric, and full release of code, prompts, logs, and evaluation data. The authors explicitly frame the system as a user-initiated reference platform, not a certified navigation aid.","tokens_in":14099,"tokens_out":1421,"duration_ms":15500,"significance":"If the latency and systems results hold—as the instrumentation, failure attribution, and open artifacts strongly suggest—the paper is a useful, reproducible reference for local-first egocentric MLLM assistance under realistic wireless capture. Strengths that should be credited include: (i) clean separation of backend replay latency from real Wi-Fi E2E (Tables 1–4); (ii) concrete identification of resize/image slices as the dominant lever (Tables 3, 5); (iii) failure attribution showing local TTS is not the bottleneck; (iv) safety-oriented abstention design and auditable logs; and (v) full open-source release of firmware, prompts, evaluation set, and logs. The contribution is primarily systems engineering and measurement rather than a new model architecture, but that is appropriate for the stated goal of practical local deployment.","major_comments":[{"comment":"The strongest claim (Table 4: 993 ms median user→audio, 97.5% <2 s under real ESP32 Wi-Fi with resize) is well supported by stage definitions, ablations, and open logs. The load-bearing soft spot is packaging that latency result with “practical assistive quality / safety-aware behavior.” Quality, Success, Abstain, and HCE in Tables 1–2 rest on a 120-instance author-collected ESP32 set scored with an author-defined 0/1/2 rubric, with no inter-annotator agreement, independent raters, or BLV participant study (Limitations; §3.3, §3.10). Task-level numbers already show risk: T3A HCE 37.5%, T3B HCE 66.7% (Table 2), so usable rates depend heavily on counting safe abstention as Success. Please either (a) add independent scoring / IAA and clarify that Success includes abstention, or (b) narrow abstract/intro claims so quality is presented as exploratory rubric evidence for a reference platform,","section":null},{"comment":"§3.1 and Table 4 use a Lenovo Legion laptop with RTX 5060 as the “nearby consumer-grade device.” The paper acknowledges this is an edge-host reference rather than glasses-only deployment (Limitations), but the abstract and introduction still frame everyday BLV assistance. A short, explicit deployment envelope—required host class, power, and that phone/NPU targets are future work—should appear near the main latency claim so readers do not over-generalize the 993 ms result to fully wearable or always-on mobile hosts.","section":null},{"comment":"Table 2 and §3.5 identify sign/QR (T3A/T3B) as the main quality bottleneck with high HCE. The manuscript motivates tool-based QR decoding and stricter abstention, but does not implement or evaluate them. Because fabricated sign/QR content is defined as high-risk (rubric §3.3), either add a minimal tool-decode baseline for T3B or state more sharply that current T3 performance is insufficient for deployment and that Success on T3 should not be read as reliable reading.","section":null}],"minor_comments":[{"comment":"Abstract and §1: “local MLLM-driven” is accurate for vision–text MiniCPM via llama.cpp, but speech I/O is modular Whisper + SAPI5/pyttsx3 because GGUF does not support omni audio tokens (§2.1). One clarifying sentence in the abstract would prevent readers from assuming native omni streaming.","section":null},{"comment":"Table 1: MiniCPM-o 4.5 Quality 1.283 vs MiniCPM-V 4.5 Quality 1.325; the text says latency is comparable but does not discuss the small quality gap or when to prefer V vs o as a vision–text backend.","section":null},{"comment":"Figure 1 stage numbering is clear; consider adding approximate median ms per stage (from logs) so the 993 ms decomposition is visible without cross-referencing Tables 1 and 4.","section":null},{"comment":"§3.4 estimated ≈1.11 s from combining medians is helpful but slightly inconsistent with Table 4’s measured 993 ms; note that medians do not add and prefer the measured E2E as primary.","section":null},{"comment":"References: llama.cpp and MiniCPM-o are cited as GitHub repos with access dates; ensure version pins match the released configs for long-term reproducibility.","section":null},{"comment":"Typos / polish: “OPENGLASS” vs “OpenGlass” casing is inconsistent in places; “chinese cloud” → “Chinese cloud” (Baselines); ensure Table 6/7 are clearly labeled supplementary in the main text cross-refs.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Fit is solid for a systems/CV venue that values open assistive prototypes and careful latency measurement. Novelty is co-design and measurement rather than a new model; that is fine if the journal accepts systems papers. The MiniCPM lineage overlap with an author is disclosed via model choice and citations and does not appear circular for the latency claims. I would not reject on evaluation size alone given the explicit “reference platform” framing, but I would insist the authors not oversell BLV efficacy without user studies."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that they actually measured a full glasses-to-host loop under real ESP32 Wi-Fi and got 993 ms median user-to-audio with resize (97.5% under 2 s), 1625 ms raw, with open code, prompts, logs, and a clear non-navigation disclaimer. That latency result is the real contribution.\n\nWhat is new is not ESP32, MiniCPM, llama.cpp, Whisper, or TTS—those are known—but the co-design and instrumentation: sensing-computing split that keeps raw egocentric frames local by default, replay vs real Wi-Fi separation, resize/streaming/TTS ablations, resolution-slice trade-off, and failure attribution that correctly pins rare misses on capture/inference tails rather than TTS. Cloud baselines on the same capture path make the comparison fair. Shipping the stack and auditable logs is genuine systems work, not a demo video.\n\nSoft spots are real but proportionate. Quality, Success, Abstain, and HCE rest on 120 author-collected frames scored with an author 0/1/2 rubric, no IAA, no independent raters, no BLV users. Table 2 already shows the risk: T3A 37.5% HCE, T3B 66.7%, and T2 is mostly abstention counted as usable. So the assistive-quality half of the framing is weaker than the latency half; the paper is honest about this in Limitations, but the abstract still packages both. The everyday substrate is also a nearby RTX-class laptop, not a phone-only or glasses-only path—fine as a reference platform if stated that way.\n\nMath is not the point; wall-clock stages and open artifacts are. Citations cover VizWiz, cloud MLLMs, edge VLMs, and MiniCPM lineage without pretending the split is a new theorem. For people building local assistive or edge multimodal systems, this is useful. I would send it to referees: accept the engineering claim, push for clearer host constraints and stronger human evaluation. Worth engaging if you care about deployable local MLLM assistance.","headline":"Solid open systems paper on real Wi-Fi sub-2s local assistive latency; quality/safety half is thinner than the latency half.","tokens_in":14697,"tokens_out":534,"would_cite":true,"duration_ms":5377,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A sensing-computing split puts local MLLM visual assistance under 1 s median for glasses users.","keywords":["local multimodal LLM","visual assistance","blind and low-vision","sensing-computing split","ESP32 wearable","end-to-end latency","safety-aware abstention","privacy-oriented inference"],"falsifier":"Run the same ESP32 Wi-Fi capture pipeline with a phone-class host (no laptop GPU) or with independent blind participants on the T1–T3 tasks; if median user-to-audio exceeds 2 s or high-confidence sign/QR fabrications rise sharply, the practicality claim fails.","tokens_in":14565,"feed_emoji":"👁️","tokens_out":680,"duration_ms":5988,"temperature":0.7,"pith_summary":"OpenGlass argues that cloud multimodal assistants are a poor default for blind and low-vision users: they force upload of first-person images and often add multi-second network delay. Wearable glasses are good at sensing but cannot host large models. The paper shows a practical middle path: split sensing from computing. A low-cost ESP32 glasses unit captures a frame on demand; a nearby consumer laptop runs a quantized local vision-language model and streams speech locally. Under real Wi-Fi capture the full pipeline reaches 993 ms median time from query-ready to first audio with resized frames (97.5% of trials under 2 s) and 1625 ms with raw 1280×720 frames (93.3% under 2 s), while raw egocentric images stay on devices the user controls. The system is framed as a user-initiated reference platform for hazard awareness, object and sign queries, and image-quality self-checks with safety-aware abstention—not a certified navigation aid—and the authors release hardware, code, prompts, and logs so others can reproduce the loop.","feed_headline":"Local glasses AI answers in under a second without cloud upload","feed_subtitle":"ESP32 sensing plus nearby laptop MiniCPM hits 993 ms median speech, keeps first-person frames local","key_machinery":"The sensing-computing split architecture: lightweight glasses-side capture (ESP32-S3 OV5640, on-demand JPEG over Wi-Fi) plus host-side packing, local llama.cpp VLM, sentence-level streaming flush into local TTS, and safety-first prompts that force abstention and retake guidance when evidence is weak.","core_discovery":"A sensing-computing split that pairs ESP32 glasses capture with nearby consumer-grade local MiniCPM inference and streaming TTS can deliver practical sub-2 s, mostly sub-second user-to-audio visual assistance while keeping raw first-person images off the cloud by default, outperforming both overseas and domestic cloud VLMs on end-to-end latency under the same wearable capture path.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["ESP32 glasses plus local MiniCPM hit 993 ms speech without cloud","Sensing-computing split keeps frames local for sub-second visual answers","Nearby device runs MiniCPM on wearable capture for private under-2s aid","OpenGlass pairs ESP32 sensing with local inference for real-time help","Raw first-person images stay on-device as local MLLM replies in 1 s"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That a nearby consumer GPU laptop and local Wi-Fi are an acceptable everyday substrate for assistive use, and that a 120-frame author-collected rubric set is enough to claim practical quality and safety without blind-user studies.","fun_headline_variants_meta":{"raw":{"variants":["ESP32 glasses plus local MiniCPM hit 993 ms speech without cloud","Sensing-computing split keeps frames local for sub-second visual answers","Nearby device runs MiniCPM on wearable capture for private under-2s aid","OpenGlass pairs ESP32 sensing with local inference for real-time help","Raw first-person images stay on-device as local MLLM replies in 1 s"]},"model":"grok-4.5","effort":"low","cost_usd":0.00504,"raw_usage":{"total_tokens":1450,"prompt_tokens":819,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":50400000,"prompt_tokens_details":{"text_tokens":819,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":545,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":819,"tokens_out":86,"duration_ms":5277,"temperature":1.0,"reasoning_tokens":545,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:04:56.851648+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same ESP32 Wi-Fi capture pipeline with a phone-class host (no laptop GPU) or with independent blind participants on the T1–T3 tasks; if median user-to-audio exceeds 2 s or high-confidence sign/QR fabrications rise sharply, the practicality claim fails.","supporting_citations":[],"review_version":1}