{"id":"3d8f151d-5e74-495f-b04a-22ffc68f40f2","arxiv_id":"2508.16143","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MIEL integrates sound-source localization and one round of GPT-4o clarifying questions into exophora resolution, achieving about 2x higher success for out-of-view users in 90 real-world trials.","lead":"This paper introduces MIEL, a robot system that uses sound localization, a semantic map, and one GPT-4o clarifying question to interpret ambiguous commands like 'take that,' even when the user is out of view. In a simulated-home test, it roughly doubled success over a prior method for invisible users, though humans were still nearly twice as accurate.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Not-visible 2.0x gain is assumed, not measured: SSL success within 29° is treated as guaranteeing usable skeleton after reorientation, making Table II identical to Table I.","rationale":"The reader's weakest assumption identifies the same underlying issue: the not-visible condition is evaluated with pre-recorded skeletal/SSL data and an untested all-success assumption about physically reorienting the robot. This is the main threat to the central quantitative claim because the reported equality of visible and not-visible MIEL performance (0.53 vs 0.53) is a direct consequence of that assumption, not an observed outcome. The concern is concrete and falsifiable: a live closed-loop run with actual SSL and skeleton detection after turning would settle it. It does not require rejecting the architecture or the visible-user results; it only means the 2.0x not-visible improvement is conditional on a step the paper never measures. The reader's CONDITIONAL verdict already captures this appropriately, so no verdict change is needed.","tokens_in":11895,"tokens_out":5025,"duration_ms":57294,"concrete_test":"Run the closed-loop not-visible experiment exactly as specified in Section V-B: robot starts with the user out of view, uses the ReSpeaker array to localize the user, rotates to the estimated angle, then attempts MediaPipe skeleton detection and pointing-feature extraction before executing MIEL's estimators and optional question. Report the fraction of trials with usable skeleton data and the resulting Top-1 SR, with SSL angular errors. If the SR drops below about 0.54 (or if the skeleton acquisition rate is materially below 100%), the claimed 2.0x improvement over ECRAP is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is the pre-recorded closed-loop equivalence in the out-of-view condition (Section V-B1). MIEL's not-visible success rate is not measured end-to-end: user skeletal data and SSL results were collected in advance; when the user is initially out of view, the system assumes that an SSL angular error below 29° (half the camera FOV) means the robot can turn and obtain usable MediaPipe skeleton/pointing data. This assumption does all the work in Table II: the paper states SSL succeeded in all trials, so the not-visible SR is identical to the visible SR (0.53 Top-1). Thus the '2.0x over ECRAP' claim is not an empirical result of physically reorienting the robot; it is a conditional extrapolation from visible-condition data. If real closed-loop SSL/MediaPipe succeeds with probability p<1, the expected not-visible SR is at best p*0.53+(1-p)*0.32 (using the no-SSL Q&A-only ablation, Table III), and matching 2.0x ECRAP (0.54) requires p≈1. Any realistic skeleton-detection failure after reorientation weakens the headline, and the paper provides no data on that step. The authors' own Section V-F concedes SSL SR can decrease due to noise, yet the evaluation never exercises such a case.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MIEL, a multimodal exophora-resolution framework for ambiguous demonstrative instructions such as \"Take that for me.\" MIEL combines sound source localization (SSL), a 3D semantic map, VLM/CLIP features, user skeletal and pointing cues, and a GPT-4o-based interactive-questioning module. Experiments are conducted in a simulated home environment with a real robot, comparing MIEL against VGPN, ECRAP, and human toplines under visible-user and not-visible-user conditions, with query-information levels 1–3 and ablations. The central claim is that SSL makes the not-visible condition perform as well as the visible condition, yielding about 1.3x and 2.0x top-1 success rates over ECRAP in the visible and not-visible conditions, respectively.","tokens_in":12220,"tokens_out":6406,"duration_ms":70695,"significance":"If the reported results hold, MIEL would be a useful step toward robots that resolve ambiguous referring expressions when the user is outside the camera's field of view. The paper's strengths are its real-world setup, the use of a reasonably large semantic map (114 objects), 90 systematically degraded linguistic queries, a human topline, and ablations separating SSL and interactive questioning. However, the headline out-of-view result is currently a conditional extrapolation rather than a measured closed-loop outcome: skeletons and SSL outputs were pre-recorded, SSL success is equated with an angular error below 29 degrees, and the paper assumes that a successful SSL turn guarantees usable skeletal data. This leaves the main contribution of SSL for out-of-view users under-supported. The paper is readable and the system idea is valuable, but the experimental evidence needs strengthening before the central claim can be accepted.","major_comments":[{"comment":"The not-visible-user success rate is not measured end-to-end. The authors state that skeletal data and SSL results were collected in advance and that, if SSL succeeds, the robot 'can detect the skeleton by turning around.' Since SSL succeeded in all trials (Section V-E), Table II's MIEL row is identical to Table I's by construction. The reported 2.0x improvement over ECRAP in the not-visible condition is therefore an extrapolation under the assumption p=1 for reorientation plus skeleton detection, not an empirical result of physically reorienting the robot. The paper's own Section V-F concedes that SSL performance can degrade with noise, but no such failure case is evaluated. At minimum, the authors should report the success rate of the full reorientation→skeleton-detection step and run the not-visible condition live, or explicitly re-frame Table II as a simulation under the p=1 assumpti","section":"Section V-B1, Tables I-II, Section V-E"},{"comment":"The 29-degree SSL success criterion equates angular accuracy of the sound-source direction with the usability of the downstream skeleton/pointing pipeline. A turn that places the user within the camera's field of view can still fail to yield a MediaPipe skeleton because of distance, occlusion, lighting, or motion; no measurement of skeleton-detection success after reorientation is reported. Thus the paper's first contribution, 'demonstrated effectiveness of SSL for acquiring user skeletal data,' overstates what was actually measured: only SSL direction error was evaluated. Please add data on skeleton-detection success from the turned viewpoint and, ideally, pointing-estimation accuracy under the not-visible condition.","section":"Section V-B1"},{"comment":"All success rates are point estimates with N=30 per condition/level and no confidence intervals or significance tests. Several differences that support the main claims are small; for example, Level-1 visible MIEL (0.63) versus ECRAP (0.57) is a difference of 2 trials out of 30, and Level-2 visible MIEL (0.60) versus ECRAP (0.53) is also 2 trials. The '1.3x' and '2.0x' claims should be accompanied by exact binomial confidence intervals or an appropriate significance test. Without this, the improvement over ECRAP, particularly within individual query levels, is not established beyond sampling noise.","section":"Section V-D, Tables I-III"},{"comment":"The not-visible comparison between MIEL and ECRAP uses mismatched denominators. ECRAP's total is 16/60 because 30 Level-3 trials are excluded, while MIEL's total is 48/90. The text also says MIEL 'outperforms ECRAP by a factor of three,' whereas 0.53/0.27 is close to 2.0. The reported factor depends on how ECRAP's unanswerable Level-3 trials are handled: if counted as failures, ECRAP becomes 16/90=0.18 and the ratio is about 2.9; if restricted to the 60 shared Level-1/2 trials, MIEL is 37/60=0.62 and ECRAP is 16/60=0.27, a ratio of about 2.3. Please specify the analysis and keep denominators consistent, and reconcile the 'factor of three' wording with the abstract's '2.0 times.'","section":"Section V-E, Table II"}],"minor_comments":[{"comment":"The success-rate formula is typeset as SR = 1NPNi=1Si; it should be SR = (1/N) Σ S_i. Please fix the formatting for clarity.","section":"Section V-D"},{"comment":"The checkmarks in the ablation table are not aligned with the module headers, and the rows are not explicitly labeled as SSL-only, Q&A-only, and full MIEL. The reader has to infer which ablation is which from the totals and the text; please make the row labels explicit.","section":"Table III"},{"comment":"The text says non-English queries are translated by GPT-4o before encoding, and the experiments were in Japanese. Please clarify whether all 90 queries were translated, whether the translations were checked, and what impact translation errors could have on the results.","section":"Section IV-A"},{"comment":"The Human (w/o Q&A) and Human (topline) protocols are described only briefly. It would be useful to know how many human subjects participated, how instructions were given, and whether the same subjects also assessed the visible and not-visible conditions.","section":"Section V-C"},{"comment":"The demonstrative-region Gaussian variances and the von Mises concentration parameter are not reported; the paper says only that the formulas are 'the same as in [3].' If the values are reused from ECRAP, say so explicitly and give the values or a pointer to where they are specified, for reproducibility.","section":"Sections IV-C and IV-D"}],"recommendation":"major_revision","confidential_remarks":"The manuscript presents a useful integrated system and the visible-condition evaluation is informative, but the not-visible headline result is currently an optimistic simulation rather than a measured closed-loop result. I would want to see either a live reorientation experiment or a clearly labeled conditional analysis before accepting the main claim. The denominator inconsistency in Table II compounds this and should be corrected. The paper is within scope for a robotics venue and the idea is worth publishing after these fixes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a credible integration paper that probably sets the right direction for exophora resolution with out-of-view users, but the headline 2.0x improvement in the not-visible condition is assumed rather than measured. Read it as a promising system with a conditional empirical claim, not as a demonstration.\n\nWhat is actually new: the specific combination of sound-source localisation, NLMap-style semantic maps, CLIP features, and one round of GPT-4o questioning for exophora resolution. The ablation cleanly separates the contributions: interactive questioning helps most when the query lacks class or feature words, and SSL is what lets the system keep performance when the user is initially out of view. The experiments use a real home-like environment with 114 objects and 90 queries, and compare against VGPN, ECRAP, and human baselines. The paper is refreshingly honest in its limitation section, listing stale semantic maps, SSL noise sensitivity, and user burden.\n\nThe main soft spot is in the evaluation of the not-visible condition. Skeletal and SSL data were recorded in advance; SSL was deemed successful whenever its angular error was within 29 degrees, and the paper states SSL succeeded in all trials, so Table II is identical to Table I. That means the 2.0x gain over ECRAP for out-of-view users comes from the assumption that reorienting the robot toward the SSL-estimated direction will always let it detect the user's skeleton and pointing. The paper never tests that closed-loop step. If skeleton detection fails after reorientation with any real probability, the headline gain shrinks, and the authors' own limitations mention SSL can degrade with noise. There are also smaller issues: no confidence intervals or significance tests on 30-trial point estimates, and no code or data.\n\nNone of this breaks the central architecture, and the visible-condition results are informative. But the out-of-view claim should be treated as a hypothesis until an end-to-end run is done or the claim is explicitly conditioned on SSL success. I'd send this to peer review; a revision that either runs the live loop or softens the claim would make it solid. For my own work, I'd cite it cautiously as related work.","headline":"A promising integration paper whose out-of-view 2.0x result is an assumption, not a measured outcome.","tokens_in":12744,"tokens_out":3028,"would_cite":false,"duration_ms":30265,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A home robot can identify the object behind 'that' even when the user is out of view by turning toward the speaker's voice and, if still unsure, asking one clarifying question.","keywords":["exophora resolution","sound source localization","interactive questioning","semantic mapping","vision-language models","demonstratives","service robots"],"falsifier":"Run MIEL live in the same apartment with an initial robot pose that hides the user and with variable acoustic noise, so SSL errors near the 29-degree threshold occur naturally. If the top-1 success rate in the not-visible condition drops below the pre-recorded 0.53, or if reorientation fails to produce a skeletal keypoint in a nontrivial share of trials, the paper's closed-loop claim is falsified.","tokens_in":11787,"feed_emoji":"🤖","tokens_out":5723,"duration_ms":61221,"temperature":0.7,"pith_summary":"The paper aims to make exophora resolution (figuring out what 'that' refers to when a person says 'take that') work in real home conditions where the user or target object is not visible to the robot. The proposed MIEL system first uses sound source localization to turn the robot toward an out-of-view user so it can capture pointing gestures and body position. It then combines those cues with a semantic map and visual-language features to rank candidate objects, and if the top candidates are still ambiguous, it asks the user one clarifying question generated by GPT-4o. In a real apartment-like test setting, MIEL achieved 0.53 top-1 success both when the user was visible and when the user was not, about 1.3 times and 2.0 times the success rate of the ECRAP baseline. The central claim is that these two additions, sound-based user localization and one round of interactive questioning, are enough to keep performance stable when observational data are incomplete.","feed_headline":"Sound-based reorienting doubles robot's success on 'that' tasks","feed_subtitle":"Turning the robot toward the speaker by sound keeps 'which object' accuracy at 53 percent in home tests.","key_machinery":"The load-bearing mechanism is the coupling of sound source localization with the rest of the pipeline: SSL turns an initially invisible user into a visible one, restoring skeletal data and pointing direction that feed two of the three estimators. The complementary mechanism is a single round of interactive questioning triggered when GPT-4o cannot identify the target from the top-five candidates, which compensates for information-poor instructions such as 'Bring me that.' A third component, the 3D semantic map built from NLMap, supplies object labels, visual features, and coordinates that let the linguistic-query estimator match both object class and visual attributes.","core_discovery":"The central discovery is that a service robot can maintain the same exophora-resolution accuracy whether or not the user is inside its camera view, provided it can locate the speaker by sound and rotate toward them to recover skeletal and pointing data, and ask a single GPT-4-generated clarifying question when the query lacks object class or attribute information. MIEL combines three probabilistic estimators (linguistic-query similarity, demonstrative-region Gaussian, and pointing-direction von Mises) whose probabilities are multiplied and given to GPT-4o for top-5 ranking and optional interactive questioning. The reported experiments show top-1 success of 0.53 in both visible and non-visibl","pith_inferences":["Left implicit: the 29-degree SSL acceptance threshold implies a decision rule: if sound-direction uncertainty exceeds half the camera field of view, the robot should request a repetition or another cue rather than rotate blindly.","A testable extension: vary the number of question rounds or let the robot ask about object location instead of attributes; the paper's single-round limit is a design choice, not a demonstrated optimum.","The stability across visible and non-visible conditions suggests SSL could be replaced by any user-localization modality (for example, voice identification from multiple microphones) that supplies the user's bearing when vision fails.","In a multi-user home, the framework would need speaker diarization or voice identity to know which sound source to reorient toward; the paper's single-user setup leaves this unaddressed."],"forward_implications":["Robot exophora resolution need not degrade when a user moves out of the camera's field of view; sound source localization can substitute for visual user localization.","A single round of interactive questioning is enough to make information-poor instructions (just 'that') usable, doubling the success rate on such queries relative to no questioning.","Semantic mapping and visual-language features let the system exploit object attributes such as color, which the earlier ECRAP baseline could not handle.","Even with these additions, robot performance (0.53 top-1) remains well below human performance (0.86 to 0.98), so further cues such as gaze and better question generation are needed to close the gap.","The Top-5 success of 0.79 means the target is usually in the robot's shortlist even when top-1 misses, suggesting that downstream interaction or confirmation could recover many failures."],"supporting_citations":[{"why":"Supplies the formulas for the demonstrative-region Gaussian and pointing-direction von Mises estimators used by MIEL.","marker":"[3]"},{"why":"CLIP text and image encoders provide the learned features that match the linguistic query to object visual attributes.","marker":"[6]"},{"why":"ECRAP is the primary baseline; its lack of SSL and interactive questioning isolates MIEL's contribution.","marker":"[13]"},{"why":"GPT-4o is used to extract demonstratives, rank top-five candidates, identify the target, and generate clarifying questions.","marker":"[14]"},{"why":"CLARA motivates the interactive-questioning design used when exophora resolution remains ambiguous.","marker":"[18]"},{"why":"Sentence-BERT encodes the label features and linguistic query for the linguistic-query estimator.","marker":"[19]"},{"why":"NLMap is extended into the 3D semantic map that supplies object labels, visual features, and coordinates.","marker":"[20]"},{"why":"MediaPipe provides the skeletal detection that yields eye and wrist coordinates for pointing estimation.","marker":"[21]"},{"why":"Detic supplies object detections from exploration images for building the semantic map.","marker":"[24]"},{"why":"Objects365 provides the label vocabulary used by the object detector during semantic-map construction.","marker":"[25]"}],"fun_headline_variants":["Sound-guided robot doubles accuracy on 'that' tasks","Turning by sound keeps robot's 'that' resolution at 53%","MIEL: robot locates speaker by sound to resolve 'that'","Hearing helps: robot asks to disambiguate out-of-view 'that'"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The experiments used skeleton and sound-source data recorded in advance, and SSL was counted as successful whenever its angular error was within 29 degrees; the results assume that physically rotating the robot to the estimated direction will always capture usable skeletal and pointing data in live operation.","fun_headline_variants_meta":{"raw":{"variants":["Sound-guided robot doubles accuracy on 'that' tasks","Turning by sound keeps robot's 'that' resolution at 53%","MIEL: robot locates speaker by sound to resolve 'that'","Hearing helps: robot asks to disambiguate out-of-view 'that'"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000475,"raw_usage":{"total_tokens":2202,"prompt_tokens":757,"completion_tokens":1445,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":1368}},"tokens_in":501,"tokens_out":1445,"duration_ms":11339,"temperature":1.0,"reasoning_tokens":1368,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:29:15.155658+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run MIEL live in the same apartment with an initial robot pose that hides the user and with variable acoustic noise, so SSL errors near the 29-degree threshold occur naturally. If the top-1 success rate in the not-visible condition drops below the pre-recorded 0.53, or if reorientation fails to produce a skeletal keypoint in a nontrivial share of trials, the paper's closed-loop claim is falsified.","supporting_citations":[{"cited_title":"Exophora Resolution of Linguistic Instructions with a Demonstrative based on Real-World Multimodal Information,","cited_arxiv_id":null,"evidence_quote":"Supplies the formulas for the demonstrative-region Gaussian and pointing-direction von Mises estimators used by MIEL."},{"cited_title":"Learning Transferable Visual Models from Natural Language Supervision,","cited_arxiv_id":null,"evidence_quote":"CLIP text and image encoders provide the learned features that match the linguistic query to object visual attributes."},{"cited_title":"ECRAP: Exophora Resolution and Classifying User Commands for Robot Action Planning by Large Language Models,","cited_arxiv_id":null,"evidence_quote":"ECRAP is the primary baseline; its lack of SSL and interactive questioning isolates MIEL's contribution."},{"cited_title":"CLARA: Classifying and Disambiguating User Com- mands for Reliable Interactive Robotic Agents,","cited_arxiv_id":null,"evidence_quote":"CLARA motivates the interactive-questioning design used when exophora resolution remains ambiguous."},{"cited_title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,","cited_arxiv_id":null,"evidence_quote":"Sentence-BERT encodes the label features and linguistic query for the linguistic-query estimator."},{"cited_title":"Open-V ocabulary Queryable Scene Representations for Real World Planning,","cited_arxiv_id":null,"evidence_quote":"NLMap is extended into the 3D semantic map that supplies object labels, visual features, and coordinates."},{"cited_title":"Detecting Twenty-Thousand Classes using Image- Level Supervision,","cited_arxiv_id":null,"evidence_quote":"Detic supplies object detections from exploration images for building the semantic map."},{"cited_title":"Objects365: A Large-Scale, High-Quality Dataset for Object Detection,","cited_arxiv_id":null,"evidence_quote":"Objects365 provides the label vocabulary used by the object detector during semantic-map construction."}],"review_version":1}