{"id":"9639f8ff-2baa-41c6-8b6c-739f3064b439","arxiv_id":"2607.06165","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.","lead":"EAGOR is a training-free framework that lets robots with 360-degree cameras maintain a consistent sense of target direction by representing spatial beliefs directly on a sphere rather than on a flattened panoramic image. A smart generalist might read it because it offers a geometric fix for why vision-language models fail at navigation when the robot turns, potentially improving robotic search and map-free navigation without requiring model retraining.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Headline gains on HOS/OSR-Bench may stem from temporal accumulation rather than the spherical harmonic representation specifically, as no ERP-space temporal accumulation baseline is included for those benchmarks.","rationale":"The reader correctly identifies the likelihood calibration concern, and it is a real theoretical issue. However, the empirical results—particularly the failure mode analysis (Table 4) showing that failures stem from VLM perception errors rather than systematic directional bias—suggest that uncalibrated likelihoods are not the binding constraint in practice. The system works despite uncalibrated scores because the VLM response maps do carry useful directional signal, and temporal accumulation provides robustness to per-frame noise. The more load-bearing concern is that the headline benchmark gains cannot be cleanly attributed to the spherical harmonic representation because the Active Visual Search experiments lack an ERP-space temporal accumulation baseline. The navigation experiments do include such a baseline ('Grid'), and EAGOR outperforms it, which provides some evidence that the spherical representation matters. But the headline numbers (+34.4%, +45.6%) come from HOS and OSR-Bench, where no such baseline exists. This is a gap in experimental design rather than a fundamental flaw in the approach. The mathematical formulation is sound: the SH representation does provide equivariant rotation propagation (Eq. 4 via Wigner-D matrices), eliminates ERP seam discontinuities, and the log-space additive update (Eq. 5) is a valid Bayesian accumulation. The concern is about whether the experimental evidence supports the specific attribution of gains to these properties. I recommend UNCHANGED because the reader's CONDITIONAL verdict already accounts for insufficient evidence (missing code, qualitative real-world validation, no error bars), and this concern falls under the same category. However, the reader should note that the headline gains' attribution to the spherical representation is unestablished for the primary benchmarks, which is a more specific and actionable gap than the general calibration concern.","tokens_in":14685,"tokens_out":2834,"duration_ms":141963,"concrete_test":"Implement a simple temporal accumulation baseline in ERP space for HOS and OSR-Bench: at each view, extract the VLM response map, rotate it into a common egocentric frame using the known agent rotation, accumulate log-response maps additively across views, and extract the peak direction. Compare its accuracy against EAGOR on the same HOS/OSR-Bench splits. If the ERP-accumulation baseline achieves within 5% of EAGOR's accuracy, the spherical harmonic representation's contribution to the headline gains is marginal and the novelty claim weakens significantly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's concern about VLM likelihood calibration is valid but partially addressed by empirical results showing the system works despite uncalibrated scores. The more load-bearing issue is experimental attribution. The headline claims (+34.4% on HOS, +45.6% on OSR-Bench) compare EAGOR—which combines (a) spherical harmonic belief representation, (b) equivariant rotation propagation, and (c) temporal evidence accumulation—against standalone single-frame VLMs that lack any temporal accumulation. For the navigation tasks (Tables 2–3), the paper includes a 'Grid' baseline that does temporal accumulation in ERP space, and EAGOR outperforms it. But for Active Visual Search (Table 1), no ERP-space temporal accumulation baseline is included. This means the gains on HOS and OSR-Bench could be largely attributable to temporal evidence accumulation (which any simple scheme could provide) rather than to the spherical harmonic representation specifically. The paper's core novelty claim—that maintaining belief on the sphere via SH is superior to ERP-based approaches—is not directly tested on the benchmarks that produce the headline numbers. The SH bandlimit ablation (Fig. 7) only varies L within EAGOR; it does not compare against non-SH temporal accumulation. Without an ERP-accumulation baseline on HOS/OSR-Bench, the contribution of the spherical representation to the headline gains is unestablished.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The paper introduces EAGOR, a training-free framework for embodied omnidirectional reasoning that formulates target-direction estimation as recursive Bayesian filtering on the sphere. The core technical contribution is the Spherical Harmonic Belief Field (SH-BF), which represents directional belief using spherical harmonics, propagates it equivariantly under agent rotation via Wigner-D matrices, and decodes the target direction via the Fréchet mean. The approach is evaluated on waypoint following, map-free navigation (Habitat-Sim), active visual search (HOS, OSR-Bench), and real-world dynamic tracking on a legged robot. The central claim is that maintaining belief directly on the sphere avoids ERP seam discontinuities, latitude distortions, and interpolation errors, yielding consistent directional estimates under ego-motion.","tokens_in":15428,"tokens_out":1315,"duration_ms":225854,"significance":"The paper addresses a genuine representation gap in embodied 360° reasoning: treating VLM outputs as directional likelihoods on S² rather than as ERP pixel coordinates. The SH-BF formulation is clean and parameter-free in its core derivation (Eqs. 2–6), grounding the belief representation in well-established mathematical frameworks (spherical harmonics, Wigner-D rotations, Bayesian filtering). The equivariant propagation via Wigner-D matrices in coefficient space is a principled solution to the motion-consistency problem. The real-world deployment on a legged robot and the evaluation across multiple VLM backbones (Qwen2.5-VL, Gemma-3) strengthen the practical relevance. The framework is training-free and model-agnostic, which is a notable strength for adoptability. However, the experimental attribution of headline gains to the spherical harmonic representation specifically—rather than to temporal evidence accumulation more generally—is not fully established on the benchmarks producing the headline numbers.","major_comments":[{"comment":"§4.2, Table 1 (HOS and OSR-Bench): The headline gains (+34.4% on HOS, +45.6% on OSR-Bench) compare EAGOR—which combines (a) spherical harmonic belief representation, (b) equivariant rotation propagation, and (c) temporal evidence accumulation—against standalone single-frame VLM baselines that lack any temporal accumulation. For the navigation tasks (Tables 2–3), an ERP-space temporal accumulation baseline ('Grid') is included, and EAGOR outperforms it. However, for Active Visual Search (Table 1), no ERP-space temporal accumulation baseline is included. This means the gains on HOS and OSR-Bench could be largely attributable to temporal evidence accumulation (which any simple scheme could provide) rather than to the spherical harmonic representation specifically. The paper's core novelty claim—that maintaining belief on the sphere via SH is superior to ERP-based approaches—is not directly ","section":null},{"comment":"§3.1: The interpretation of the VLM's target-conditioned response map ℓ_t(u,v) as a directional likelihood field proportional to the log-likelihood of the target existing in that viewing direction is the foundational assumption of the entire framework. VLM attention scores are uncalibrated heuristics, and the paper does not validate that ℓ_t(ω) behaves as a proper likelihood (e.g., that its peaks correspond to higher target probability, that its relative magnitudes are meaningful). While empirical results show the system works in practice, the paper would benefit from either (i) a calibration analysis showing that VLM response maps correlate with target presence probability, or (ii) an explicit acknowledgment that the log-likelihood interpretation is an approximation and a discussion of conditions under which it may fail. This is load-bearing because the Bayesian update (Eq. 5) and theFr","section":null}],"minor_comments":[{"comment":"§3.2, Eq. (5): The additive update c^(t) = c̃^(t) + b^(t) corresponds to log-space Bayesian fusion with equal weighting of prior and observation. This implicitly assumes that the observation noise is stationary and uniform across directions. A brief discussion of why equal weighting is appropriate (or whether a discount factor on the prior would be beneficial) would strengthen the presentation.","section":null},{"comment":"§3.2: The statement 'L=7 is the SH bandlimit' is introduced without justification in the main text. Fig. 7 provides the ablation, but the choice of L=7 should be cross-referenced when first introduced.","section":null},{"comment":"§4.1: The baselines for HOS and OSR-Bench are described as 'fine-tuned and zero-shot VLM baselines,' but Table 1 only shows zero-shot VLM results. If fine-tuned baselines were evaluated, their results should be reported; if not, the description should be corrected.","section":null},{"comment":"Fig. 7: The x-axis label 'Angular Separation δ' and the curve labels 'L=12 (azimuth. min)' are unclear. Clarifying what 'azimuth. min' refers to would help the reader.","section":null},{"comment":"§4.2, Table 2: The 'Grid' baseline shows a MAE of 138.9° on segment L2, which is extremely high and suggests a systematic failure (e.g., 180° ambiguity). A brief note explaining this failure mode would help the reader interpret the comparison fairly.","section":null},{"comment":"The abstract states 'reducing step count by 17.7%,' but Table 3 shows steps reduced from 61.0 (Centroid) to 50.2 (EAGOR), which is a 17.7% reduction relative to Centroid. Clarifying which baseline is the reference for each percentage would improve precision.","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the missing ERP-accumulation baseline on HOS/OSR-Bench is the most important issue. The navigation experiments (Tables 2–3) do include a Grid baseline with temporal accumulation, and EAGOR outperforms it—but the headline numbers come from Table 1, where no such baseline exists. If the authors add even a simple ERP-space temporal accumulation baseline to Table 1 and EAGOR still outperforms it, the contribution is substantially strengthened. Without it, the reviewer community will question whether the spherical representation or the temporal accumulation is doing the heavy lifting. The VLM likelihood calibration concern is valid but secondary; the empirical results suggest the system is robust to miscalibration in practice."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The referee correctly identifies two important gaps in our experimental design and theoretical framing. We address both below and commit to revisions.","responses":[{"response":"The referee is correct. For the navigation tasks (Tables 2–3), we included the 'Grid' baseline, which performs temporal accumulation in ERP pixel space, and EAGOR outperforms it. However, for Active Visual Search (Table 1), we did not include an ERP-space temporal accumulation baseline, which means the attribution of the +34.4% and +45.6% gains to the spherical harmonic representation specifically is not directly supported by the current experiments on those benchmarks. This is a valid gap in our experimental design. We will address it in the revision by adding an ERP-space temporal accumulation baseline (analogous to the 'Grid' baseline used in Tables 2–3) to the HOS and OSR-Bench experiments in Table 1. This will allow a direct comparison between temporal accumulation in ERP space versus temporal accumulation on the sphere via SH-BF, isolating the contribution of the spherical representation from the contribution of evidence accumulation per se. We expect the spherical representation to retain an advantage—particularly under seam crossings and rotation, as demonstrated in Tables 2–3—but we agree the reader should be able to see this directly for the active visual search benchmarks as well. We will also temper the language in the abstract and main text to clarify that the gains reflect the combination of (a) spherical belief representation, (b) equivariant propagation, and (c) temporal accumulation, and that the relative contribution of each component is partially—but not fully—disentangled across all benchmarks.","revision_made":"yes","referee_comment":"§4.2, Table 1 (HOS and OSR-Bench): No ERP-space temporal accumulation baseline is included for Active Visual Search, so headline gains could be attributable to temporal evidence accumulation rather than the spherical harmonic representation specifically."},{"response":"The referee raises a legitimate concern. We do not claim that VLM response maps are calibrated probabilities, and our use of the log-likelihood formulation is best understood as a modeling assumption: we treat the VLM's spatial response as a directional evidence signal that, when accumulated recursively, yields a useful posterior over target directions. We agree that this should be stated more explicitly rather than left implicit. In the revision, we will: (i) add an explicit statement in §3.1 that the log-likelihood interpretation is an approximation, not a claim of calibrated probability; (ii) add a brief calibration analysis showing the empirical correlation between VLM response map peaks and target presence on a subset of HOS episodes, reporting rank correlation between response intensity and ground-truth target direction proximity; and (iii) expand the limitations discussion (currently in §5, Table 4) to explicitly address conditions under which the likelihood assumption breaks down—namely, multi-instance confusion (where multiple peaks of similar intensity exist), rare targets (where the response map may be flat or dominated by false positives), and fine-grained text/OCR tasks (where spatial attention may not reflect directional likelihood at all). Table 4 already provides some evidence for these failure modes; we will connect it more directly to the likelihood assumption. We note that the recursive Bayesian formulation is somewhat robust to miscalibration because it accumulates evidence over multiple views, which partially mitigates single-frame noise—a point we will also make explicit.","revision_made":"yes","referee_comment":"§3.1: The interpretation of VLM response maps as directional likelihood fields is unvalidated. VLM attention scores are uncalibrated heuristics; the paper should either provide a calibration analysis or explicitly acknowledge the approximation and discuss failure conditions."}],"tokens_in":14532,"tokens_out":945,"duration_ms":71134,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"EAGOR reformulates 360° directional reasoning as recursive Bayesian estimation on S² using spherical harmonics, which is a genuinely new application. The math is clean: VLM response maps are lifted to directional log-likelihoods on the sphere, projected into the SH basis, and propagated under agent rotation via Wigner-D matrices. Belief updates happen in coefficient space, and the final direction is decoded analytically via the degree-1 coefficients (Fréchet mean). This is parameter-free in its core derivation, and the decoupling of frozen VLM semantic evidence from geometric state estimation is a sensible architectural choice. The qualitative results (Figs. 4–5) convincingly show geometric consistency under seam crossings and ego-motion where ERP-based methods drift. For the navigation tasks (Tables 2–3), the paper includes a Grid baseline that does temporal accumulation in ERP space, and EAGOR beats it — this is the right comparison and it holds up. Credit is earned here. The SH bandlimit ablation (Fig. 7) is reasonable, and the runtime overhead is minimal. The failure mode analysis (Table 4) is honest about where the system breaks. The main problem is experimental attribution on the headline benchmarks. For HOS and OSR-Bench (Table 1), the baselines are standalone single-frame VLMs with no temporal accumulation. EAGOR combines three things: spherical representation, equivariant rotation propagation, and temporal evidence accumulation. On the benchmarks producing the +34.4% and +45.6% gains, there is no ERP-space temporal accumulation baseline to isolate the contribution of the SH representation specifically. Any simple temporal smoothing scheme in ERP space could capture much of that gain. The paper's core novelty claim — that maintaining belief on the sphere via SH is superior to ERP-based approaches — is directly tested in navigation but not on the active visual search benchmarks that generate the headline numbers. This is a real gap, not a minor one. The reader's concern about VLM likelihood calibration is valid in principle but less load-bearing in practice — the system works despite uncalibrated scores, and the paper acknowledges this. The lack of error bars and purely qualitative real-world validation are secondary issues. This paper deserves a serious referee. The formulation is novel and the math is sound, but the reviewer should require an ERP-accumulation baseline on HOS/OSR-Bench to establish that the gains come from the spherical representation rather than from temporal smoothing alone.","headline":"Novel SH-based spherical belief formulation for 360° reasoning; headline gains on HOS/OSR-Bench lack the right baseline to attribute gains to the spherical representation.","tokens_in":15356,"tokens_out":1076,"would_cite":false,"duration_ms":87277,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Spherical belief field lets robots reason in 360° without training","keywords":[],"falsifier":"If the VLM's target-conditioned response map does not correlate monotonically with true target direction probability—e.g., if attention peaks on distractors or is distorted by ERP artifacts before lifting—then the SH-BF Bayesian update will accumulate biased evidence and the belief will diverge from the true target direction over time, producing worse estimates than single-frame prediction.","tokens_in":14980,"feed_emoji":"🧭","tokens_out":823,"duration_ms":199513,"temperature":0.7,"pith_summary":"EAGOR reframes embodied directional reasoning with 360° cameras as a recursive Bayesian estimation problem on the sphere rather than pixel-coordinate prediction in a flat equirectangular projection (ERP) image. The central object is the Spherical Harmonic Belief Field (SH-BF), which represents the agent's continuous belief over possible target directions using spherical harmonic coefficients. At each timestep, a frozen vision-language model (VLM) produces a spatial response map over the panorama; EAGOR lifts this map onto the sphere as a directional log-likelihood, projects it into the spherical harmonic basis, and fuses it additively with a rotated copy of the previous belief. When the agent moves, the belief is propagated via Wigner-D rotation matrices that act directly on the harmonic coefficients, maintaining geometric consistency without re-querying the VLM. The final target direction is decoded analytically from the degree-1 coefficients as a spherical Fréchet mean. Because the belief lives on the sphere throughout, the framework eliminates the seam discontinuities, latitude-dependent distortions, and interpolation errors inherent to ERP representations. The entire pipeline requires no task-specific training of the backbone VLM. The paper claims that this geometric decoupling—VLM for semantic evidence, SH-BF for directional state estimation—yields substantial gains: +34.4% and +45.6% relative improvement on the HOS and OSR-Bench benchmarks respectively, a 14.6% navigation success increase, 17.7% fewer steps, and 24.5% lower angular error, with a 3B-parameter model outperforming a standalone 72B model on spatial reasoning.","feed_headline":"Spherical belief field lets robots reason in 360° without training","feed_subtitle":"By moving directional estimation off flat panoramas and onto the sphere, a 3B model beats a 72B model at knowing which way to turn.","key_machinery":"The Spherical Harmonic Belief Field (SH-BF) is a continuous belief representation over target directions on the unit sphere, expressed in the real spherical harmonic basis up to bandlimit L=7. It supports three operations in coefficient space: (1) projection of per-frame VLM response maps as directional log-likelihood observations, (2) Wigner-D rotation of the prior belief under agent ego-motion, and (3) additive Bayesian fusion of prior and observation. The MAP direction is decoded analytically from the degree-1 coefficients as a spherical Fréchet mean, avoiding grid search.","core_discovery":"The paper's central discovery is that treating VLM spatial attention as a directional likelihood on the sphere and accumulating it through a spherical-harmonic Bayesian filter produces geometrically consistent, motion-equivariant directional estimates that substantially outperform direct pixel-coordinate prediction from ERP images. The Spherical Harmonic Belief Field is the mechanism that makes this work: it provides a continuous, globally defined, rotation-aware representation that supports additive evidence accumulation in coefficient space and analytical direction decoding, all without training the VLM backbone.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["EAGOR: Training-free 360° reasoning via spherical Bayesian filtering","Spherical harmonics give VLMs consistent 360° directional reasoning","Moving 360° reasoning from flat projections to the sphere","Spherical harmonic belief field fixes 360° directional estimates","EAGOR: Geometry-aware spherical reasoning for 360° navigation"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The framework treats the VLM's spatially distributed attention patch as a calibrated directional likelihood field proportional to the log-probability of the target existing in each viewing direction. VLM attention scores are uncalibrated heuristics; if they systematically misrepresent true target probability, the recursive Bayesian update will propagate and amplify that bias rather than converge to the correct direction.","fun_headline_variants_meta":{"raw":{"variants":["EAGOR: Training-free 360° reasoning via spherical Bayesian filtering","Spherical harmonics give VLMs consistent 360° directional reasoning","Moving 360° reasoning from flat projections to the sphere","Spherical harmonic belief field fixes 360° directional estimates","EAGOR: Geometry-aware spherical reasoning for 360° navigation"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1029,"prompt_tokens":629,"completion_tokens":400,"prompt_tokens_details":null},"tokens_in":629,"tokens_out":400,"duration_ms":26222,"temperature":1.0,"reasoning_tokens":366,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T14:29:51.925407+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the VLM's target-conditioned response map does not correlate monotonically with true target direction probability—e.g., if attention peaks on distractors or is distorted by ERP artifacts before lifting—then the SH-BF Bayesian update will accumulate biased evidence and the belief will diverge from the true target direction over time, producing worse estimates than single-frame prediction.","supporting_citations":[],"review_version":1}