{"id":"01a85303-1501-4aa8-a4c6-65566148c320","arxiv_id":"2607.02565","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Yaw–pitch and Euler conformal scores redistribute coverage near coordinate singularities; geodesic scores restore slice-conditional reliability without retraining.","lead":"Chart-based conformal scores for gaze and head pose silently undercover near poles and gimbal lock by 30–50 points while overall coverage looks fine. Switching to geodesic scores fixes the distortion without retraining, which matters for safety-critical vision systems.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies the strongest claim (Proposition 2 + multi-dataset collapse) and the only real soft spot (first-order linearisation). That soft spot is not load-bearing: the impossibility result is purely algebraic once the score is radial in chart coordinates, and the controlled isotropic experiment already demonstrates the coverage redistribution without any appeal to the local Gaussian model. The paper’s own robustness sweeps (anisotropy up to 4:1, multiple σ, stronger backbones) further reduce residual doubt. Consequently the ACCEPT / HIGH-confidence verdict stands; no adjustment is warranted.","tokens_in":15507,"tokens_out":481,"duration_ms":4912,"concrete_test":"Recompute the |β|>60° row of Table 2 after replacing the isotropic noise with a fixed 4:1 anisotropic covariance deliberately anti-aligned with the metric eigenvectors; if Euler L2 coverage remains below ~70 % while geodesic stays near 90 %, the shape-distortion claim is confirmed even under the most favourable model anisotropy the paper itself considers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that chart-coordinate nonconformity scores (yaw–pitch L2 / Euler L2) produce acceptance sets whose local axis ratios are fixed by the metric-tensor eigenvalues, so slice-conditional coverage near singularities collapses even when marginal coverage is correct; geodesic scoring removes the distortion. Proposition 2 is elementary and does not rely on the linearisation of Proposition 1: any acceptance set that is a sublevel set of a monotone function of the chart residual norm is the preimage of a chart ball, whose pullback is an ellipsoid with fixed axis ratios √(λ_max/λ_min). The controlled BIWI experiment (isotropic Lie-algebra noise, Table 2) isolates score geometry from model error and already shows the 57.5 % collapse. Real-model tables (3–5) and the three-layer decomposition further separate scale from shape. The reader’s weakest assumption (finite-radius validity of the linearisation) is already flagged by the authors and is not load-bearing for the structural claim or the empirical pattern. No internal inconsistency or unacknowledged confound overturns the result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper shows that conformal prediction for manifold-valued vision outputs (gaze on S², head pose on SO(3)) can suffer large slice-conditional undercoverage near coordinate singularities when nonconformity scores are defined in chart coordinates (yaw–pitch or Euler L2), even when marginal coverage is correctly calibrated near 90%. Across ETH-XGaze, Gaze360, BIWI, and AFLW2000-3D, coverage drops by 30–50 pp near poles and gimbal lock. The authors prove (Proposition 2) that scalar adaptive methods (normalised CP, CQR-style threshold modulation) only rescale chart-space balls and cannot change axis ratios fixed by the metric-tensor eigenvalues. They propose the Riemannian volume density as a diagnostic and show that coordinate-free geodesic (or monotone-equivalent) scoring removes the chart-induced distortion without retraining and with negligible cost. A controlled SO(3) experiment with isotropic Lie-algebra noise isolates score geometry, and a three-layer decomposition separates chart distortion from heteroscedastic model error.","tokens_in":15739,"tokens_out":1094,"duration_ms":40717,"significance":"If the claims hold—and the evidence is strong—this is a practically important and theoretically clean contribution for reliability in geometric vision. Conformal prediction is increasingly used for distribution-free guarantees; the paper identifies a failure mode that is invisible under marginal metrics yet concentrated in safety-relevant regimes (extreme pitch, profile poses). Proposition 2 is an elementary but load-bearing geometric identity that cleanly rules out a natural class of remedies. The controlled BIWI experiment isolates score geometry from model error; the four-dataset audit, backbone ablations, and multi-chart/Mahalanobis comparisons make the empirical case robust. The proposed fix is immediately actionable (no retraining, ≤0.02 µs/sample). Credit is due for the clean impossibility result, the geometry-isolating controlled experiment, the three-layer decomposition, and the practical scoring protocol.","major_comments":[],"minor_comments":[{"comment":"In Sec. 4.3 the text reports normalised Euler coverage of 62.8±4.0% on AFLW2000-3D, but this number does not appear in Table 5. Adding it (or a short note) would make the Layer-2 vs Layer-3 comparison fully self-contained in the main table.","section":null},{"comment":"Sec. 3.2 (“Validity of the linearisation”) already notes that Prop. 1’s quantitative bounds can be loose at extreme angles. A single clarifying sentence earlier in Sec. 3.2 stating that Prop. 1 supplies local intuition and first-order predictions, while the structural claim rests on Prop. 2 and the finite-radius experiments, would help readers weight the two results correctly.","section":null},{"comment":"Fig. 3 is very effective; the shaded “degraded / severe / collapsed” bands are useful. Consider stating the exact coverage thresholds used for those bands in the caption so the figure is fully self-contained.","section":null},{"comment":"Sec. 5 mentions GazeTR-ViT near-pole numbers (30.3% YP, 81.3% norm. geodesic) that support the “stronger models amplify distortion” claim. A one-row summary in the main text (or a small table) would avoid forcing readers into the supplement for a headline ablation.","section":null},{"comment":"Notation: ρ(ξ) is introduced as volume density and later used as a correlation diagnostic. A brief reminder that for the standard charts ρ reduces to |cos(·)| (already stated) could be repeated once near Tables 2–5 where the ρ–coverage correlations are reported.","section":null},{"comment":"Minor typography: “T able” appears with a space in several table captions (e.g., “T able 1”, “T able 2”); fix to “Table”. Also “F unctions” in the Sec. 3.4 heading.","section":null},{"comment":"Sec. 6’s practical protocol is clear. A short decision note on when to prefer plain geodesic vs normalised geodesic vs Mondrian (once the base score is intrinsic) would help practitioners operationalise the three-layer decomposition.","section":null}],"recommendation":"accept","confidential_remarks":"Strong fit for a CV/ML venue that cares about reliability and geometric deep learning. The impossibility result is elementary but correctly scoped and well supported by a clean controlled experiment; I would not ask for major new theory. No novelty or citation concerns stood out."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: if you run split conformal on yaw–pitch or Euler residuals for gaze or head pose, you can keep 90% marginal coverage while dropping to ~39–57% in the high-pitch / near-gimbal-lock slices that matter for safety. The paper shows this on four datasets, proves that any scalar function of the chart residual norm cannot fix the local axis ratios (Prop. 2), and shows that switching to geodesic (or a monotone surrogate) removes the distortion with no retraining.\n\nWhat is actually new is not the classical chart distortion itself, nor the recent existence of geodesic conformal validity on manifolds. It is the quantitative slice-conditional collapse under standard vision scoring, the clean separation of scale vs shape (normalised chart still fails by tens of points; normalised geodesic recovers), and the controlled BIWI experiment with isotropic Lie-algebra noise that isolates geometry as the sole cause. The three-layer decomposition and the volume-density diagnostic are useful packaging. Math is elementary Riemannian geometry; the impossibility argument does not even need the linearisation of Prop. 1. Empirical pattern is consistent across backbones, including stronger models that actually make the chart failure worse.\n\nSoft spots are real but secondary. The first-order bounds get loose at extreme angles (authors say so). Everything is S²/SO(3); SE(3) and articulated pose are left as future work. No public code. Slice thresholds and the k-NN bandwidth for normalisation are free parameters, but the qualitative story does not hinge on them. None of this overturns the structural claim or the practical recommendation.\n\nThis is for people who put conformal wrappers on geometric vision outputs, and for anyone who still calibrates in Euler/yaw–pitch while reporting angular error. I would bring it to reading group, cite the impossibility + the coverage tables when I next touch manifold-valued UQ, and send it to peer review. The central result holds.","headline":"Chart-based conformal scores silently undercover near singularities by 30–50 pp; the impossibility result and controlled experiment make the claim solid, and geodesic scoring is a free fix.","tokens_in":16366,"tokens_out":494,"would_cite":true,"duration_ms":5330,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Chart-based conformal scores for gaze and head pose systematically undercover near singularities even when overall coverage looks correct.","keywords":["conformal prediction","gaze estimation","head pose estimation","coordinate singularities","Riemannian geometry","geodesic scoring","SO(3)","S2"],"falsifier":"On any of the four datasets, replace the chart-norm score with geodesic scoring while keeping the identical model and calibration split; if near-pole or near-gimbal-lock slice coverage does not rise substantially toward the nominal 90 percent target while marginal coverage stays controlled, the geometric claim is falsified.","tokens_in":16381,"feed_emoji":"👁️","tokens_out":673,"duration_ms":6202,"temperature":0.7,"pith_summary":"Conformal prediction is supposed to give distribution-free reliability guarantees for vision systems, but those guarantees depend on how error is measured. For gaze on the sphere and head pose in SO(3), many pipelines still score residuals in yaw–pitch or Euler charts. Those charts compress distances near poles and gimbal lock, so a single global threshold carves out vanishingly small intrinsic regions there. The paper shows that slice-conditional coverage at a nominal 90 percent target collapses by 30–50 percentage points in those regions across four standard datasets, while marginal coverage stays near target. The failure is structural: scalar adaptive methods can only resize the set, not fix its distorted shape. Switching the nonconformity score to a coordinate-free geodesic (or a monotone equivalent) removes the geometric distortion without retraining and at negligible cost. The practical message is that reliability for manifold-valued outputs is not only a calibration problem; it is also a geometry-of-the-score problem.","feed_headline":"Chart scores silently break gaze and pose coverage near poles","feed_subtitle":"Nominal 90% guarantees drop 30–50 points at singularities; geodesic scoring fixes it without retraining.","key_machinery":"Proposition 2: any acceptance set defined by scalar thresholding of a chart-coordinate residual norm is the pre-image of a chart-space ball; its local axis ratios are fixed by the eigenvalues of the metric tensor and therefore cannot be corrected by any scalar radius adaptation (normalised CP, CQR, etc.). The Riemannian volume density supplies a simple diagnostic that tracks where the collapse occurs.","core_discovery":"When the conformal nonconformity score is the Euclidean norm of chart residuals (yaw–pitch L2 or Euler L2), the resulting prediction sets inherit the chart’s metric distortion. Near coordinate singularities the same fixed threshold therefore covers far less manifold volume than it does in well-conditioned regions, redistributing coverage so that poles and near-gimbal-lock poses systematically undercover even though marginal coverage remains correctly calibrated. Geodesic scoring restores intrinsic isotropy and recovers the missing coverage.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Chart residuals collapse conformal coverage at gaze poles","Euler scores undercover near gimbal lock despite 90% marginal","Coordinate singularities silently break gaze and pose guarantees","Geodesic scoring restores coverage that chart thresholds lose","Riemannian density predicts where conformal gaze sets fail"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The first-order linearisation of the chart map and the local tangent-space error model remain qualitatively predictive at the finite radii and extreme angles actually present in the real datasets.","fun_headline_variants_meta":{"raw":{"variants":["Chart residuals collapse conformal coverage at gaze poles","Euler scores undercover near gimbal lock despite 90% marginal","Coordinate singularities silently break gaze and pose guarantees","Geodesic scoring restores coverage that chart thresholds lose","Riemannian density predicts where conformal gaze sets fail"]},"model":"grok-4.5","effort":"low","cost_usd":0.00437,"raw_usage":{"total_tokens":1363,"prompt_tokens":856,"num_sources_used":0,"completion_tokens":75,"cost_in_usd_ticks":43700000,"prompt_tokens_details":{"text_tokens":856,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":432,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":856,"tokens_out":75,"duration_ms":4099,"temperature":1.0,"reasoning_tokens":432,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T10:36:20.948345+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On any of the four datasets, replace the chart-norm score with geodesic scoring while keeping the identical model and calibration split; if near-pole or near-gimbal-lock slice coverage does not rise substantially toward the nominal 90 percent target while marginal coverage stays controlled, the geometric claim is falsified.","supporting_citations":[],"review_version":1}