REVIEW 7 minor 46 references
Chart-based conformal scores for gaze and head pose systematically undercover near singularities even when overall coverage looks correct.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 10:36 UTC pith:RHLXAJT3
load-bearing objection Chart-based conformal scores silently undercover near singularities by 30–50 pp; the impossibility result and controlled experiment make the claim solid, and geodesic scoring is a free fix.
Coordinate Singularities Break Conformal Coverage for Gaze and Head Pose
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When the conformal nonconformity score is the Euclidean norm of chart residuals (yaw–pitch L2 or Euler L2), the resulting prediction sets inherit the chart’s metric distortion. Near coordinate singularities the same fixed threshold therefore covers far less manifold volume than it does in well-conditioned regions, redistributing coverage so that poles and near-gimbal-lock poses systematically undercover even though marginal coverage remains correctly calibrated. Geodesic scoring restores intrinsic isotropy and recovers the missing coverage.
What carries the argument
Proposition 2: any acceptance set defined by scalar thresholding of a chart-coordinate residual norm is the pre-image of a chart-space ball; its local axis ratios are fixed by the eigenvalues of the metric tensor and therefore cannot be corrected by any scalar radius adaptation (normalised CP, CQR, etc.). The Riemannian volume density supplies a simple diagnostic that tracks where the collapse occurs.
Load-bearing premise
The first-order linearisation of the chart map and the local tangent-space error model remain qualitatively predictive at the finite radii and extreme angles actually present in the real datasets.
What would settle it
On any of the four datasets, replace the chart-norm score with geodesic scoring while keeping the identical model and calibration split; if near-pole or near-gimbal-lock slice coverage does not rise substantially toward the nominal 90 percent target while marginal coverage stays controlled, the geometric claim is falsified.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper shows that conformal prediction for manifold-valued vision outputs (gaze on S², head pose on SO(3)) can suffer large slice-conditional undercoverage near coordinate singularities when nonconformity scores are defined in chart coordinates (yaw–pitch or Euler L2), even when marginal coverage is correctly calibrated near 90%. Across ETH-XGaze, Gaze360, BIWI, and AFLW2000-3D, coverage drops by 30–50 pp near poles and gimbal lock. The authors prove (Proposition 2) that scalar adaptive methods (normalised CP, CQR-style threshold modulation) only rescale chart-space balls and cannot change axis ratios fixed by the metric-tensor eigenvalues. They propose the Riemannian volume density as a diagnostic and show that coordinate-free geodesic (or monotone-equivalent) scoring removes the chart-induced distortion without retraining and with negligible cost. A controlled SO(3) experiment with isotropic Lie-algebra noise isolates score geometry, and a three-layer decomposition separates chart distortion from heteroscedastic model error.
Significance. If the claims hold—and the evidence is strong—this is a practically important and theoretically clean contribution for reliability in geometric vision. Conformal prediction is increasingly used for distribution-free guarantees; the paper identifies a failure mode that is invisible under marginal metrics yet concentrated in safety-relevant regimes (extreme pitch, profile poses). Proposition 2 is an elementary but load-bearing geometric identity that cleanly rules out a natural class of remedies. The controlled BIWI experiment isolates score geometry from model error; the four-dataset audit, backbone ablations, and multi-chart/Mahalanobis comparisons make the empirical case robust. The proposed fix is immediately actionable (no retraining, ≤0.02 µs/sample). Credit is due for the clean impossibility result, the geometry-isolating controlled experiment, the three-layer decomposition, and the practical scoring protocol.
minor comments (7)
- In Sec. 4.3 the text reports normalised Euler coverage of 62.8±4.0% on AFLW2000-3D, but this number does not appear in Table 5. Adding it (or a short note) would make the Layer-2 vs Layer-3 comparison fully self-contained in the main table.
- Sec. 3.2 (“Validity of the linearisation”) already notes that Prop. 1’s quantitative bounds can be loose at extreme angles. A single clarifying sentence earlier in Sec. 3.2 stating that Prop. 1 supplies local intuition and first-order predictions, while the structural claim rests on Prop. 2 and the finite-radius experiments, would help readers weight the two results correctly.
- Fig. 3 is very effective; the shaded “degraded / severe / collapsed” bands are useful. Consider stating the exact coverage thresholds used for those bands in the caption so the figure is fully self-contained.
- Sec. 5 mentions GazeTR-ViT near-pole numbers (30.3% YP, 81.3% norm. geodesic) that support the “stronger models amplify distortion” claim. A one-row summary in the main text (or a small table) would avoid forcing readers into the supplement for a headline ablation.
- Notation: ρ(ξ) is introduced as volume density and later used as a correlation diagnostic. A brief reminder that for the standard charts ρ reduces to |cos(·)| (already stated) could be repeated once near Tables 2–5 where the ρ–coverage correlations are reported.
- Minor typography: “T able” appears with a space in several table captions (e.g., “T able 1”, “T able 2”); fix to “Table”. Also “F unctions” in the Sec. 3.4 heading.
- Sec. 6’s practical protocol is clear. A short decision note on when to prefer plain geodesic vs normalised geodesic vs Mondrian (once the base score is intrinsic) would help practitioners operationalise the three-layer decomposition.
Circularity Check
No significant circularity: geometric identities and held-out empirical measurements stand independently of any fitted inputs or self-referential definitions.
full rationale
The paper's load-bearing claims do not reduce to their inputs by construction. Proposition 2 is an elementary geometric identity: any acceptance set that is a sublevel set of a monotone function of the chart residual norm is the preimage of a chart ball, whose pullback under the chart differential is an ellipsoid whose axis ratios equal sqrt(lambda_max/lambda_min) of the metric tensor and are therefore independent of the scalar threshold or normalisation function h. This identity does not depend on data, fitted parameters, or linearisation. Proposition 1 supplies only local first-order bounds under a tangent-space error model; the authors themselves flag that the quantitative bounds loosen when the metric varies rapidly, yet the qualitative directional collapse is confirmed by finite-radius experiments. The controlled BIWI experiment injects isotropic Lie-algebra noise by construction, so any coverage variation is attributable solely to score geometry; the real-model tables (ETH-XGaze, Gaze360, AFLW2000-3D) measure slice-conditional coverage on held-out subject-disjoint splits after ordinary split conformal calibration. The volume-density diagnostic is a closed-form geometric quantity (rho = |cos theta| or |cos beta|) whose correlation with coverage is an observed statistic, not a fitted prediction. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled via self-citation, and no parameter fitted on one slice is re-presented as a prediction of the same slice. The three-layer decomposition cleanly separates marginal validity, heteroscedastic scale, and chart-induced shape; the persistent gap between normalised chart scores and normalised geodesic scores is therefore an independent empirical confirmation of the geometric claim rather than a circular restatement. The derivation chain is self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
free parameters (2)
- slice thresholds (|pitch|>70°, |β|>60°)
- k-NN bandwidth for normalised scores
axioms (4)
- domain assumption Exchangeability of calibration and test points (standard split conformal assumption)
- domain assumption Local error model Y = exp_p(ε) with ε ~ N(0, Σ_p) small enough to remain inside the chart domain
- standard math Metric tensors of the standard yaw–pitch chart on S² and ZYX Euler chart on SO(3)
- ad hoc to paper Acceptance sets of the form h(∥φ(y)−φ(ŷ)∥₂, x) ≤ q with h non-decreasing
read the original abstract
Conformal prediction provides distribution-free reliability guarantees for vision systems, but these guarantees depend on how prediction errors are measured in the output space. Many vision tasks produce outputs on curved spaces (e.g. gaze directions on the sphere or 3D head rotations), yet intermediate prediction heads, residuals, uncertainty estimates, or conformal scores are often defined in flat coordinate charts such as yaw-pitch or Euler angles. We show that this scoring choice introduces systematic geometric distortion near coordinate singularities (large pitch angles on the sphere and poses approaching gimbal lock in 3D rotations). Across four datasets (ETH-XGaze, Gaze360, BIWI, AFLW2000-3D), slice-conditional coverage at a nominal 90% target drops by 30-50 percentage points in these regions, falling to 38.9% on ETH-XGaze and 42.0% on Gaze360 at gaze pitch above 70 degrees, and to 57.5% on BIWI and 55.2% on AFLW2000-3D at head pose pitch above 60 degrees near gimbal lock, despite marginal coverage remaining near 90%. We prove that this is structural. Scalar thresholding changes the size of chart-coordinate prediction sets but leaves their distorted axis ratios unchanged. To diagnose this hidden failure mode, we show that a simple geometric quantity, the Riemannian volume density, strongly correlates with where coverage collapse occurs. Finally, we show that coordinate-free geodesic scoring removes this distortion. It requires no retraining and adds negligible computational cost.
Figures
Reference graph
Works this paper leans on
-
[1]
In: 2023 8th International Conference on Frontiers of Signal Processing (ICFSP)
Abdelrahman, A.A., Hempel, T., Khalifa, A., Al-Hamadi, A., Dinges, L.: L2cs- net: Fine-grained gaze estimation in unconstrained environments. In: 2023 8th International Conference on Frontiers of Signal Processing (ICFSP). pp. 98–102. IEEE (2023)
2023
-
[2]
In: Pro- ceedings of the IEEE/CVF Winter Conference on applications of computer vision
Cantarini, G., Tomenotti, F.F., Noceti, N., Odone, F.: Hhp-net: A light het- eroscedastic neural network for head pose estimation with uncertainty. In: Pro- ceedings of the IEEE/CVF Winter Conference on applications of computer vision. pp. 3521–3530 (2022)
2022
-
[3]
In: 2022 26th International Conference on Pattern Recognition (ICPR)
Cheng, Y., Lu, F.: Gaze estimation using transformer. In: 2022 26th International Conference on Pattern Recognition (ICPR). pp. 3341–3347. IEEE (2022)
2022
-
[4]
Electronic Journal of Statistics19(2), 6141–6166 (2025)
Cholaquidis, A., Gamboa, F., Moreno, L.: Conformal inference for regression on riemannian manifolds. Electronic Journal of Statistics19(2), 6141–6166 (2025)
2025
-
[5]
Do Carmo, M.P., Flaherty Francis, J.: Riemannian geometry, vol. 393. Springer (1992)
1992
-
[6]
International journal of computer vision101(3), 437–458 (2013)
Fanelli, G., Dantone, M., Gall, J., Fossati, A., Van Gool, L.: Random forests for real time 3d face analysis. International journal of computer vision101(3), 437–458 (2013)
2013
-
[7]
Clinical Biomechanics14(7), 462–470 (1999)
Feipel, V., Rondelet, B., Le Pallec, J.P., Rooze, M.: Normal global motion of the cervical spine:: an electrogoniometric study. Clinical Biomechanics14(7), 462–470 (1999)
1999
-
[8]
Bernoulli29(1), 1–23 (2023)
Fontana, M., Zeni, G., Vantini, S.: Conformal prediction: a unified review of theory and new challenges. Bernoulli29(1), 1–23 (2023)
2023
-
[9]
Information and Inference: A Journal of the IMA10(2), 455–482 (2021)
Foygel Barber, R., Candes, E.J., Ramdas, A., Tibshirani, R.J.: The limits of distribution-free conditional predictive inference. Information and Inference: A Journal of the IMA10(2), 455–482 (2021)
2021
-
[10]
IET Computer Vision10(4), 308– 314 (2016)
Fridman, L., Lee, J., Reimer, B., Victor, T.: ‘owl’and ‘lizard’: Patterns of head pose and eye pose in driver gaze classification. IET Computer Vision10(4), 308– 314 (2016)
2016
-
[11]
In: international conference on machine learn- ing
Gal, Y., Ghahramani, Z.: Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In: international conference on machine learn- ing. pp. 1050–1059. PMLR (2016)
2016
-
[12]
In: Advances in Neural Information Processing Systems (NeurIPS) (2021)
Gibbs, I., Candès, E.: Adaptive conformal inference under distribution shift. In: Advances in Neural Information Processing Systems (NeurIPS) (2021)
2021
-
[13]
Cambridge university press (2003)
Hartley, R., Zisserman, A.: Multiple view geometry in computer vision. Cambridge university press (2003)
2003
-
[14]
In: 2022 IEEE International Conference on image processing (ICIP)
Hempel, T., Abdelrahman, A.A., Al-Hamadi, A.: 6d rotation representation for unconstrained head pose estimation. In: 2022 IEEE International Conference on image processing (ICIP). pp. 2496–2500. IEEE (2022)
2022
-
[15]
Journal of Math- ematical Imaging and Vision35(2), 155–164 (2009)
Huynh, D.Q.: Metrics for 3d rotations: Comparison and analysis. Journal of Math- ematical Imaging and Vision35(2), 155–164 (2009)
2009
-
[16]
In: Conformal and Probabilistic Prediction and Applications
Johnstone, C., Cox, B.: Conformal uncertainty sets for robust optimization. In: Conformal and Probabilistic Prediction and Applications. pp. 72–90. PMLR (2021)
2021
-
[17]
In: Proceedings of the IEEE/CVF international conference on computer vision
Kellnhofer, P., Recasens, A., Stent, S., Matusik, W., Torralba, A.: Gaze360: Physi- cally unconstrained gaze estimation in the wild. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 6912–6921 (2019)
2019
-
[18]
(eds.) Advances Coordinate Singularities Break Conformal Coverage 17 in Neural Information Processing Systems
Kendall, A., Gal, Y.: What uncertainties do we need in bayesian deep learning for computer vision? In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances Coordinate Singularities Break Conformal Coverage 17 in Neural Information Processing Systems. vol. 30. Curran Associates, Inc. (2017),https : / ...
2017
-
[19]
In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R
Lakshminarayanan, B., Pritzel, A., Blundell, C.: Simple and scalable predic- tive uncertainty estimation using deep ensembles. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Ad- vances in Neural Information Processing Systems. vol. 30. Curran Associates, Inc. (2017),https://proceedings.neurips.cc/pa...
2017
-
[20]
Lee, J.M.: Introduction to Riemannian manifolds, vol. 2. Springer (2018)
2018
-
[21]
Eye33(7), 1192 (2019)
Lee, W.J., Kim, J.H., Shin, Y.U., Hwang, S., Lim, H.W.: Correction: Differences in eye movement range based on age and gaze direction. Eye33(7), 1192 (2019)
2019
-
[22]
Journal of the American Statistical Association 113(523), 1094–1111 (2018)
Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R.J., Wasserman, L.: Distribution-free predictive inference for regression. Journal of the American Statistical Association 113(523), 1094–1111 (2018)
2018
-
[23]
ACM Computing Surveys 56(2), 1–38 (2023)
Lei, Y., He, S., Khamis, M., Ye, J.: An end-to-end review of gaze estimation and its interactive applications on handheld mobile devices. ACM Computing Surveys 56(2), 1–38 (2023)
2023
-
[24]
arXiv preprint arXiv:2502.10570 (2025)
Lei, Y., Wang, Y., Buchanan, F., Zhao, M., Sugano, Y., He, S., Khamis, M., Ye, J.: Quantifying the impact of motion on 2d gaze estimation in real-world mobile interactions. arXiv preprint arXiv:2502.10570 (2025)
Pith/arXiv arXiv 2025
-
[25]
arXiv preprint arXiv:2505.22769 (2025)
Lei, Y., Zhao, M., Wang, Y., He, S., Sugano, Y., Khamis, M., Ye, J.: Mac- gaze: Motion-aware continual calibration for mobile gaze tracking. arXiv preprint arXiv:2505.22769 (2025)
Pith/arXiv arXiv 2025
-
[26]
Ma, Y., Soatto, S., Kosecka, J., Sastry, S.S.: An invitation to 3-d vision: from images to geometric models, vol. 26. Springer Science & Business Media (2012)
2012
-
[27]
In: Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition
Martyniuk, T., Kupyn, O., Kurlyak, Y., Krashenyi, I., Matas, J., Sharmanska, V.: Dad-3dheads: A large-scale dense, accurate and diverse dataset for 3d head alignment from a single image. In: Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition. pp. 20942–20952 (2022)
2022
-
[28]
In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H
Mohlin, D., Sullivan, J., Bianchi, G.: Probabilistic orientation estimation with ma- trix fisher distributions. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems. vol. 33, pp. 4884–4893. Curran Associates, Inc. (2020),https://proceedings.neurips.cc/ paper _ files / paper / 2020 / fil...
2020
-
[29]
In: NeuRIPS 2022 Workshop on Gaze Meets ML (2022)
Nikan, S., Upadhyay, D.: Appearance-based gaze estimation for driver monitoring. In: NeuRIPS 2022 Workshop on Gaze Meets ML (2022)
2022
-
[30]
In: European conference on machine learning
Papadopoulos, H., Proedrou, K., Vovk, V., Gammerman, A.: Inductive confidence machines for regression. In: European conference on machine learning. pp. 345–356. Springer (2002)
2002
-
[31]
ACM Transactions On Graphics (TOG)35(6), 1–12 (2016)
Patney, A., Salvi, M., Kim, J., Kaplanyan, A., Wyman, C., Benty, N., Luebke, D., Lefohn, A.: Towards foveated rendering for gaze-tracked virtual reality. ACM Transactions On Graphics (TOG)35(6), 1–12 (2016)
2016
-
[32]
In: Proceedings of the European conference on computer vision (ECCV)
Prokudin, S., Gehler, P., Nowozin, S.: Deep directional statistics: Pose estimation with uncertainty quantification. In: Proceedings of the European conference on computer vision (ECCV). pp. 534–551 (2018)
2018
-
[33]
In: Ad- vances in Neural Information Processing Systems (NeurIPS) (2019) 18 M
Romano, Y., Patterson, E., Candès, E.: Conformalized quantile regression. In: Ad- vances in Neural Information Processing Systems (NeurIPS) (2019) 18 M. Jamalifard et al
2019
-
[34]
In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops
Ruiz, N., Chong, E., Rehg, J.M.: Fine-grained head pose estimation without key- points. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops. pp. 2074–2083 (2018)
2074
-
[35]
arXiv preprint arXiv: 2107.07511 (2021)
Stephen, B., et al.: A gentle introduction to conformal prediction and distribution- free uncertainty quantification. arXiv preprint arXiv: 2107.07511 (2021)
Pith/arXiv arXiv 2021
-
[36]
Advances in neural information processing systems32(2019)
Tibshirani, R.J., Foygel Barber, R., Candes, E., Ramdas, A.: Conformal prediction under covariate shift. Advances in neural information processing systems32(2019)
2019
-
[37]
Springer (2005)
Vovk, V., Gammerman, A., Shafer, G.: Algorithmic learning in a random world. Springer (2005)
2005
-
[38]
In: Proceedings of the 2025 Symposium on Eye Tracking Research and Applications
Wang, Y., Yan, R., Lei, Y., Fu, X.: Ptgaze: Cross-domain gaze estimation via proxy tuning. In: Proceedings of the 2025 Symposium on Eye Tracking Research and Applications. pp. 1–2 (2025)
2025
-
[39]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Yang, T.Y., Chen, Y.T., Lin, Y.Y., Chuang, Y.Y.: Fsa-net: Learning fine-grained structure aggregation for head pose estimation from a single image. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 1087–1096 (2019)
2019
-
[40]
In: European conference on computer vision
Zhang, X., Park, S., Beeler, T., Bradley, D., Tang, S., Hilliges, O.: Eth-xgaze: A large scale dataset for gaze estimation under extreme head pose and gaze variation. In: European conference on computer vision. pp. 365–381. Springer (2020)
2020
-
[41]
In: 31st British Machine Vision Conference (BMVC 2020)
Zhang, X., Sugano, Y., Bulling, A., Hilliges, O.: Learning-based region selection for end-to-end gaze estimation. In: 31st British Machine Vision Conference (BMVC 2020). p. 86. British Machine Vision Association (2020)
2020
-
[42]
arXiv preprint arXiv:2303.10062 (2023)
Zheng, Q., Zhang, X.: Confidence-aware 3d gaze estimation and evaluation metric. arXiv preprint arXiv:2303.10062 (2023)
Pith/arXiv arXiv 2023
-
[43]
IEEE Transactions on Image Processing33, 2851–2866 (2024)
Zhong, W., Xia, C., Zhang, D., Han, J.: Uncertainty modeling for gaze estimation. IEEE Transactions on Image Processing33, 2851–2866 (2024)
2024
-
[44]
ACM computing surveys58(2), 1–37 (2025)
Zhou, X., Chen, B., Gui, Y., Cheng, L.: Conformal prediction: A data perspective. ACM computing surveys58(2), 1–37 (2025)
2025
-
[45]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Zhou, Y., Barnes, C., Lu, J., Yang, J., Li, H.: On the continuity of rotation rep- resentations in neural networks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 5745–5753 (2019)
2019
-
[46]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Zhu, X., Lei, Z., Liu, X., Shi, H., Li, S.Z.: Face alignment across large poses: A 3d solution. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 146–155 (2016)
2016
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.