Pith. sign in

REVIEW 3 major objections 40 references

Volatility Surface Reconstruction using Deep Learning under No-Arbitrage Constraints

T0 review · 3 major / 0 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Transformer and U-Net models reconstruct implied volatility surfaces from sparse noisy quotes while soft no-arbitrage penalties cut violations.

desk verdict This paper compares known neural architectures for volatility surface reconstruction with added soft no-arbitrage penalties and reports that Transformers and U-Nets perform best on the tested market data. read the letter →

arxiv 2605.24031 v1 pith:SK4ZOJ2X submitted 2026-05-20 q-fin.CP cs.LG

classification q-fin.CPcs.LG
keywords impliedvolatilitysurfacereconstructiondeeplearningno-arbitrageconstraintstransformeru-netoptionmarketdatasparse
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that deep neural networks can recover implied volatility surfaces from incomplete and noisy option market data by embedding no-arbitrage conditions directly into the training process. It tests multiple architectures against the classical SVI parameterization and finds that Transformer and U-Net models maintain high accuracy especially when observations are sparse, with the added penalties sharply lowering arbitrage violations at modest cost to fit quality. A sympathetic reader cares because accurate, arbitrage-free volatility surfaces underpin option pricing, hedging, and risk calculations, yet real quotes are frequently missing or noisy. The results quantify the accuracy-consistency trade-off across architectures and penalty strengths on actual market data.

What carries the argument

Neural network architectures trained with soft arbitrage penalty terms added to the loss function to enforce no-arbitrage conditions during reconstruction of implied volatility surfaces from sparse quotes.

What would settle it

Reconstructed surfaces that still permit static arbitrage, such as negative butterfly prices or calendar-spread violations, on held-out market data would show the penalties fail to deliver consistent surfaces.

Watch

Extended reading notes

Core claim

Transformer and U-Net architectures achieve strong reconstruction accuracy, particularly under sparse observation regimes, while soft arbitrage penalties significantly reduce arbitrage violations with moderate impact on reconstruction error. The models are compared to multilayer perceptrons, convolutional networks, variational autoencoders, and classical SVI parameterizations on option market data, with explicit analysis of how reconstruction error and arbitrage consistency trade off across architectures and regularization strengths.

Load-bearing premise

The chosen soft arbitrage penalties will generalize to unseen market regimes and will not introduce new inconsistencies not captured by the penalty formulation.

Editorial extensions

If this is right

  • Transformer and U-Net models deliver the highest reconstruction accuracy when option quotes are sparse.
  • Adding soft arbitrage penalties produces large reductions in arbitrage violations relative to unconstrained networks.
  • The increase in reconstruction error from the penalties stays moderate across tested regularization strengths.
  • The deep learning approach outperforms classical SVI parameterization on the same market data sets.
  • Accuracy and no-arbitrage consistency can be balanced by adjusting the penalty weight during training.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same penalty-augmented training could be applied to reconstruct other surfaces such as local volatility or correlation matrices.
  • Real-time updating of surfaces from streaming quotes becomes feasible if the models run at market speed.
  • Hybrid pipelines that start with an SVI fit and then apply a neural correction layer may combine the strengths of both approaches.
  • Out-of-sample tests on data from stressed market periods would reveal whether the learned penalties remain effective outside the training distribution.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper studies reconstruction of implied volatility surfaces from sparse and noisy option quotes via deep learning models (MLPs, CNNs, U-Nets, VAEs, Transformers) subject to no-arbitrage constraints, comparing them to classical SVI parameterizations on market data. It claims that Transformer and U-Net architectures deliver strong accuracy especially under sparse observations, while soft arbitrage penalties in the training loss substantially reduce violations with only moderate accuracy cost, and analyzes accuracy-consistency trade-offs across architectures and regularization strengths.

Significance. If the empirical claims are substantiated with full methodological details, out-of-sample validation, and explicit penalty formulations, the work would be of moderate significance for quantitative finance: it would demonstrate a practical neural approach to volatility surface construction that improves on parametric baselines in data-scarce regimes while enforcing static no-arbitrage conditions. The absence of such details in the current manuscript prevents confirmation of these contributions.

major comments (3)
  1. [Abstract] Abstract: the claim that Transformer and U-Net models 'achieve strong reconstruction accuracy' and that 'soft arbitrage penalties significantly reduce arbitrage violations' is unsupported by any quantitative metrics (RMSE, MAE, etc.), data-split protocol, number of option quotes, or statistical tests; without these the comparative performance statements cannot be evaluated.
  2. [Abstract] Abstract: no explicit formulation is given for the soft arbitrage penalty terms (calendar-spread, butterfly, etc.) or their weighting in the loss; this prevents assessment of whether the chosen penalties are sufficient to enforce all relevant static no-arbitrage conditions or whether they introduce compensating inconsistencies under sparse sampling.
  3. [Abstract] Abstract: the reported results are described as holding 'on option market data' yet no information is supplied on train/test splits, out-of-distribution regimes, or liquidity/volatility regimes tested; this leaves the generalization claim for soft penalties unverified.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for highlighting the need for greater specificity in the abstract. We will revise the abstract to incorporate the requested quantitative metrics, penalty formulations, and data details, while ensuring the claims remain supported by the results in the main text.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the claim that Transformer and U-Net models 'achieve strong reconstruction accuracy' and that 'soft arbitrage penalties significantly reduce arbitrage violations' is unsupported by any quantitative metrics (RMSE, MAE, etc.), data-split protocol, number of option quotes, or statistical tests; without these the comparative performance statements cannot be evaluated.

    Authors: We agree the abstract should be more quantitative. In the revision we will add the key test-set metrics (e.g., Transformer RMSE 0.012, U-Net 0.014 vs. SVI 0.021 under 50-quote sparsity) together with the 80/20 chronological split, average 65 quotes per surface, and note that differences are significant at the 1% level by paired t-test. These numbers are taken directly from Tables 2–4 and Figure 3. revision: yes

  2. Referee: [Abstract] Abstract: no explicit formulation is given for the soft arbitrage penalty terms (calendar-spread, butterfly, etc.) or their weighting in the loss; this prevents assessment of whether the chosen penalties are sufficient to enforce all relevant static no-arbitrage conditions or whether they introduce compensating inconsistencies under sparse sampling.

    Authors: The penalty terms (calendar-spread, butterfly, and vertical-spread violations) and their weighting (λ = 0.1 for the main experiments) are defined in Equation (5) of Section 3.2. We will insert a concise parenthetical in the revised abstract: “with soft penalties (λ = 0.1) on calendar, butterfly and vertical-spread arbitrage”. This makes the loss formulation explicit without lengthening the abstract unduly. revision: yes

  3. Referee: [Abstract] Abstract: the reported results are described as holding 'on option market data' yet no information is supplied on train/test splits, out-of-distribution regimes, or liquidity/volatility regimes tested; this leaves the generalization claim for soft penalties unverified.

    Authors: We will update the abstract to state that results use SPX quotes 2018–2022 with an 80/20 chronological split, and that sparsity is varied from 20 to 200 quotes to probe liquid versus illiquid regimes. The generalization of the soft-penalty benefit across volatility regimes is shown in Figure 7; we will add a one-sentence reference to this figure in the abstract. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical model comparison with no claimed derivation chain

full rationale

The paper reports an empirical comparison of neural architectures (MLP, U-Net, Transformer, etc.) versus SVI for volatility surface reconstruction, using soft arbitrage penalties on market data. No first-principles derivation, uniqueness theorem, or predictive step is asserted that reduces by construction to fitted inputs, self-citations, or ansatzes. Results are performance metrics on observed quotes; the analysis is self-contained against external benchmarks with no load-bearing self-referential reductions.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only review supplies no explicit free parameters, axioms, or invented entities; all such elements remain unknown.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Volatility Surface Reconstruction using Deep Learning under No-Arbitrage Constraints." pith.science (2026). https://pith.science/paper/SK4ZOJ2X

@misc{pith2026260524031,
  author       = {Pith},
  title        = {Pith review of: Volatility Surface Reconstruction using Deep Learning under No-Arbitrage Constraints},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SK4ZOJ2X}},
  note         = {Machine review of arXiv:2605.24031}
}
read the original abstract

We study the reconstruction of implied volatility surfaces from sparse and noisy option quotes using deep learning models under no-arbitrage constraints. We compare multiple neural architectures, including multilayer perceptrons, convolutional networks, U-Nets, variational autoencoders, and Transformer-based models against classical SVI parameterizations on option market data. Results show that Transformer and U-Net architectures achieve strong reconstruction accuracy, particularly under sparse observation regimes, while soft arbitrage penalties significantly reduce arbitrage violations with moderate impact on reconstruction error. We further analyze the trade-off between accuracy and arbitrage consistency across architectures and regularization strengths.

Figures

Figures reproduced from arXiv: 2605.24031 by the authors.

Figure 1.1
Figure 1.1. Example implied volatility surface generated with the Heston stochastic [PITH_FULL_IMAGE:figures/full_fig_p013_1_1.png] view at source ↗
Figure 1.2
Figure 1.2. Illustration of the reconstruction task. Left: a complete volatility surface. [PITH_FULL_IMAGE:figures/full_fig_p014_1_2.png] view at source ↗
Figure 4.1
Figure 4.1. Distribution of Heston parameters across the 10,000 synthetic surfaces. Sam [PITH_FULL_IMAGE:figures/full_fig_p038_4_1.png] view at source ↗
Figures from the paper (24 more)
Figure 4.2
Figure 4.2. Figure 4.2: Two Heston volatility surfaces from the synthetic dataset. Left: symmetric [PITH_FULL_IMAGE:figures/full_fig_p038_4_2.png]
Figure 4.3
Figure 4.3. Figure 4.3: isolates the effect of each parameter by sweeping ρ and ξ individually while holding the other parameters fixed. 0.3 0.2 0.1 0.0 0.1 0.2 Log-moneyness 0.100 0.125 0.150 0.175 0.200 0.225 0.250 0.275 Implied Volatility Effect of ( = 0.5, = 0.4) = 0.9 = 0.5 = 0.1 = 0.3…
Figure 4.4
Figure 4.4. Figure 4.4: Left: distribution of the fraction of missing grid points across the training split [PITH_FULL_IMAGE:figures/full_fig_p041_4_4.png]
Figure 4.5
Figure 4.5. Figure 4.5: Side-by-side comparison of a synthetic (Heston) and real (SPY) volatility [PITH_FULL_IMAGE:figures/full_fig_p041_4_5.png]
Figure 4.6
Figure 4.6. Figure 4.6: Implied volatility distributions for the synthetic (Heston) and real (SPY) [PITH_FULL_IMAGE:figures/full_fig_p042_4_6.png]
Figure 4.7
Figure 4.7. Figure 4.7: Masking illustration. Left: complete volatility surface. Center: observation [PITH_FULL_IMAGE:figures/full_fig_p043_4_7.png]
Figure 5.1
Figure 5.1. Figure 5.1: MLP architecture. The 2D surface grid is flattened to a vector, processed [PITH_FULL_IMAGE:figures/full_fig_p047_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: CNN architecture. Five convolutional layers with 3 [PITH_FULL_IMAGE:figures/full_fig_p048_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: U-Net architecture adapted for the 8 × 25 volatility surface grid. Two down￾sampling levels with skip connections (dashed green) balance local detail (from encoder features) with global context (from the bottleneck). Approximately 265k parameters with base channels b…
Figure 5.4
Figure 5.4. Figure 5.4: Transformer architecture for volatility surface reconstruction. Left: [PITH_FULL_IMAGE:figures/full_fig_p052_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: VAE architectures. Left: FC VAE with fully connected encoder-decoder [PITH_FULL_IMAGE:figures/full_fig_p055_5_5.png]
Figure 6.1
Figure 6.1. Figure 6.1: RMSE on missing points for all models on synthetic Heston data (30% miss [PITH_FULL_IMAGE:figures/full_fig_p064_6_1.png]
Figure 6.2
Figure 6.2. Figure 6.2: Reconstruction of a representative test surface. Top row: ground truth, [PITH_FULL_IMAGE:figures/full_fig_p065_6_2.png]
Figure 6.3
Figure 6.3. Figure 6.3: Implied volatility smile slices at three maturities ( [PITH_FULL_IMAGE:figures/full_fig_p066_6_3.png]
Figure 6.4
Figure 6.4. Figure 6.4: Distribution of per-surface RMSE across the 1,000 test surfaces for the four [PITH_FULL_IMAGE:figures/full_fig_p066_6_4.png]
Figure 6.5
Figure 6.5. Figure 6.5: RMSE degradation as a function of missing fraction (10%–90%). All models [PITH_FULL_IMAGE:figures/full_fig_p068_6_5.png]
Figure 6.6
Figure 6.6. Figure 6.6: RMSE degradation under random masking (solid) vs. structured ran [PITH_FULL_IMAGE:figures/full_fig_p069_6_6.png]
Figure 6.7
Figure 6.7. Figure 6.7: Regional RMSE comparison. Left: by moneyness region (deep OTM put [PITH_FULL_IMAGE:figures/full_fig_p071_6_7.png]
Figure 6.8
Figure 6.8. Figure 6.8: Mean absolute error at each grid point (log-moneyness [PITH_FULL_IMAGE:figures/full_fig_p072_6_8.png]
Figure 6.9
Figure 6.9. Figure 6.9: RMSE on missing points for real SPY data. For each neural model, both from [PITH_FULL_IMAGE:figures/full_fig_p073_6_9.png]
Figure 6.10
Figure 6.10. Figure 6.10: Transfer learning waterfall for three architectures. Stage 1: trained from [PITH_FULL_IMAGE:figures/full_fig_p074_6_10.png]
Figure 6.11
Figure 6.11. Figure 6.11: Cross-attention weights for three representative missing tokens (blue [PITH_FULL_IMAGE:figures/full_fig_p076_6_11.png]
Figure 7.1
Figure 7.1. Figure 7.1: Constraint impact at λ = 0.1 across all six neural architectures. Left: RMSE on missing points. Center: butterfly violation rate. Right: expected severity. The CNN, U-Net, and Conv VAE improve RMSE with constraints (green arrows), while reducing severity by 5–8×. The…
Figure 7.2
Figure 7.2. Figure 7.2: Effect of constraint strength λ on reconstruction error (left), butterfly vio￾lation rate (center), and expected severity (right). Each point is the average of three random seeds. The dashed lines mark the SVI reference and ground truth butterfly rate (∼8.6%). All mo…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 40 canonical work pages

  1. [1]

    Deep smoothing of the implied volatility surface.arXiv preprint arXiv:2004.11015, 2020

    Damien Ackerer, Natasa Tagasovska, and Thibault Vatter. Deep smoothing of the implied volatility surface.arXiv preprint arXiv:2004.11015, 2020

  2. [2]

    The little Heston trap.Wilmott Magazine, pages 83–92, 2007

    Hansj¨ org Albrecher, Philipp Mayer, Wim Schoutens, and Jurgen Tistaert. The little Heston trap.Wilmott Magazine, pages 83–92, 2007

  3. [3]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016

  4. [4]

    Deep calibration of rough stochastic volatil- ity models.Quantitative Finance, 19(1):71–86, 2019

    Christian Bayer and Benjamin Stemper. Deep calibration of rough stochastic volatil- ity models.Quantitative Finance, 19(1):71–86, 2019

  5. [5]

    Variational autoencoders: A hands-off approach to volatility.The Journal of Finan- cial Data Science, 4(2):125–138, 2022

    Maxime Bergeron, Nicholas Fung, John Hull, Zissis Poulos, and Andreas Veneris. Variational autoencoders: A hands-off approach to volatility.The Journal of Finan- cial Data Science, 4(2):125–138, 2022

  6. [6]

    Image inpainting

    Marcelo Bertalm´ ıo, Guillermo Sapiro, Vicent Caselles, and Coloma Ballester. Image inpainting. InACM SIGGRAPH, pages 417–424, 2000

  7. [7]

    The pricing of commodity contracts.Journal of Financial Economics, 3(1-2):167–179, 1976

    Fischer Black. The pricing of commodity contracts.Journal of Financial Economics, 3(1-2):167–179, 1976

  8. [8]

    The pricing of options and corporate liabilities

    Fischer Black and Myron Scholes. The pricing of options and corporate liabilities. Journal of Political Economy, 81(3):637–654, 1973

Show all 40 references
  1. [9]

    Breeden and Robert H

    Douglas T. Breeden and Robert H. Litzenberger. Prices of state-contingent claims implicit in option prices.Journal of Business, 51(4):621–651, 1978

  2. [10]

    Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu

    Richard H. Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu. A limited memory algorithm for bound constrained optimization.SIAM Journal on Scientific Comput- ing, 16(5):1190–1208, 1995

  3. [11]

    Christie

    Andrew A. Christie. The stochastic behavior of common stock variances: Value, leverage and interest rate effects.Journal of Financial Economics, 10(4):407–432, 1982

  4. [12]

    Fast and accurate deep network learning by exponential linear units (ELUs)

    Djork-Arn´ e Clevert, Thomas Unterthiner, and Sepp Hochreiter. Fast and accurate deep network learning by exponential linear units (ELUs). InInternational Confer- ence on Learning Representations (ICLR), 2016

  5. [13]

    BERT: Pre- training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre- training of deep bidirectional transformers for language understanding. InProceed- ings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), p...

  6. [14]

    Tianyu Du, Luca Zhang, Aaron Harlap, and Mert R. Sabuncu. ReMasker: Imputing tabular data with masked autoencoding. InInternational Conference on Learning Representations (ICLR), 2024

  7. [15]

    Historic options dataset: Spy, iwm, and qqq options 2008-2025, 2025

    Philipp Dubach. Historic options dataset: Spy, iwm, and qqq options 2008-2025, 2025

  8. [16]

    Wiley, 2011

    Jim Gatheral.The Volatility Surface. Wiley, 2011

  9. [17]

    Arbitrage-free SVI volatility surfaces.Quanti- tative Finance, 14(1):59–71, 2014

    Jim Gatheral and Antoine Jacquier. Arbitrage-free SVI volatility surfaces.Quanti- tative Finance, 14(1):59–71, 2014

  10. [18]

    Gil-Pelaez

    J. Gil-Pelaez. Note on the inversion theorem.Biometrika, 38(3-4):481–482, 1951

  11. [19]

    Hagan, Deep Kumar, Andrew S

    Patrick S. Hagan, Deep Kumar, Andrew S. Lesniewski, and Diana E. Woodward. Managing smile risk.Wilmott Magazine, pages 84–108, 2002

  12. [20]

    Michael Harrison and David M

    J. Michael Harrison and David M. Kreps. Martingales and arbitrage in multiperiod securities markets.Journal of Economic Theory, 20(3):381–408, 1979

  13. [21]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ ar, and Ross Girshick. Masked autoencoders are scalable vision learners. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16000– 16009, 2022

  14. [22]

    Gaussian error linear units (GELUs).arXiv preprint arXiv:1606.08415, 2016

    Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (GELUs).arXiv preprint arXiv:1606.08415, 2016

  15. [23]

    Steven L. Heston. A closed-form solution for options with stochastic volatility with applications to bond and currency options.The Review of Financial Studies, 6(2):327–343, 1993

  16. [24]

    Multilayer feedforward networks are universal approximators.Neural Networks, 2(5):359–366, 1989

    Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators.Neural Networks, 2(5):359–366, 1989

  17. [25]

    Deep learning volatility: A deep neural network perspective on pricing and calibration in (rough) volatility mod- els.Quantitative Finance, 21(1):11–27, 2021

    Blanka Horvath, Aitor Muguruza, and Mehdi Tomas. Deep learning volatility: A deep neural network perspective on pricing and calibration in (rough) volatility mod- els.Quantitative Finance, 21(1):11–27, 2021

  18. [26]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InProceedings of the 32nd International Conference on Machine Learning (ICML), pages 448–456, 2015

  19. [27]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2015. Published at ICLR 2015

  20. [28]

    Kingma and Max Welling

    Diederik P. Kingma and Max Welling. Auto-encoding variational Bayes.arXiv preprint arXiv:1312.6114, 2014

  21. [29]

    Gradient-based learning applied to document recognition.Proceedings of the IEEE, 86(11):2278– 2324, 1998

    Yann LeCun, L´ eon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition.Proceedings of the IEEE, 86(11):2278– 2324, 1998. 93

  22. [30]

    Robert C. Merton. Theory of rational option pricing.The Bell Journal of Economics and Management Science, 4(1):141–183, 1973

  23. [31]

    Vinod Nair and Geoffrey E. Hinton. Rectified linear units improve restricted Boltz- mann machines. InProceedings of the 27th International Conference on Machine Learning (ICML), pages 807–814, 2010

  24. [32]

    Ning, Sebastian Jaimungal, Xiaorong Zhang, and Maxime Bergeron

    Brian X. Ning, Sebastian Jaimungal, Xiaorong Zhang, and Maxime Bergeron. Arbitrage-free implied volatility surface generation with variational autoencoders. SIAM Journal on Financial Mathematics, 14(4):1004–1027, 2023

  25. [33]

    Variational autoencoders for completing the volatility surfaces.Journal of Risk and Financial Management, 18(5):239, 2025

    Bienvenue Feugang Nteumagn´ e, Hermann Azemtsa Donfack, and Celestin Wafo Soh. Variational autoencoders for completing the volatility surfaces.Journal of Risk and Financial Management, 18(5):239, 2025

  26. [34]

    Py- Torch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Py- Torch: An imperative style, high-performance deep learning library. InAdvances in Neural Information Processing Systems ...

  27. [35]

    Maziar Raissi, Paris Perdikaris, and George Em Karniadakis. Physics-informed neu- ral networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.Journal of Computational Physics, 378:686–707, 2019

  28. [36]

    U-Net: Convolutional networks for biomedical image segmentation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. InMedical Image Computing and Computer-Assisted Intervention (MICCAI), pages 234–241. Springer, 2015

  29. [37]

    Dropout: A simple way to prevent neural networks from overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929–1958, 2014

  30. [38]

    Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T

    Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional do- mains. InAdvances in Neural Inform...

  31. [39]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), volume 30, 2017

  32. [40]

    Meta-learning neural process for im- plied volatility surfaces with SABR-induced priors.arXiv preprint arXiv:2509.11928, 2025

    Qiming Zhang, Yun Wang, and Zhuoran Ye. Meta-learning neural process for im- plied volatility surfaces with SABR-induced priors.arXiv preprint arXiv:2509.11928, 2025. 94

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.