Pith. sign in

REVIEW 3 major objections 4 minor 39 references

The shape of TabPFN's hidden geometry predicts, without test labels, when its confidence stops matching true probability.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 16:31 UTC pith:C26K56CQ

load-bearing objection A careful, honest empirical study that transfers known zigzag machinery to TabPFN, but the headline reliability correlations are not separated from the hand-built difficulty ladder, so the diagnostic claim overreaches. the 3 major comments →

arxiv 2607.17962 v1 pith:C26K56CQ submitted 2026-07-20 cs.LG cs.AIstat.ML

Topological Signatures of Context-Level Reliability in TabPFN

classification cs.LG cs.AIstat.ML MSC 55N3168T07
keywords TabPFNin-context learningzigzag persistencepersistent homologytopological data analysisrepresentation geometrymodel calibrationoverconfidence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that the geometry of TabPFN's internal representations — tracked across the model's 12 layers with zigzag persistence, a technique that records when clusters and loops are born and die as the point cloud moves through depth — carries a label-free signal of when the model's predictions can be trusted. On synthetic tabular tasks with known true probabilities (warped circles, tori, spheres, Hopf links, trefoil knots, Swiss rolls), the count of fragmented clusters in the hidden geometry correlates with the model's mean absolute residual (MAR) at ρ=0.92 and with Bayes error at ρ=0.94, while the total lifetime of those clusters correlates at ρ=−0.94. The authors describe the result as a 'scissors' pattern: as task difficulty rises, loop-like structure proliferates while stable cluster structure erodes, and both movements coincide with worse calibration and more overconfidence. If the claim is right, a whole-dataset reliability score can be computed from the model's own activations before ground-truth labels exist, and pretraining for tabular foundation models should diversify the topology of synthetic tasks rather than merely scale them up. The paper's unifying reading is that zigzag persistence diagnoses the reliability of the in-context task geometry TabPFN constructs.

Core claim

The claim: topology of TabPFN's hidden representations strongly tracks dataset-level reliability. TabPFN is a transformer that predicts from labeled and unlabeled rows in one pass; the authors run zigzag persistence over layerwise point clouds and extract topological descriptors. Harder geometries induce a dual signature: more loops, more H0 fragmentation (splintered clusters), shorter-lived durable structure. In a high-resolution warped-circle study (5,000 rows), H0 fragmentation correlates with mean absolute residual at ρ=0.92 and Bayes error at ρ=0.94, and H0 total persistence at ρ=−0.94. Their reading: these descriptors diagnose the reliability of the inferred in-context task geometry.

What carries the argument

Zigzag persistence: at each of TabPFN's 12 layers a k-nearest-neighbor graph (k=4) on the representation cloud is expanded into a simplicial complex, and intersections between consecutive layers form a 23-position zigzag filtration recording feature births and deaths across depth. H0 (connected components) tracks whether rows coalesce into long-lived clusters or splinter into fragments; H1 (loops) tracks entangled cyclic structure. Key descriptors: H0 fragmentation count, H0 total persistence, H1 area, durable-H1 persistence. The signature is the H0/H1 'scissors': difficulty raises loop activity while eroding durable cluster structure — more features, each living shorter.

Load-bearing premise

The load-bearing premise is that the topological descriptors carry reliability information beyond the shared difficulty ladder: because noise, nuisance dimensions, warp, and family-specific knobs were escalated together by hand, both the descriptors and the reliability metrics are monotone in input hardness, and the paper never partials difficulty out to show the topology–reliability link is not just 'harder input, worse everything.'

What would settle it

Partially correlate H0 fragmentation count with mean absolute residual across the 45 warped-circle runs, controlling for difficulty (σ, m, w, τ). If the association collapses to zero, the descriptors are pure difficulty proxies, adding nothing beyond flagging hard inputs. If a large association survives at fixed difficulty, the readout is genuine. Complementary check: two families matched on every difficulty knob but with clearly different mean absolute residual; the claim predicts their H0/H1 descriptors must differ too, so identical descriptors with divergent reliability would refute it.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A new tabular context can be pre-screened for reliability risk from TabPFN's own activations — before any ground-truth labels arrive — and the signal persists across query, support-feature, and support-label token scopes (ρ≈0.71–0.83), so it is not tied to one embedding surface.
  • Deployed monitoring: a context that drifts toward more fragmented H0, more active H1, or shorter-lived durable structure is drifting toward worse calibration and more overconfidence — a shift the descriptors would catch without labels.
  • Fine-tuning is not an obvious remedy: small-scale fine-tuning left observed error, Bayes error, and MAR statistically unchanged (p≈0.6–0.8) and slightly increased overconfidence at the hardest difficulty, suggesting the failure is bound to the pretrained prior, not to insufficient adaptation.
  • Pretraining strategy: diversifying the intrinsic topology of synthetic pretraining tasks — loops, links, knots, and curved but loop-free sheets — may matter more than scaling row counts for reliable transfer to unseen complex geometries.
  • The direction of the diagnostic is regime-dependent: five of six families show unreliability as added topological complexity, while the trefoil knot inverts the link (ρ=−0.47) into a collapse mode, so a deployer must know which regime a context is in.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A check the paper does not run: because noise, nuisance dimensions, warp, and family knobs escalate together, the descriptors may be pure proxies for input hardness; partialing difficulty level out of the ρ=0.92–0.94 warped-circle correlations would decide whether topology adds anything beyond flagging the input directly.
  • Deployment ambiguity: the trefoil's sign inversion means a monotone 'stressed topology = unreliable' rule cannot be applied blindly; a practical monitor would need a regime classifier or sign-calibration step first, which the paper leaves implicit.
  • Descriptor substitution: the paper's own saturation results suggest a production diagnostic should downweight raw H1 area at dense sampling and lean on H0 fragmentation plus durable-H1 persistence, the two descriptors the authors show have no raw ceiling.
  • External validity test: since the benchmark is fully synthetic with fixed label sharpness and coarse global descriptors — limitations the paper states — the natural next experiment is whether the same zigzag descriptors flag reliability shifts on real tables with categorical features, missingness, and heterogeneous types.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper uses zigzag persistent homology on TabPFN v2's 12 layerwise hidden representations to ask whether the topology of the internal representation geometry is associated with dataset-level reliability. Representations from each layer are converted to k-NN graphs and clique complexes, with intersection complexes inserted between layers; H0 and H1 descriptors (counts, total persistence, durable fractions, layer-timing histograms) are computed from the resulting zigzag intervals. The authors build six synthetic classification families with known true probabilities — warped circle, torus, sphere, Hopf link, trefoil knot, and Swiss roll — and vary difficulty through a shared ladder of noise, nuisance dimensions, warp, plus family-specific knobs. Reliability is measured via MAR, Bayes error, observed error, and overconfidence rates computed against the known generative labels. The central empirical claim is that H0 fragmentation and H1 activity correlate strongly with reliability, with a large-sample warped-circle study reporting H0 fragmentation vs MAR ρ=0.92 and vs Bayes error ρ=0.94, H0 total persistence ρ=−0.94. A secondary claim is the 'H0/H1 scissors': difficulty increases H1 area and H0 counts while decreasing durable H0 persistence, with durable H1 persistence reversing sign at high sampling density. The trefoil family is reported as an exception, showing a negative H1-area/error relationship interpreted as representational collapse.

Significance. The paper is potentially useful: it combines a controlled, known-truth benchmark with a modern topological tool and avoids circularity because the homology descriptors are computed from hidden representations while reliability targets are computed from predictions versus the generative labels. The large-sample warped-circle study is a good design for reducing sampling noise, the Swiss roll negative control is appropriate, and the explanation of H1 saturation via the finite cycle capacity of k-NN graphs is concrete and testable. If the topology–reliability association survives control for the difficulty schedule, the paper would establish a genuinely new, label-free diagnostic for TabPFN's internal task geometry. However, as written, the headline correlations are pooled across a hand-built difficulty ladder, and the descriptors are themselves nearly monotone functions of that ladder (Table 4). The trefoil sign inversion further shows that the direction of the relationship is not universal. These issues do not invalidate the raw observations, but they do mean the central diagnostic claim is not yet established.

major comments (3)
  1. [§5.1, Table 2; §3.4 and Table 4] The headline correlations (H0 fragmentation vs MAR ρ=0.92, H0 total persistence vs Bayes error ρ=−0.94) are pooled over 45 runs spanning 9 difficulty levels. Difficulty is a common cause: Table 4 shows H0 count correlates 0.94–0.98 with difficulty, H0 total persistence −0.91 to −0.94, while MAR and Bayes error rise with difficulty by construction through Eq. (5)–(6). The paper never reports within-level correlations or partial Spearman coefficients controlling for difficulty level. Without this, the claim 'topology diagnoses reliability' is indistinguishable from 'difficulty drives both.' Please add partial correlations, within-level analyses, or a formal mediation-style comparison, and state what remains after controlling the ladder.
  2. [§5.3, Table 3; §5.5, Table 5] The pooled six-family correlations (ρ≈0.25–0.51) inherit the same confound, and family identity is entangled with the family-specific difficulty knobs (Tables 6–7). The trefoil's inversion in Table 5 (ρ=−0.47 for H1 area vs observed error) demonstrates that even the sign of the topology–reliability relationship is not fixed across geometries. The paper should report family-stratified analyses with difficulty level as a covariate, and should not claim a general diagnostic until it can either define the regime in which the positive association holds or provide a way to identify the collapse regime from observable quantities alone.
  3. [§6.3, Practical implications] The proposed use of topology as a dataset-level reliability diagnostic requires that the descriptors carry information beyond the hand-chosen input difficulty variables (σ, m, w, and family knobs). The manuscript provides no comparison against a baseline that uses the input difficulty schedule, or simple input statistics such as noise level and nuisance count, to predict MAR/Bayes error. If H0 fragmentation is just a proxy for σ and m, the §6.3 diagnostic adds nothing over flagging the input directly. Please add an experiment where topological descriptors are evaluated for incremental predictive value after controlling for the generator's difficulty parameters.
minor comments (4)
  1. [§3.3] 'H0 fragmentation count' and 'durable H1 persistence' are used throughout but never formally defined. Please give exact definitions, e.g., whether fragmentation count is the number of H0 bars, the number of bars above a persistence threshold, or a separate fragmentation index, and specify the threshold used for 'durable.'
  2. [§3.2 and §6.4] The k-NN graph uses k=4 and a maximum simplex dimension that is not stated in the main text. The limitations section mentions dependence on these choices, but the actual value of the maximum simplex dimension should be reported in the experimental setup, not left implicit.
  3. [Table 4] The table reports n=45 runs per sample size for the warped circle, which is consistent with 9 levels × 5 seeds. If the n_test=300 column comes from the main suite plus the extreme suite, please state this explicitly; otherwise a reader may expect 30 runs for the main suite alone.
  4. [§5.2, Figure 3] The hard-versus-easy tertile contrast compares 'easiest and hardest difficulty tertiles' but the split criteria and the number of runs per tertile are not fully specified. Please clarify how tertiles are formed when there are 9 levels and 6 families.

Circularity Check

0 steps flagged

No significant circularity: topological descriptors and reliability targets are computed independently, self-citations are not load-bearing, and the difficulty-ladder proxy concern is a confounding caveat rather than a derivation loop.

full rationale

The topological descriptors are derived from zigzag persistence on TabPFN hidden representations (Sections 3.2–3.3), while MAR, Bayes error, and observed error are computed from predicted probabilities against generator-known true probabilities (Eqs. 3–6). No parameter of the topological pipeline is fitted to any reliability target; k=4, homology dimensions, and effective-layer mapping are fixed before the correlations are measured. The central correlations are therefore not self-definitional: H0 fragmentation count is not defined in terms of MAR, and MAR is not defined in terms of any homology descriptor. The main caveat—that descriptors and reliability metrics both track the hand-built difficulty ladder (Tables 6–7), and the paper does not partial out difficulty—is a confounding/validity limitation, not circularity, because the reported associations are empirical and could in principle fail; indeed they invert for the trefoil (Table 5) and weaken for raw H1 area at scale (Table 2). The trefoil sign inversion is direct evidence that no definitional tie forces the topology–reliability direction. Self-citations [21] and [17] are contextual or setup references only; the load-bearing methodological choices (zigzag construction, k=4) come from external prior work [5, 12, 23]. The paper's own §6.4 limitations label the descriptors as a coarse context-level diagnostic rather than a mechanistic attribution method. Nothing in the text exhibits an equation or fitted parameter that reduces the claimed prediction to its own input, so no circular step meets the evidentiary standard.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 1 invented entities

The central claim rests on representational faithfulness (k-NN homology of one embedding scope ≈ 'task geometry'), on the generators realizing their stated homotopy types at all difficulty levels, and on unstated graph-construction parameters. No entity is invented; the 'scissors' is a descriptive label for an observed correlation, not an independent postulate. The hand-built difficulty ladder is the main free-parameter surface: descriptors track it at ρ≥0.94, so it, not the model, may be the actual common cause of the reported associations.

free parameters (7)
  • k (k-NN graph neighbors) = 4
    §3.2: graph size for every layer's simplicial complex; inherited from LLM zigzag work [12], never varied. Homology counts (H0/H1) are direct functions of k.
  • Maximum simplex dimension (clique filling) = not stated
    §3.2: 'a fixed maximum simplex dimension' is referenced but never given a value; changing it changes which H1 cycles survive.
  • 'Durable feature' threshold = not stated
    §3.3, §5.1: durable-feature fractions drive the headline scale-invariance result (durable H1 sign flip, Table 4) but the durability cutoff is never defined.
  • π_hist weighting exponent α = 1
    §3.3, §5.8: 'unless otherwise noted we report α=1'; the timing analysis (peak-layer shifts) is reported for this one choice.
  • Label sharpness κ = 4.0
    Appendix A.1 Eq. (10): fixed across all runs to decouple Bayes noise from geometric difficulty; sets the MAR dynamic range and is not otherwise justified.
  • Difficulty ladder (σ, m, w per level + family knobs τ, δ, s_f, n_turns) = Tables 6–7, levels 0–8
    Hand-chosen schedule of 9 difficulty levels. It is the common driver of both descriptors (H0 count ρ≥0.94 with level, Table 4) and every reliability metric; this is the confound that a difficulty-partialing baseline would resolve.
  • Overconfidence-rate threshold = r > 0.10
    §3.4: defines the overconfidence and wrong-overconfidence metrics; arbitrary cutoff.
axioms (5)
  • domain assumption Query label-token representations, turned into k=4 k-NN clique complexes, faithfully represent TabPFN's 'inferred in-context task geometry'.
    §3.1–3.2. All interpretations (fragmentation, loops, scissors) are read off this construction; embedding-scope robustness checks (§5.10) are partial support, but no ground-truth check of the representation topology exists.
  • domain assumption The six generators realize the asserted homotopy types (S1, T2, S2, Hopf, trefoil, R2 sheet) at every difficulty level.
    §4, Appendix A.1. No verification is reported that, e.g., the level-8 trefoil (scale factor 1.005) or the noisy Hopf link retains the intended topology; the trefoil inversion may partly reflect topology change, not model collapse.
  • standard math Zigzag persistence intervals, mapped to 'effective model-layer intervals,' are stable summaries that remove artifacts.
    §3.2, following [5, 12]; relies on persistence stability [7, 29]. The effective-interval mapping is a modeling choice inherited from LLM work.
  • domain assumption Pooled Spearman correlations across the hierarchical design (family × level × seed) are valid.
    §3.4. Pooling 180 runs confounds family identity with the difficulty schedule; no mixed-effects or within-family pooled analysis is reported, so pooled ρ values may be inflated by between-family differences.
  • domain assumption TabPFN v2 with feature group size=1, one estimator, and feature shuffling disabled is representative of TabPFN.
    Appendix A.2: configuration chosen to remove ensembling artifacts, but the topology of a single estimator may differ from the ensemble behavior users actually deploy.
invented entities (1)
  • H0/H1 'scissors' pattern and 'fragmentation principle' no independent evidence
    purpose: Interpretive labels summarizing observed joint movement of two homology descriptors across the synthetic difficulty ladder.
    These are descriptive names for a correlation structure within the designed benchmark, not independently falsifiable postulates; no new physical or mathematical entities are introduced. The 'scissors' has no handle outside this paper's synthetic setting.

pith-pipeline@v1.3.0-alltime-deepseek · 15306 in / 22178 out tokens · 178296 ms · 2026-08-01T16:31:05.045920+00:00 · methodology

0 comments
read the original abstract

TabPFN is a transformer-based foundation model for tabular prediction that performs inference without task-specific training by conditioning on a support set and query inputs. Despite its strong empirical performance, its internal behavior on structurally difficult tabular geometries remains poorly understood. We study this behavior using zigzag persistent homology, treating TabPFN layer representations as evolving point clouds. We construct a controlled benchmark of synthetic tabular tasks with known true probabilities and varied intrinsic topology, including warped circles, tori, spheres, Hopf links, trefoil knots, and Swiss rolls. Across these tasks, we find that the topology of TabPFN's internal representation geometry is strongly associated with dataset-level reliability; for example, the zeroth homology group $H_0$ fragmentation count correlates positively with mean absolute residual across controlled tasks, and this association strengthens in a high-resolution warped circle case study at large sample size. Harder geometries induce a dual topological signature: increased $H_1$ loop activity and increased $H_0$ fragmentation, while the $H_1$ persistence becomes shorter-lived. These descriptors correlate with Bayes error, mean absolute residuals, and overconfidence. Our results suggest that zigzag persistence diagnoses the reliability of the inferred in-context task geometry and provides a context-level view of when TabPFN operates in topologically stressed regimes.

Figures

Figures reproduced from arXiv: 2607.17962 by James Hu, Mahdi Ghelichi.

Figure 1
Figure 1. Figure 1: The six controlled topology families, shown as the ambient point cloud of a representative [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: warped circle case study at scale (45 runs, [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: (a) Embedding-topology descriptors in the easy vs hard difficulty tertiles; all loop and [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Dataset-level relationships between embedding topology and reliability (each point is [PITH_FULL_IMAGE:figures/full_fig_p023_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Standardized effect sizes (Cohen’s d) for the hard-versus-easy tertile contrast. Loop activity and the H0 fragmentation count rise with difficulty while Bayes error increases; H0 total persistence (Section 5.4) moves in the opposite direction. (a) (b) (c) [PITH_FULL_IMAGE:figures/full_fig_p024_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The H0/H1 scissors across difficulty levels (mean ± s.d. per family). (a) H1 area rises with difficulty; (b) H0 total persistence falls; (c) H0 count rises. Loop structure proliferates while stable cluster structure fragments. 24 [PITH_FULL_IMAGE:figures/full_fig_p024_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Layerwise timing of H1 resolution. (a) Persistence-weighted birth histogram πhist(α = 1) for easy, medium, and hard datasets. (b) The πhist peak layer as a function of difficulty; entangled families resolve their loops deeper in the network. Each individual run’s peak falls on one of the 12 discrete TabPFN layers; the plotted curve is the mean peak layer over the five seeds, so it can take fractional value… view at source ↗
Figure 8
Figure 8. Figure 8: Birth–persistence heatmaps for the Hopf link (log count of cycles by birth layer and [PITH_FULL_IMAGE:figures/full_fig_p025_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 8 linked inside Pith

  1. [1]

    Tabnet: Attentive interpretable tabular learning

    Sercan ¨O Arik and Tomas Pfister. Tabnet: Attentive interpretable tabular learning. InPro- ceedings of the AAAI conference on artificial intelligence, volume 35, pages 6679–6687, 2021

  2. [2]

    Orion-msp: Multi-scale sparse attention for tabular in-context learning.arXiv preprint arXiv:2511.02818, 2025

    Mohamed Bouadi, Pratinav Seth, Aditya Tanna, and Vinay Kumar Sankarapu. Orion-msp: Multi-scale sparse attention for tabular in-context learning.arXiv preprint arXiv:2511.02818, 2025

  3. [3]

    Random forests.Machine learning, 45(1):5–32, 2001

    Leo Breiman. Random forests.Machine learning, 45(1):5–32, 2001

  4. [4]

    Topology and data.Bulletin of the American Mathematical Society, 46(2):255–308, 2009

    Gunnar Carlsson. Topology and data.Bulletin of the American Mathematical Society, 46(2):255–308, 2009

  5. [5]

    Zigzag persistence.Foundations of Computational Math- ematics, 10(4):367–405, 2010

    Gunnar Carlsson and Vin de Silva. Zigzag persistence.Foundations of Computational Math- ematics, 10(4):367–405, 2010

  6. [6]

    Xgboost: A scalable tree boosting system

    Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. InProceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pages 785–794, 2016

  7. [7]

    Stability of persistence diagrams

    David Cohen-Steiner, Herbert Edelsbrunner, and John Harer. Stability of persistence diagrams. Discrete & Computational Geometry, 37(1):103–120, 2007

  8. [8]

    A survey on in-context learning

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, et al. A survey on in-context learning. InProceedings of the 2024 conference on empirical methods in natural language processing, pages 1107–1128, 2024

  9. [9]

    American Mathematical Society, 2010

    Herbert Edelsbrunner and John Harer.Computational Topology: An Introduction. American Mathematical Society, 2010

  10. [10]

    Topological persistence and simplification.Discrete & Computational Geometry, 28(4):511–533, 2002

    Herbert Edelsbrunner, David Letscher, and Afra Zomorodian. Topological persistence and simplification.Discrete & Computational Geometry, 28(4):511–533, 2002

  11. [11]

    Tabarena: A living benchmark for machine learning on tabular data.arXiv preprint arXiv:2506.16791, 2025

    Nick Erickson, Lennart Purucker, Andrej Tschalzev, David Holzm¨ uller, Prateek Mutalik Desai, David Salinas, and Frank Hutter. Tabarena: A living benchmark for machine learning on tabular data.arXiv preprint arXiv:2506.16791, 2025

  12. [12]

    Persistent topological features in large language models

    Yuri Gardinazzi, Karthik Viswanathan, Giada Panerai, Alessio Ansuini, Alberto Cazzaniga, and Matteo Biagetti. Persistent topological features in large language models. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceedings of Machine Learning Research, pages 18811–18830. PMLR, 2025

  13. [13]

    Revisiting deep learning models for tabular data.Advances in neural information processing systems, 34:18932– 18943, 2021

    Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Revisiting deep learning models for tabular data.Advances in neural information processing systems, 34:18932– 18943, 2021

  14. [14]

    Tabpfn-2.5: Advancing the state of the art in tabular foundation models.arXiv preprint arXiv:2511.08667, 2025

    L´ eo Grinsztajn, Klemens Fl¨ oge, Oscar Key, Felix Birkel, Philipp Jund, Brendan Roof, Ben- jamin J¨ ager, Dominik Safaric, Simone Alessi, Adrian Hayler, et al. Tabpfn-2.5: Advancing the state of the art in tabular foundation models.arXiv preprint arXiv:2511.08667, 2025

  15. [15]

    Tabpfn-3: Technical report.arXiv preprint arXiv:2605.13986, 2026

    L´ eo Grinsztajn, Klemens Fl¨ oge, Oscar Key, Felix Birkel, Philipp Jund, Brendan Roof, Mihir Manium, Shi Bin Hoo, Magnus B¨ uhler, Anurag Garg, et al. Tabpfn-3: Technical report.arXiv preprint arXiv:2605.13986, 2026. 17

  16. [16]

    Why do tree-based models still out- perform deep learning on typical tabular data?Advances in neural information processing systems, 35:507–520, 2022

    L´ eo Grinsztajn, Edouard Oyallon, and Ga¨ el Varoquaux. Why do tree-based models still out- perform deep learning on typical tabular data?Advances in neural information processing systems, 35:507–520, 2022

  17. [17]

    Generative ai enhanced financial risk management information retrieval.arXiv:2504.06293, 2025

    Amin Haeri, Jonathan Vitrano, and Mahdi Ghelichi. Generative ai enhanced financial risk management information retrieval.arXiv:2504.06293, 2025

  18. [18]

    Topological feature-driven tabpfn model for prediction of enlarged hemorrhage and edema after tumor resection in meningiomas.Frontiers in Medicine, 13:1808831, 2026

    Wenjing Han, Guirong Tan, Lijia Li, Zhenyang Feng, Chen Zhou, Xiang Liu, and Lingjing Hu. Topological feature-driven tabpfn model for prediction of enlarged hemorrhage and edema after tumor resection in meningiomas.Frontiers in Medicine, 13:1808831, 2026

  19. [19]

    TabPFN: A transformer that solves small tabular classification problems in a second

    Noah Hollmann, Samuel M¨ uller, Katharina Eggensperger, and Frank Hutter. TabPFN: A transformer that solves small tabular classification problems in a second. InInternational Conference on Learning Representations, 2023

  20. [20]

    Accurate predictions on small data with a tabular foundation model.Nature, 637(8045):319–326, 2025

    Noah Hollmann, Samuel M¨ uller, Lennart Purucker, Arjun Krishnakumar, Max K¨ orfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter. Accurate predictions on small data with a tabular foundation model.Nature, 637(8045):319–326, 2025

  21. [21]

    Noise immunity in in-context tabular learning: An empirical robustness analysis of tabpfn’s attention mechanisms.ArXiv:2604.04868, 2026

    James Hu and Mahdi Ghelichi. Noise immunity in in-context tabular learning: An empirical robustness analysis of tabpfn’s attention mechanisms.ArXiv:2604.04868, 2026

  22. [22]

    Robustness of random forest-based gene selection methods.BMC bioinformatics, 15(1):8, 2014

    Miron Bartosz Kursa. Robustness of random forest-based gene selection methods.BMC bioinformatics, 15(1):8, 2014

  23. [23]

    Persistent homology with k-nearest-neighbor filtrations reveals topological convergence of pagerank.arXiv preprint arXiv:2206.04725, 2022

    Minh Quang Le and Dane Taylor. Persistent homology with k-nearest-neighbor filtrations reveals topological convergence of pagerank.arXiv preprint arXiv:2206.04725, 2022

  24. [24]

    Generalization can emerge in tabular foundation models from a single table.arXiv preprint arXiv:2511.09665, 2025

    Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L Caterini, and Valentin Thomas. Generalization can emerge in tabular foundation models from a single table.arXiv preprint arXiv:2511.09665, 2025

  25. [25]

    Tab- dpt: Scaling tabular foundation models on real data.arXiv preprint arXiv:2410.18164, 2024

    Junwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach, Hamidreza Kamkari, Jesse C Cresswell, Keyvan Golestan, Guangwei Yu, Anthony L Caterini, and Maksims Volkovs. Tab- dpt: Scaling tabular foundation models on real data.arXiv preprint arXiv:2410.18164, 2024

  26. [26]

    The GUDHI library: Simplicial complexes and persistent homology

    Cl´ ement Maria, Jean-Daniel Boissonnat, Marc Glisse, and Mariette Yvinec. The GUDHI library: Simplicial complexes and persistent homology. InMathematical Software – ICMS 2014, pages 167–174. Springer, 2014

  27. [27]

    Topology of deep neural networks

    Gregory Naitzat, Andrey Zhitnikov, and Lek-Heng Lim. Topology of deep neural networks. Journal of Machine Learning Research, 21(184):1–40, 2020

  28. [28]

    Assessing the robustness of tabular prior-data fitted network classifier

    Ali Nawaz, Amir Ahmad, and Shehroz S Khan. Assessing the robustness of tabular prior-data fitted network classifier. In1st ICML Workshop on Foundation Models for Structured Data, 2025

  29. [29]

    Finding the homology of submanifolds with high confidence from random samples.Discrete & Computational Geometry, 39(1–3):419– 441, 2008

    Partha Niyogi, Stephen Smale, and Shmuel Weinberger. Finding the homology of submanifolds with high confidence from random samples.Discrete & Computational Geometry, 39(1–3):419– 441, 2008

  30. [30]

    Catboost: unbiased boosting with categorical features.Advances in neural information processing systems, 31, 2018

    Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin. Catboost: unbiased boosting with categorical features.Advances in neural information processing systems, 31, 2018. 18

  31. [31]

    Neural persistence: A complexity measure for deep neural net- works using algebraic topology

    Bastian Rieck, Matteo Togninalli, Christian Bock, Michael Moor, Max Horn, Thomas Gumb- sch, and Karsten Borgwardt. Neural persistence: A complexity measure for deep neural net- works using algebraic topology. InInternational Conference on Learning Representations, 2019

  32. [32]

    Samaga, Gilberto Gonzalez Arroyo, and Tamal K

    Shreyas N. Samaga, Gilberto Gonzalez Arroyo, and Tamal K. Dey. HalluZig: Hallucination detection using zigzag persistence. InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, pages 3466–3482. Association for Computational Linguistics, 2026

  33. [33]

    Does tabpfn understand causal structures?arXiv preprint arXiv:2511.07236, 2025

    Omar Swelam, Lennart Purucker, Jake Robertson, Hanne Raum, Joschka Boedecker, and Frank Hutter. Does tabpfn understand causal structures?arXiv preprint arXiv:2511.07236, 2025

  34. [34]

    Exploring fine-tuning for tabular foundation models.arXiv preprint arXiv:2601.09654, 2026

    Aditya Tanna, Pratinav Seth, Mohamed Bouadi, and Vinay Kumar Sankarapu. Exploring fine-tuning for tabular foundation models.arXiv preprint arXiv:2601.09654, 2026

  35. [35]

    Why tabular foundation models should be a research priority.arXiv preprint arXiv:2405.01147, 2024

    Boris Van Breugel and Mihaela Van Der Schaar. Why tabular foundation models should be a research priority.arXiv preprint arXiv:2405.01147, 2024

  36. [36]

    Attention is all you need.Advances in neural information processing systems, 30, 2017

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in neural information processing systems, 30, 2017

  37. [37]

    Transformers learn in-context by gra- dient descent

    Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, Jo˜ ao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. Transformers learn in-context by gra- dient descent. InInternational Conference on Machine Learning, pages 35151–35174. PMLR, 2023

  38. [38]

    A closer look at TabPFN v2: Understanding its strengths and extending its capabilities.arXiv:2502.17361, 2025

    Han-Jia Ye, Si-Yang Liu, and Wei-Lun Chao. A closer look at TabPFN v2: Understanding its strengths and extending its capabilities.arXiv:2502.17361, 2025

  39. [39]

    Computing persistent homology.Discrete & Compu- tational Geometry, 33(2):249–274, 2005

    Afra Zomorodian and Gunnar Carlsson. Computing persistent homology.Discrete & Compu- tational Geometry, 33(2):249–274, 2005. 19 This appendix provides additional details about the experimental setups used in the main paper. A Experimental Details A.1 Synthetic Data Generation All six topology families are produced by a common pipeline; they differ only in...