Schema-1 is the first Data Language Model that natively understands raw tabular data and outperforms gradient-boosted ensembles, AutoML, and prior tabular foundation models on row-level prediction and imputation tasks.
and Schmidt, Ludwig , year =
8 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
representative citing papers
FT-MDN-Transformer improves transfer learning for loan recovery rate prediction under covariate, conditional, and label shifts with heterogeneous features, outperforming baselines when target data is limited.
LLM embeddings from policy text outperform hand-engineered features in a GLM for French motor insurance claim frequency, with larger gains at small sample sizes and further improvement from insurance-specific fine-tuning.
ADIGen generates counterfactuals under general interventions via Riesz regression, causal invariance, and orthogonal learning, with excess-risk bounds featuring product-bias remainder and invariant risk across environments.
For tabular in-context learning models, recourse is well-defined, its cost is bounded and converges to classical linear recourse as context grows; ASR-ICL finds sparse recourse with fewer queries.
LLMTabBench evaluates LLMs on zero- and few-shot binary tabular classification and reports that zero-shot can outperform few-shot due to example conflicts with model priors while performance drops beyond a complexity threshold.
TabICL scales in-context learning to large tabular data via column-then-row attention for row embeddings followed by a transformer, matching TabPFNv2 speed and performance while outperforming it and CatBoost on datasets over 10K samples.
PPE framework with T3+OCSVM one-class detector reaches 0.93+ borderline AUROC, cuts false positives 44-55 points versus Gaussian baselines, and runs at millisecond latency on synthetic multi-domain data.
citing papers explorer
-
Data Language Models: A New Foundation Model Class for Tabular Data
Schema-1 is the first Data Language Model that natively understands raw tabular data and outperforms gradient-boosted ensembles, AutoML, and prior tabular foundation models on row-level prediction and imputation tasks.
-
Transfer Learning for Loan Recovery Prediction under Distribution Shifts with Heterogeneous Feature Spaces
FT-MDN-Transformer improves transfer learning for loan recovery rate prediction under covariate, conditional, and label shifts with heterogeneous features, outperforming baselines when target data is limited.
-
Semantic insurance pricing with large language models
LLM embeddings from policy text outperform hand-engineered features in a GLM for French motor insurance claim frequency, with larger gains at small sample sizes and further improvement from insurance-specific fine-tuning.
-
Automatic, Debiased, and Invariant Counterfactual Generation under General Interventions
ADIGen generates counterfactuals under general interventions via Riesz regression, causal invariance, and orthogonal learning, with excess-risk bounds featuring product-bias remainder and invariant risk across environments.
-
Algorithmic Recourse of In-Context Learning for Tabular Data
For tabular in-context learning models, recourse is well-defined, its cost is bounded and converges to classical linear recourse as context grows; ASR-ICL finds sparse recourse with fewer queries.
-
LLMTabBench: Evaluating LLMs on Binary Tabular Classification From Zero to Few Shots
LLMTabBench evaluates LLMs on zero- and few-shot binary tabular classification and reports that zero-shot can outperform few-shot due to example conflicts with model priors while performance drops beyond a complexity threshold.
-
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
TabICL scales in-context learning to large tabular data via column-then-row attention for row embeddings followed by a transformer, matching TabPFNv2 speed and performance while outperforming it and CatBoost on datasets over 10K samples.
-
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
PPE framework with T3+OCSVM one-class detector reaches 0.93+ borderline AUROC, cuts false positives 44-55 points versus Gaussian baselines, and runs at millisecond latency on synthetic multi-domain data.