REVIEW 3 major objections 4 minor 73 references
Three small specialist language models — one for knowledge, one for reasoning, one for code — fill missing spreadsheet cells more accurately than frontier reasoning models, at under 1% of the cost.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 11:31 UTC pith:OIN3JWIA
load-bearing objection Auto-Fill is a strong systems paper with a real calibration gap: the ensemble confidence isn't recalibrated after max-selection, so the precision-at-threshold guarantee isn't supported, but the empirical rank-based results are credible. the 3 major comments →
Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
This paper claims that high-precision missing-value prediction in tables is best solved by a calibrated ensemble of three specialist small language models. A knowledge specialist answers directly from memory and formats output to match the column; a reasoning specialist is distilled on chain-of-thought traces; a coding specialist emits column-level Python that reconstructs the target column. Each carries a mode-appropriate confidence signal — token log-probabilities, teacher sampling agreement, execution-based column reconstruction — mapped by isotonic regression to one probability scale. The ensemble returns the most confident prediction only above a precision threshold, else abstains.
What carries the argument
The load-bearing mechanism is the specialist ensemble plus calibrated abstention. Three small models are post-trained so that each handles one mode: the knowledge specialist (direct supervised fine-tuning; confidence from geometric-mean token log-probabilities), the reasoning specialist (distilled on the teacher's shortest correct chain-of-thought trace; confidence from a two-stage product of correctness rate × mean verbalized confidence), and the coding specialist (distilled on teacher-generated column-level code; confidence from executing the code and measuring how many observed column values it reconstructs). Because the three raw scores are not comparable, isotonic regression maps each t
Load-bearing premise
The load-bearing premise is that uniformly random cell masking produces training and test cases that represent how values really go missing in real tables; if actual missingness is systematic rather than random, the calibrated precision guarantees may not hold in deployment.
What would settle it
Take naturally missing cells from real spreadsheets, recover their true values from source documents, run the released system, and measure precision on the predictions it does not abstain from: if precision falls materially below the calibrated 0.9 threshold, the random-masking assumption fails. A cheaper probe is to re-mask the same 2,200 benchmark tables with a non-random missingness model — e.g., masking cells conditional on row or column features — and compare R@P=0.9.
If this is right
- Missing-value prediction can be deployed at spreadsheet scale: the best Auto-Fill variant returns a 0.9-precision segment with mean recall 0.656 at $0.0113 per query, while o3-pro's comparable segment reaches 0.56 at $0.177 per query.
- Calibrated abstention, not raw accuracy, is what separates usable from hallucinating predictors: frontier models reach near-perfect recall on deterministic relational tables but collapse on benchmarks that require saying 'I don't know'.
- Separate specialists beat a single hybrid model: the hybrid collapses to near-zero recall under abstention, while the ensemble is best on every dataset, confirming capability interference in shared parameters.
- The three-mode decomposition is near-exhaustive: the paper's error audit attributes 89 of 100 sampled failures to capability gaps within the three specialists and only 11 to genuinely unrecoverable cells.
- Cost scales down gracefully: smaller backbones (Qwen 1.7B/4B, GPT-4.1 nano) retain most of the quality, so practitioners can trade recall for cost within the same framework.
Where Pith is reading between the lines
- Editorial: the masking assumption is the transfer risk. Training and evaluation both use uniformly random cell masking; real-world missingness is often not random — cells tend to be absent because of the very column context the models attend to. If that distribution shift is material, the advertised 0.9 precision at deployment would drop below the measured level.
- Editorial: the coding specialist's execution-based self-validation is a transferable idea — any prediction expressible as a column transformation (SQL, formula, unit conversion) can be validated against observed rows without labels, which may apply to error detection and schema repair beyond imputation.
- Editorial: the error audit points to an obvious paper-adjacent improvement: 52 of 100 sampled failures are knowledge-specialist misses, many of them long-tail facts a web lookup would settle. Augmenting the knowledge specialist with retrieval would likely raise recall, at the cost of the paper's strict per-query budget.
- Editorial: the 'exact match or abstain' framing may be the paper's most transferable contribution — it reframes imputation from statistical closeness to decision-making under a cost of error, a framing that could benefit other data-cleaning tasks such as error detection and schema matching.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses high-precision missing-value prediction in tables: a model must either predict the exact missing value or abstain, with precision above a user threshold τ. It proposes Auto-Fill, which post-trains three specialist small language models — a knowledge specialist (direct SFT, log-probability confidence), a reasoning specialist (distillation from DeepSeek-R1, two-stage sampling-based confidence), and a coding specialist (distillation of column-level code, execution-based confidence). Each specialist’s confidence is calibrated with isotonic regression, and the ensemble selects the most confident specialist or abstains if the maximum calibrated confidence is below τ. Experiments on 11 benchmarks with 2,200 tables (5 in-distribution, 6 out-of-distribution) compare Auto-Fill against frontier reasoning models, retrieval imputation, classical repair, and tabular ML baselines. The paper reports mean R@P=0.9 of 0.628 (Auto-Fill-Qwen) and 0.656 (Auto-Fill-GPT), above the best baseline o3-pro (0.56), at substantially lower cost. It also includes extensive ablations, negative results (modest RL gains, hybrid model collapse), calibration comparisons, and a manual error audit.
Significance. If the empirical claims hold, this is a practically valuable contribution: it demonstrates that a calibrated ensemble of small specialized models can achieve high-precision table repair at a fraction of the cost of frontier models, and the explicit abstention formulation is well matched to real spreadsheet use. The paper’s strengths include a publicly released benchmark suite, open-source code, systematic ablation of the design space, honest documentation of negative results, a manual failure audit, and an explicit comparison of calibration methods. The core capability-decomposition idea is well supported by the ablation tables. However, the central ‘calibrated abstention ensures precision ≥ τ’ claim is not fully verified because the ensemble selection step itself is not recalibrated, and the cost claim in the abstract is stronger than the reported tables support. These are fixable but load-bearing issues.
major comments (3)
- [§6.2, Eq. (9); §6.1, Eq. (8); Appendix D] The max-of-calibrated-confidences selection is not recalibrated. Even if each g_i is marginally calibrated (Eq. 8), the event that specialist i has the largest of three noisy calibrated scores is correlated with over-estimation of its true correctness probability. Thus P(correct | i*=i, conf_i=p) need not equal p, and the Definition 1 precision guarantee at threshold τ is not established. Table 9 reports per-specialist ECE but no ECE for the ensemble’s final selected confidence. The error audit in Appendix D is symptomatic: the 10 mis-routed failures and 11 unrecoverable failures are, by definition, cases where the ensemble made a confident wrong prediction rather than abstaining; the audit’s claim that unrecoverable cases are ‘silent deferrals’ is inconsistent with them being failures. The authors should provide a second-stage isotonic calibration of the selected output, or equivalently
- [§5, Example 4; §7.1, A.1] The training and most test protocols use uniformly random cell masking; A.1 states that for non-relational datasets cells are ‘uniformly sampled’ and only one cell per table is masked. Real spreadsheet missingness is likely not MCAR — missingness often depends on the row/column context, and the paper’s deployment target is Excel/Sheets at scale. Under MNAR, the advertised precision-at-τ guarantee, which is calibrated on random masking, may not transfer. The paper’s own Appendix D shows 11% of sampled failures are values not deducible from context, supporting the concern that real missing data can violate the learnable-rule assumption. The authors should evaluate on naturally missing cells (e.g., cells that were originally empty and held out) or provide a direct argument, with data, that random masking is representative of the missingness mechanism in their target deployment. Without this
- [Abstract; Figure 5; Table 2] The abstract and Figure 5 state that Auto-Fill operates at ‘less than 1%’ of the cost of frontier models. From Table 2, Auto-Fill-Qwen’s mean cost is $1.48e-3, which is about 0.84% of o3-pro’s $1.77e-1, but it is 3.4% of Gemini 3 Pro’s $4.30e-2. Auto-Fill-GPT, the higher-quality variant, costs $1.13e-2, which is 6.4% of o3-pro and 26% of Gemini 3 Pro. Thus the ‘less than 1%’ claim holds only for Auto-Fill-Qwen against o3-pro, not for the other frontier baselines reported in the same sentence. Please qualify the cost claim to match the data, or present the cost comparison separately by baseline.
minor comments (4)
- [§5.2, Eq. (5)] Eq. (5) mixes units: in Example 5, the verbalized confidence is an integer 0–100, producing conftrain = 76, while later sections treat confidence in [0,1]. Please state the normalization explicitly.
- [§7.4, Table 6] The hybrid model’s R@P=0.9 of 0.000 on ID while its full accuracy is 0.558 (Table 14) is striking; a sentence clarifying that this is a confidence-calibration collapse rather than a pure prediction failure would help readers interpret the zero entry.
- [Appendix B.2, Table 9] The ECE values in Table 9 are per-specialist. Please add a line reporting ECE for the ensemble’s selected confidence, since that is what governs the abstention decision.
- [Appendix D] In the residual error labels, the ‘within-M_K’ category includes cases that the authors propose to fix with retrieval augmentation, but retrieval is not implemented in the current system. This makes the ‘89/100 absorbed by the decomposition’ statement optimistic. Consider reporting the subset that is recoverable without adding new machinery.
Circularity Check
No significant circularity: the derivation is self-contained and externally evaluated.
full rationale
Auto-Fill's derivation chain is not circular. Training data are generated by masking known cell values (Example 4), each specialist is trained to predict the masked value and emit a confidence signal, and the final confidence is mapped to empirical correctness by isotonic regression on a held-out validation set (Eq. 8), disjoint from test. The high-precision evaluation then ranks predictions by this calibrated confidence on held-out masked tables with known ground truth (Section 7.1), so the reported R@P=0.9 numbers are measured against external correctness labels rather than implied by any fitted parameter. The reasoning specialist's confidence target (Eq. 5) does use the teacher's verbalized self-assessment, but it is only a training label; the same specialist's outputs are subsequently recalibrated against actual correctness, so no prediction reduces to its input by construction. Self-citations (e.g., DeepSeek-R1 as teacher, same-author benchmark sources such as Auto-Relate) are used as tools or data sources, not as unverified premises that force the conclusion. The acknowledged residual issues in Appendix D — mis-routes and unrecoverable cases — are calibration/deployment-quality concerns, and the max-selection calibration gap noted in review is a statistical validity concern, but neither is a definitional equivalence between output and input. No quoted equation or fitted parameter is renamed as a prediction. Hence no circular step meets the evidentiary bar.
Axiom & Free-Parameter Ledger
free parameters (6)
- abstention/precision threshold τ =
0.9
- teacher sample size k =
10
- column reconstruction threshold tr =
0.8
- M_R confidence composition form =
E[conf|correct] × P(correct), equal weighting
- inference sampling temperatures =
0.8 (M_R, M_C); 0.1 (M_K)
- isotonic calibration set size =
2,000 held-out cases
axioms (5)
- domain assumption Uniformly random cell masking in training and test construction is representative of real missingness
- domain assumption Exact-match-or-abstain is the correct performance criterion for the deployment scenario
- domain assumption The knowledge/reasoning/coding decomposition is complete
- domain assumption Teacher (DeepSeek-R1) traces and self-assessed confidences are a sound distillation target
- standard math Isotonic regression on 2,000 validation cases yields true probabilities for each specialist
read the original abstract
Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show great promise in predicting missing values in tables, by reasoning holistically across rows and columns, they are costly to deploy at scale and tend to be overconfident, often generating hallucinated or false-positive predictions. In this paper, we observe that achieving high-precision missing-value prediction in tables requires a distinct combination of three capabilities: (1) world knowledge, (2) text-based reasoning, and (3) code-based reasoning. We systematically explore design choices for combining these capabilities, and propose an Auto-Fill approach that post-trains three specialist small language models (SLMs), each optimized for one capability. We develop a calibrated ensemble mechanism that either dynamically selects the most confident specialist or abstains, ensuring high accuracy. Extensive experiments on 11 benchmarks with 2200 real tables drawn from diverse domains show that Auto-Fill achieves superior accuracy compared to state-of-the-art reasoning models (e.g., o3-pro, Gemini 3 Pro, and DeepSeek R1), while operating at a fraction (less than 1%) of the cost of these frontier models. Our results highlight the effectiveness of specialization and calibrated abstention in the important domain of tabular data. Auto-Fill is publicly available at https://github.com/lyrain2001/auto-fill.
Figures
Reference graph
Works this paper leans on
-
[1]
Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Mu- ralidharan, Yingyan Celine Lin, and Pavlo Molchanov. 2025. Small language models are the future of agentic ai.arXiv preprint arXiv:2506.02153(2025)
Pith/arXiv arXiv 2025
-
[2]
Alex Bogatu, Alvaro AA Fernandes, Norman W Paton, and Nikolaos Konstanti- nou. 2020. Dataset discovery in data lakes. In2020 ieee 36th international confer- ence on data engineering (icde). IEEE, 709–720
2020
-
[3]
Christine P Chai. 2020. The importance of data cleaning: Three visualization examples.Chance33, 1 (2020), 4–9
2020
-
[4]
Qixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui, Song Ge, Haidong Zhang, Dongmei Zhang, and Surajit Chaudhuri. 2025. Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables.Pro- ceedings of the ACM on Management of Data3, 3 (2025), 1–27
2025
-
[5]
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023. Pro- gram of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.Transactions on Machine Learning Research(2023)
2023
-
[6]
Xu Chu, Ihab F Ilyas, Sanjay Krishnan, and Jiannan Wang. 2016. Data cleaning: Overview and emerging challenges. InProceedings of the 2016 international conference on management of data. 2201–2206
2016
-
[7]
Mehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld, Leshem Choshen, Yoon Kim, and Jacob Andreas. 2025. Beyond binary rewards: Training lms to reason about their uncertainty.arXiv preprint arXiv:2507.16806(2025)
Pith/arXiv arXiv 2025
-
[8]
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong, Bany- atsang Mphago, and Oteng Tabona. 2021. A survey on missing data in machine learning.Journal of Big data8, 1 (2021), 140
2021
-
[9]
Pedro J García-Laencina, José-Luis Sancho-Gómez, and Aníbal R Figueiras-Vidal
-
[10]
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024. A survey of confidence estimation and calibration in large lan- guage models. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6577–6595
2024
-
[11]
Google. 2020. Connected Sheets is Generally Available. Retrieved February 2026 from https://workspace.google.com/blog/product-announcements/connected- sheets-is-generally-available
2020
-
[12]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. 2025. Deepseek-r1: Incen- tivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948(2025)
Pith/arXiv arXiv 2025
-
[13]
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yifan Wu, YK Li, et al. 2024. DeepSeek-Coder: when the large language model meets programming–the rise of code intelligence.arXiv preprint arXiv:2401.14196(2024)
Pith/arXiv arXiv 2024
-
[14]
Ziyan Han, Yeye He, Shuyuan Kang, Min Xie, Weiwei Cui, Song Ge, Haidong Zhang, Dongmei Zhang, Surajit Chaudhuri, Rui Mao, et al. 2026. Auto-Relate: A Unified Approach to Discovering Reliable Functional Relationships Leveraging Statistical Tests.arXiv preprint arXiv:2606.07060(2026)
Pith/arXiv arXiv 2026
-
[15]
Xinrui He, Yikun Ban, Jiaru Zou, Tianxin Wei, Curtiss Cook, and Jingrui He
-
[16]
Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter. 2022. Tabpfn: A transformer that solves small tabular classification problems in a second.arXiv preprint arXiv:2207.01848(2022)
Pith/arXiv arXiv 2022
-
[17]
Ming Hua and Jian Pei. 2007. Cleaning disguised missing data: a heuristic approach. InProceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining. 950–958
2007
-
[18]
Anil Jadhav, Dhanya Pramod, and Krishnan Ramanathan. 2019. Comparison of performance of data imputation methods for numeric dataset.Applied Artificial Intelligence33, 10 (2019), 913–933
2019
-
[19]
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024. Openai o1 system card.arXiv preprint arXiv:2412.16720(2024)
Pith/arXiv arXiv 2024
-
[20]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran- Johnson, et al. 2022. Language models (mostly) know what they know.arXiv preprint arXiv:2207.05221(2022)
Pith/arXiv arXiv 2022
-
[21]
Ehud Karpas, Omri Abend, Yonatan Belinkov, Barak Lenz, Opher Lieber, Nir Ratner, Yoav Shoham, Hofit Bata, Yoav Levine, Kevin Leyton-Brown, et al. 2022. MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.arXiv preprint arXiv:2205.00445(2022)
Pith/arXiv arXiv 2022
-
[22]
Franklin, Ken Goldberg, and Eugene Wu
Sanjay Krishnan, Michael J. Franklin, Ken Goldberg, and Eugene Wu. 2017. BoostClean: Automated Error Detection and Repair for Machine Learning.arXiv preprint arXiv:1711.01299(2017)
Pith/arXiv arXiv 2017
-
[23]
Meelis Kull, Telmo Silva Filho, and Peter Flach. 2017. Beta calibration: a well- founded and easily implemented improvement on logistic calibration for binary classifiers. InArtificial intelligence and statistics. PMLR, 623–631
2017
-
[24]
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Deni- son, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. 2023. Measuring faithfulness in chain-of-thought reasoning.arXiv preprint arXiv:2307.13702(2023)
Pith/arXiv arXiv 2023
-
[25]
Jixuan Leng, Chengsong Huang, Banghua Zhu, and Jiaxin Huang. 2024. Taming overconfidence in llms: Reward calibration in rlhf.arXiv preprint arXiv:2410.09724 (2024)
Pith/arXiv arXiv 2024
-
[26]
Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2024. Table-gpt: Table fine-tuned gpt for diverse table tasks.Proceedings of the ACM on Management of Data2, 3 (2024), 1–28
2024
-
[27]
Yiming Lin, Yeye He, and Surajit Chaudhuri. 2023. Auto-bi: Automatically build bi-models leveraging local join prediction and global schema graph.arXiv preprint arXiv:2306.12515(2023)
Pith/arXiv arXiv 2023
-
[28]
Ryan Liu, Jiayi Geng, Addison J Wu, Ilia Sucholutsky, Tania Lombrozo, and Thomas L Griffiths. 2024. Mind your step (by step): Chain-of-thought can reduce performance on tasks where thinking makes humans worse.arXiv preprint arXiv:2410.21333(2024)
Pith/arXiv arXiv 2024
-
[29]
Fábio MF Lobato, Vincent W Tadaiesky, Igor M Araújo, and Ádamo L de Santana
-
[30]
Shane Culpepper, Shazia Sadiq, and Xiaoli Wang
Feng Luo, Hai Lan, Hui Luo, Zhifeng Bao, J. Shane Culpepper, Shazia Sadiq, and Xiaoli Wang. 2026. Missing Value Imputation in Tabular Data Lakes Unleashed: A Hybrid Approach.The VLDB Journal35, 2 (2026), 11
2026
-
[31]
Mohammad Mahdavi and Ziawasch Abedjan. 2020. Baran: Effective error cor- rection via a unified context representation and transfer learning.Proceedings of the VLDB Endowment13, 12 (2020), 1948–1961. Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Han, Dongmei Zhang, and Surajit Chaudhuri
2020
-
[32]
Giansalvatore Mecca, Paolo Papotti, Donatello Santoro, and Enzo Veltri. 2024. BUNNI: Learning Repair Actions in Rule-driven Data Cleaning.Journal of Data and Information Quality16, 2, Article 12 (2024), 31 pages
2024
-
[33]
Microsoft. 2026. Clean Data in Excel with Copilot. Retrieved February 2026 from https://support.microsoft.com/en-us/office/clean-data-in-excel-7fe20d89- 3f57-46d3-b659-e8f3ee853bda?ns=XLWAENDUSER&version=16
2026
-
[34]
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015. Obtain- ing well calibrated probabilities using bayesian binning. InProceedings of the AAAI conference on artificial intelligence, Vol. 29
2015
-
[35]
Preetum Nakkiran, Arwen Bradley, Adam Goliński, Eugene Ndiaye, Michael Kirchhof, and Sinead Williamson. 2025. Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs.arXiv preprint arXiv:2511.04869(2025)
arXiv 2025
-
[36]
Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu, Shuwei Liang, and Jianwei Yin. 2024. Automatic Data Repair: Are We Ready to Deploy?Proceedings of the VLDB Endowment17, 10 (2024), 2617–2630
2024
-
[37]
OpenAI. 2025. Introducing GPT-4.1 in the API. Retrieved February 2026 from https://openai.com/index/gpt-4-1/
2025
-
[38]
OpenAI. 2025. Introducing GPT-5. Retrieved February 2026 from https://openai. com/index/introducing-gpt-5/
2025
-
[39]
John Platt et al. 1999. Probabilistic outputs for support vector machines and com- parisons to regularized likelihood methods.Advances in large margin classifiers 10, 3 (1999), 61–74
1999
-
[40]
Jakub Podolak and Rajeev Verma. 2025. Read Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs. InProceedings of the 2nd Workshop on Uncertainty-A ware NLP (UncertaiNLP 2025). 247–258
2025
-
[41]
Erhard Rahm, Hong Hai Do, et al. 2000. Data cleaning: Problems and current approaches.IEEE Data Eng. Bull.23, 4 (2000), 3–13
2000
-
[42]
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiao- qing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al. 2023. Code llama: Open foundation models for code.arXiv preprint arXiv:2308.12950 (2023)
Pith/arXiv arXiv 2023
-
[43]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300(2024)
Pith/arXiv arXiv 2024
-
[44]
Leyang Shen, Gongwei Chen, Rui Shao, Weili Guan, and Liqiang Nie. 2024. Mome: Mixture of multimodal experts for generalist multimodal large language models. Advances in neural information processing systems37 (2024), 42048–42070
2024
-
[45]
Jie Song and Yeye He. 2021. Auto-validate: Unsupervised data validation us- ing data-domain patterns inferred from data lakes. InProceedings of the 2021 International Conference on Management of Data. 1678–1691
2021
-
[46]
Paul Stangel, David Bani-Harouni, Chantal Pellegrini, Ege Özsoy, Kamilia Zaripova, Matthias Keicher, and Nassir Navab. 2025. Rewarding doubt: A re- inforcement learning approach to calibrated confidence expression of large language models.arXiv preprint arXiv:2503.02623(2025)
arXiv 2025
-
[47]
Shreyas Subramanian, Vikram Elango, and Mecit Gungor. 2025. Small language models (slms) can still pack a punch: A survey.arXiv preprint arXiv:2501.05465 (2025)
Pith/arXiv arXiv 2025
-
[48]
Kai Sun, Yifan Xu, Hanwen Zha, Yue Liu, and Xin Luna Dong. 2024. Head-to-tail: How knowledgeable are large language models (LLMs)? AKA will LLMs replace knowledge graphs?. InProceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 311–325
2024
-
[49]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language models don’t always say what they think: Unfaithful explanations in chain-of- thought prompting.Advances in Neural Information Processing Systems36 (2023), 74952–74965
2023
-
[50]
Jianwei Wang, Kai Wang, Ying Zhang, Wenjie Zhang, Xiwei Xu, and Xuemin Lin. 2025. On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing.Proc. VLDB Endow.18, 10 (June 2025), 3421–3434. https: //doi.org/10.14778/3748191.3748205
arXiv 2025
-
[51]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models.arXiv preprint arXiv:2203.11171(2022)
Pith/arXiv arXiv 2022
-
[52]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reason- ing in large language models.Advances in neural information processing systems 35 (2022), 24824–24837
2022
-
[53]
Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Lingjiao Chen, Dong- mei Zhang, Surajit Chaudhuri, and H. V. Jagadish. 2025. MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Bench- marks Track. https://openreview.net/forum?id=ryUzgwD6UQ
2025
-
[54]
Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Dongmei Zhang, and Surajit Chaudhuri. 2024. Table-llm-specialist: Language model special- ists for tables using iterative generator-validator fine-tuning.arXiv preprint arXiv:2410.12164(2024)
arXiv 2024
-
[55]
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms.arXiv preprint arXiv:2306.13063(2023)
Pith/arXiv arXiv 2023
-
[56]
Mohamed Yakout, Laure Berti-Équille, and Ahmed K Elmagarmid. 2013. Don’t be scared: use scalable automatic repairing with maximal likelihood and bounded changes. InProceedings of the 2013 ACM SIGMOD International Conference on Management of Data. 553–564
2013
-
[57]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)
Pith/arXiv arXiv 2025
-
[58]
Chenyu Yang, Yuyu Luo, Chuanxuan Cui, Ju Fan, Chengliang Chai, and Nan Tang
-
[59]
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2024. Alignment for honesty.Advances in Neural Information Processing Systems37 (2024), 63565–63598
2024
-
[60]
Ioannidis, and Huzefa Rangwala
Yuqing Yang, Qi Zhu, Zhen Han, Boran Han, Zhengyuan Shen, Shuai Wang, Vassilis N. Ioannidis, and Huzefa Rangwala. 2026. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United ...
2026
-
[61]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InInternational Conference on Learning Representations (ICLR)
2023
-
[62]
VLDB Endow.18, 10 (June 2025), 3354–3367
Data Imputation with Limited Data Redundancy Using Data Lakes.Proc. VLDB Endow.18, 10 (June 2025), 3354–3367. https://doi.org/10.14778/3748191. 3748200
doi:10.14778/3748191 2025
-
[63]
Bianca Zadrozny and Charles Elkan. 2002. Transforming classifier scores into ac- curate multiclass probability estimates. InProceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining. 694–699
2002
-
[64]
Haochen Zhang, Yuyang Dong, Chuan Xiao, and Masafumi Oyamada. 2024. Jellyfish: Instruction-tuning local large language models for data preprocessing. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 8754–8782
2024
-
[65]
Mozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo, Chong Peng, Peng Yan, Yaqian Zhou, and Xipeng Qiu. 2024. Calibrating the confidence of large language models by eliciting fidelity. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2959–2979
2024
-
[66]
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. 2020. Gradient surgery for multi-task learning.Advances in neural information processing systems33 (2020), 5824–5836
2020
-
[67]
Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, and Yingda Chen. 2024. SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning. arXiv:2408.05517 [cs.CL] https://arxiv.org/abs/2408.05517
Pith/arXiv arXiv 2024
-
[68]
Wei Zhou, Jun Zhou, Haoyu Wang, Zhenghao Li, Qikang He, Shaokun Han, Guoliang Li, Xuanhe Zhou, Yeye He, Chunwei Liu, et al. 2026. Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs. arXiv preprint arXiv:2601.17058(2026). Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models Table 8: R@P...
arXiv 2026
-
[70]
Tianshu Zhang, Xiang Yue, Yifei Li, and Huan Sun. 2024. Tablellama: Towards open large generalist models for tables. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6024–6044
2024
-
[443]
Let me do the math: 443 - 270 = 173. So the missing UnitCount is 173. Let me check the other entries to confirm. BLI3440 only has one entry with SummedUnitCount 235, which matches the single UnitCount. So the pattern holds. Therefore, the missing value is 173. </think> {"value": "173", "confidence": 98} Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Ha...
arXiv 1997
-
[2010]
Pattern classification with missing data: a review.Neural Computing and Applications19, 2 (2010), 263–282
2010
-
[2015]
InProceedings of the companion publication of the 2015 annual conference on genetic and evolutionary computation
An evolutionary missing data imputation method for pattern classification. InProceedings of the companion publication of the 2015 annual conference on genetic and evolutionary computation. 1013–1019
2015
-
[2025]
InFindings of the Association for Computational Linguistics: ACL 2025
LLM-forest: Ensemble learning of LLMs with graph-augmented prompts for data imputation. InFindings of the Association for Computational Linguistics: ACL 2025. 6921–6936
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.