{"id":"1b6236d2-e444-4e6a-837f-74055374a820","arxiv_id":"2606.03681","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A speedrun benchmark for nanoTabPFN pretraining reports a record of 0.92 minutes to target performance, an 81x speedup over the 74.32-minute baseline using 22x fewer synthetic datasets.","lead":"The paper introduces a community speedrun for pretraining nanoTabPFN where contributors edit one script to hit a fixed ROC AUC target faster on subsampled TabArena with one L40S GPU. A smart generalist might read it to see how standardized benchmarks can let many people stack small efficiency gains in AI training.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Fixed ROC AUC target on subsampled TabArena may not proxy pretraining quality stably under arbitrary script changes","rationale":"The reader's weakest assumption directly identifies the same load-bearing point. Because the supplied abstract supplies no further evidence that the metric is robust to arbitrary modifications, the claim remains conditional on that verification step. No other internal inconsistency is visible from the given text.","tokens_in":1641,"tokens_out":348,"duration_ms":11670,"concrete_test":"Take the winning 0.92-minute checkpoint and the 74.32-minute baseline; retrain both from the same random seed on a fresh 50 % subsample of TabArena drawn with a different seed, then measure ROC AUC on the original speedrun target plus two held-out TabArena tasks. If the relative ordering or the effective speedup reverses by more than 20 %, the proxy is unstable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline claim is an 81x wall-clock speedup (0.92 min vs 74.32 min baseline) to reach a fixed downstream ROC AUC while using 22x fewer synthetic datasets. This only demonstrates a genuine pretraining improvement if the chosen target remains a stable, comparable measure of model quality when the single-file script is modified arbitrarily. Because any change to data generation, optimization, architecture, or regularization is allowed, it is possible for a record to exploit idiosyncrasies of the particular TabArena subsample, the exact ROC AUC threshold, or the single L40S GPU timing without producing a foundation model that would perform better under a different downstream distribution or evaluation protocol. No independent verification that the proxy holds across modifications is described in the abstract.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces a community 'speedrun' benchmark for pretraining nanoTabPFN tabular foundation models. Participants modify a single-file training script to reach a fixed downstream ROC AUC target on a subsampled TabArena dataset using one NVIDIA L40S GPU. The current record achieves the target in 0.92 minutes (81x speedup over the 74.32-minute baseline) while using 22x fewer synthetic datasets. The format is intended to enable the community to add, verify, and stack pretraining improvements via an open leaderboard, with code available at the provided GitHub link.","tokens_in":1791,"tokens_out":408,"duration_ms":22236,"significance":"If the evaluation protocol proves robust, the speedrun format could meaningfully lower iteration costs for tabular foundation model research by providing a simple, low-resource, community-verifiable benchmark that accumulates incremental gains. The open code and explicit empirical record (wall-clock time and dataset count) are strengths that support reproducibility and stacking of improvements.","major_comments":[{"comment":"Abstract: The headline claim of an 81x speedup and a useful community protocol rests on the assumption that a fixed ROC AUC target on the subsampled TabArena remains a stable, comparable proxy for pretraining quality under arbitrary modifications to data generation, optimization, architecture, or regularization. No analysis, ablation, or verification is provided that this specific target and subsample do not admit exploits of idiosyncrasies (e.g., the exact threshold, GPU timing, or distribution shift) without producing generally better foundation models; this is load-bearing for the central empirical contribution.","section":"Abstract"}],"minor_comments":[{"comment":"The manuscript would benefit from an explicit section detailing the precise measurement protocol, baseline implementation details, dataset subsampling procedure, and any exclusion rules for the speedrun to allow independent reproduction and extension.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for highlighting the importance of validating the evaluation protocol. The concern about the fixed ROC AUC target serving as a robust proxy is substantive and directly relevant to the benchmark's long-term value. We address it point-by-point below and commit to revisions.","responses":[{"response":"We agree that the stability of the chosen target and subsample is load-bearing and that the manuscript provides no explicit ablations or verification against potential exploits. The target was selected in preliminary runs to be reachable by the baseline yet require non-trivial improvements; the subsample size was chosen for computational feasibility on a single L40S. However, we did not test sensitivity to the precise AUC threshold, timing variance across GPU runs, or correlation with performance on held-out datasets or shifted distributions. In revision we will add a dedicated subsection under 'Evaluation Protocol' that (1) reports the exact target selection procedure, (2) includes a small set of sanity checks (re-evaluating the current record on two additional TabArena splits and on a different downstream metric), and (3) explicitly discusses known limitations and the role of the open leaderboard in surfacing future exploits. We view these additions as necessary to support the central claim.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The headline claim of an 81x speedup and a useful community protocol rests on the assumption that a fixed ROC AUC target on the subsampled TabArena remains a stable, comparable proxy for pretraining quality under arbitrary modifications to data generation, optimization, architecture, or regularization. No analysis, ablation, or verification is provided that this specific target and subsample do not admit exploits of idiosyncrasies (e.g., the exact threshold, GPU timing, or distribution shift) without producing generally better foundation models; this is load-bearing for the central empirical contribution."}],"tokens_in":1293,"tokens_out":393,"duration_ms":14311,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper sets up a speedrun format for pretraining nanoTabPFN so the community can compete on reaching a fixed downstream ROC AUC target faster. The single-file script and open leaderboard are the actual new elements, letting contributors modify the code and stack verified improvements.\n\nIt does a clean job keeping the setup minimal and reproducible. The baseline numbers are stated plainly (74.32 minutes, full synthetic data), the record is given (0.92 minutes, 22x fewer datasets), and the GitHub link makes the current implementation easy to inspect. For researchers who need faster iteration cycles on tabular models, this lowers the barrier to testing small changes.\n\nThe main soft spot is the evaluation proxy. A fixed ROC AUC on one subsampled TabArena split with one L40S GPU may not stay stable when arbitrary script changes are allowed. Nothing in the description shows that the target remains a reliable stand-in for general pretraining quality across data generation, optimization, or architecture tweaks. The claim of genuine speedup therefore rests on an untested assumption.\n\nThe work is aimed at tabular ML practitioners who want a shared benchmark rather than a new method or theory. A reader looking for practical protocols will find it useful; someone expecting formal analysis or broad validation will not.\n\nIt deserves peer review because the protocol is concrete, the code is public, and the empirical record can be checked directly by referees.","headline":"The paper introduces a simple community speedrun protocol with a single-file script and leaderboard for tabular foundation model pretraining, reporting an 81x wall-clock record.","tokens_in":2265,"tokens_out":362,"would_cite":false,"duration_ms":19125,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A community speedrun protocol reaches tabular foundation model pretraining targets in 0.92 minutes, an 81x improvement over the baseline.","keywords":["tabular foundation models","pretraining speedups","speedrun","nanoTabPFN","TabArena","ROC AUC target","community leaderboard"],"falsifier":"A script that hits the target faster yet produces models with lower performance on a wider collection of tabular tasks or real datasets outside the speedrun benchmark.","tokens_in":2541,"feed_emoji":"⏱️","tokens_out":629,"duration_ms":27449,"temperature":0.7,"pith_summary":"The paper introduces a speedrun format for nanoTabPFN pretraining in which contributors edit a single-file script to reach a fixed downstream ROC AUC target faster on subsampled TabArena with one L40S GPU. This setup creates an open leaderboard where speedups can be added, verified, and stacked by the community. The current record uses 22x fewer synthetic datasets while cutting time from 74.32 minutes to 0.92 minutes. A sympathetic reader would care because pretraining cost currently limits how often researchers can try new architectures, priors, or optimizers. The protocol aims to shorten that iteration cycle through simple, comparable benchmarks.","feed_headline":"Speedrun cuts tabular pretraining to 0.92 minutes","feed_subtitle":"Contributors edit one script to hit the quality target 81x faster while using 22x less data.","key_machinery":"The speedrun protocol: a fixed downstream performance target, single-file training script, and public leaderboard that enables verification and stacking of pretraining modifications.","core_discovery":"By establishing a speedrun challenge with a fixed ROC AUC target on subsampled TabArena and one NVIDIA L40S GPU, the authors create a standardized protocol that lets participants modify the training script and compete directly on pretraining time, with the best entry currently achieving the target in 0.92 minutes versus the 74.32-minute baseline while requiring 22x fewer synthetic datasets.","pith_inferences":["The format could be adapted to other foundation-model domains where pretraining cost is the main bottleneck.","The winning entry's use of far less data points to data efficiency as a major route to speedups.","Widespread adoption might shift emphasis from scaling compute to measuring and improving training efficiency."],"forward_implications":["New ideas for architectures, priors, or optimization can be tested by submitting modified scripts and measuring time to target.","Successful modifications can be combined over successive leaderboard updates.","Pretraining experiments become feasible on modest hardware, shortening research cycles.","The community gains a shared, auditable record of efficiency gains."],"fun_headline_variants":["Tabular pretraining speedrun hits target in 0.92 minutes","81x speedup in tabular foundation model pretraining","Speedrun achieves tabular pretraining target in 0.92 minutes","Tabular model pretraining 81x faster via speedrun","Speedrun reaches target with 22x fewer synthetic datasets"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That reaching the fixed ROC AUC target on subsampled TabArena reliably indicates overall pretraining quality regardless of how the training script is altered.","fun_headline_variants_meta":{"raw":{"variants":["Tabular pretraining speedrun hits target in 0.92 minutes","81x speedup in tabular foundation model pretraining","Speedrun achieves tabular pretraining target in 0.92 minutes","Tabular model pretraining 81x faster via speedrun","Speedrun reaches target with 22x fewer synthetic datasets"]},"model":"grok-4.3","cost_usd":0.00541,"raw_usage":{"total_tokens":2569,"prompt_tokens":595,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":54099500,"prompt_tokens_details":{"text_tokens":595,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1891,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":595,"tokens_out":83,"duration_ms":12492,"temperature":1.0,"reasoning_tokens":1891,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T11:36:59.220624+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A script that hits the target faster yet produces models with lower performance on a wider collection of tabular tasks or real datasets outside the speedrun benchmark.","supporting_citations":[],"review_version":1}