Benchmark test: TabPFN and TabICL beat tuned XGBoost on all 14 tabular datasets
An independent test compared pretrained tabular foundation models TabPFN and TabICL, which make predictions without training on new data, against a hyperparameter-tuned XGBoost model. Across 14 datasets from the Grinsztajn benchmark, using the same data splits and timing for all methods, the non-training models outperformed tuned XGBoost on every single dataset, including at scales up to 32,000 rows.
GoKawiil's interpretation of the reporting above, not reported fact.
If these results generalize, they could challenge the standard machine learning workflow for tabular data, where hyperparameter search has long been considered essential for top performance. The tester suggests this could make exhaustive tuning a discretionary step rather than a requirement, though the claim is based on one independent benchmark rather than broad industry validation. The piece also notes that TabPFN, the most-cited such model, now requires an account to download, which may affect its accessibility for reproducibility.
- TabPFN and TabICL, which predict without training on new data, outperformed tuned XGBoost across all 14 tested datasets.
- The performance advantage reportedly held even as dataset size scaled up to 32,000 rows.
- Access to TabPFN, described as the most-cited model in this category, now requires creating an account to download it.
Source: efraingaray.com, 2026-09-28
Published there as: “TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.