TabFM introduces a 400‑million‑parameter foundation model that treats supervised tabular prediction as an in‑context learning problem, enabling calibrated zero‑shot forecasts in a single forward pass without any task‑specific fine‑tuning. The model is trained exclusively on synthetic tables generated from structural causal models, which forces it to capture generalizable tabular patterns. Evaluation spans the full TabArena suite comprising 51 datasets—38 classification and 13 regression—where zero‑shot TabFM achieves the top rank among default tabular foundation models and surpasses the performance of tuned AutoML pipelines. Two post‑training enhancements built on the same frozen weights further boost results: TabFM+ applies multi‑view feature expansion coupled with ensemble averaging and a post‑hoc calibration step, while TabFM‑Auto leverages a large language model to guide dataset‑specific preprocessing and feature engineering. These extensions improve both classification and regression metrics across the benchmark, demonstrating that a foundation model trained on synthetic causal data can deliver strong, transferable tabular intelligence without per‑dataset retraining.

Read original