The paper introduces activation alignment, a technique that trains a lightweight linear map on synthetic unlabeled data to align the intermediate activations of a context‑constrained student model with those of a full‑context teacher model. The student is built from a tabular foundation model (TabPFN‑3 or TabFM) that receives only a subset of training examples as context, while the teacher processes the complete dataset. By minimizing the difference between the teacher’s and student’s activation vectors, the aligner learns a simple transformation that can be applied at inference time, requiring no GPU and converging within seconds to minutes on standard hardware. Experiments evaluate the approach on 38 classification tasks from the TabArena benchmark, using both TabPFN‑3 and TabFM. Across all context‑budget settings, the aligned students show statistically significant performance gains over unaligned baselines, and in low‑data regimes they recover nearly 50 % of the teacher’s predictive advantage. This method thus enables the inference speed of compact contexts while substantially narrowing the performance gap to full‑context training.
Read original
huggingface/daily-papers