NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass

NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If you have followed TabPFN or TabICL, the setup will look familiar. The model takes labeled rows as context and predicts new rows in one forward pass. There is no training, no hyperparameter tuning, and no feature engineering.

Kumo Tabular comes in Small, Medium, and Large versions, spanning about 28M to 215M parameters. It runs through NVIDIA’s open-source structured-data-models (SDM) library.

Is it deployable? Yes. Weights ship under the OpenMDW-1.1 license, which permits commercial use. The SDM code is Apache-2.0, and it needs Python 3.11+ and PyTorch 2.7+, with examples targeting a CUDA GPU.

What the SDM Library Adds

SDM is a GPU-native library for structured-data foundation models and preprocessing. Besides Kumo Tabular, it ships TabICLv2, Google’s TabFM, and KumoRelational for multi-table data. All models share one in-context learning interface built on a TableTensor container. The library also handles preprocessing, ensembling, and many-class prediction.

How Kumo Tabular Works

Kumo Tabular is a Transformer built around the structure of a table. It uses column, row, and in-context attention, as introduced in TabICL and TabPFN. The pipeline has 3 stages:

Cell embedding: Numerical and categorical values pass through learned Fourier features, with separate weights per type. Missing values need no imputation.

Row embedding: Column attention uses induced self-attention, so cost grows linearly with rows. Row attention, with rotary positions, learns feature interactions. 4 learnable [CLS] tokens compress each row.

In-context learning: A final Transformer runs over row embeddings. Context rows attend to each other, while query rows attend only to context rows.

Because the context never sees the queries, its keys and values are computed once and reused. The head outputs class probabilities, or 999 quantiles for regression. That gives a point prediction plus an uncertainty estimate.

One more detail matters at scale. Softmax attention spreads thin as the number of keys grows. Kumo Tabular scales each query by a temperature that grows with the log of the key count. The coefficient is learned per attention head, so attention stays sharp on larger tables.