NVIDIA unveils kumo tabular: open-source tabular models for efficient predictive analysis
NVIDIA has introduced Kumo Tabular, an open-source family of tabular foundation models (TFMs) designed for classification and regression tasks. These models predict new rows in a single forward pass without requiring training, hyperparameter tuning, or feature engineering. Available in Small, Medium, and Large versions, Kumo Tabular ranges from 28M to 215M parameters and runs through NVIDIA’s structured-data-models (SDM) library. The library also includes TabICLv2, Google’s TabFM, and KumoRelational for multi-table data. All models share an in-context learning interface built on a TableTensor container, supporting preprocessing, ensembling, and multi-class prediction.
Kumo Tabular, a Transformer-based model, uses column, row, and in-context attention mechanisms. It scales query attention with a temperature parameter that grows logarithmically with the number of keys, maintaining sharp attention on larger tables. Pretrained on synthetic tables generated from Structural Causal Models (SCMs), the models incorporate real-world patterns like missing values and conflicting rows. NVIDIA reports Kumo Tabular ranks first on TabArena with an Elo score of 1950 and runs 17x faster than LimiX-2 on an RTX 6000 Pro. The models are available under permissive licenses, making them accessible for commercial use.