New Benchmark: Muon Beats AdamW for Tabular Deep Learning, at Higher Cost
October 5, 2026 3 min read
On April 16, a preprint landed on arXiv reporting what its authors call the missing piece in tabular deep learning: a systematic comparison of optimizers. Authored by Yury Gorishniy, Ivan Rubachev, Dmitrii Feoktistov and Artem Babenko, the paper benchmarks 15 optimization methods across 17 supervised tabular-learning tasks. Architecture design has received years of attention in this field; optimizer choice, the paper argues, was simply inherited from other domains, with AdamW as the unquestioned default.
What does it establish?
Within the tested envelope, Muon performed best. It led on both a plain ReLU MLP and on stronger MLP-based architectures, including models with piecewise-linear feature embeddings and the TabM ensemble family, and the authors recommend it as a powerful modern baseline for practice and research, with one condition attached: the extra training cost must be affordable. A second, cheaper finding: maintaining an exponential moving average of model weights during training improved AdamW on vanilla MLPs, though the benefit was less consistent across the advanced variants. Schedule-free AdamW emerged as another notable runner-up.
How did they test it?
Rather than dropping each optimizer into fixed settings, the team jointly tuned model and optimizer hyperparameters for every method, so each optimizer was judged at its best configuration rather than its defaults. Training followed the standard supervised regime the authors stress makes tabular data distinct: finite, noisy datasets, early stopping, and evaluation on held-out data instead of chasing faster loss reduction. The code has been released publicly under the yandex-research GitHub organization.
Where do the results stop?
The exclusions are stated plainly in the paper: tabular foundation models, retrieval-based methods, and non-MLP paradigms are out of scope, where optimizer behavior may differ. The results are also purely empirical; the paper offers no account of why Muon works in this setting and names that question as future work. The cost side matters too. In the reported experiment, tuning Muon took roughly three times as long as tuning AdamW, and the authors concede that training is slower overall. Equally important is what the paper does not settle: whether a well-tuned AdamW with weight averaging closes the gap when both methods receive the same compute budget. That matched-cost boundary remains open. The reviewed source is the author-submitted arXiv v1 preprint.
Should you switch?
If you train MLP-based models from scratch on supervised tabular data and accuracy outweighs training time, Muon now has direct evidence behind it and deserves a place in your baseline tests. If you want a low-effort gain on a plain MLP, adding weight averaging to an existing AdamW pipeline is the paper's suggested shortcut. If your stack involves tabular foundation models, retrieval, or architectures outside the MLP family, this paper does not speak to your setting — the old default stands until someone benchmarks it.
Sources
- Benchmarking Optimizers for MLPs in Tabular Deep Learning — arXiv (author-submitted research)