Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
Parishruthi Ganesh , Gerry Dozier , Cheryl Seals
Preprint available on arXiv; submitted to AAAI.
A systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M–9B parameter range, across eight English single-label intent-classification datasets covering standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy the study analyses confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models.