Skip to content
Parishruthi Ganesh

    to navigate · to open · Esc to close

    Preprint 2026 · arXiv preprint

    Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

    Parishruthi Ganesh , Gerry Dozier , Cheryl Seals

    Preprint available on arXiv; submitted to AAAI.

    Abstract

    A systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M–9B parameter range, across eight English single-label intent-classification datasets covering standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy the study analyses confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models.

    Research areas

    BibTeX

    @misc{ganesh2026selecting,
      title         = {Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models},
      author        = {Ganesh, Parishruthi and Dozier, Gerry and Seals, Cheryl},
      year          = {2026},
      eprint        = {2607.27421},
      archivePrefix = {arXiv},
      primaryClass  = {cs.CL},
      url           = {https://arxiv.org/abs/2607.27421}
    }