🏆 Vietnamese Open ASR Leaderboard

📢 About this Leaderboard

This benchmark is designed to automatically evaluate and rank Automatic Speech Recognition (ASR) models for the Vietnamese language based on the Word Error Rate (WER) metric.

  • How to Submit: To learn more about the submission format, requirements, and how to structure your result.zip file, please read the README.md file located in the Files tab of this Space.

Models O-WER N-WER
Synthetic natural Synthetic natural
VIVOS Common Voice AVG viVoice vlsp ViMD LSVSC FOSD BUD500 GIGA-SPEECH AVG VIVOS Common Voice AVG viVoice vlsp ViMD LSVSC FOSD BUD500 GIGA-SPEECH AVG
vinai/PhoWhisper-medium 14.07 21.89 17.98 17.56 🥇 8.49 19.71 16.75 22.11 19.18 🥇 13.31 🥇 16.73 🥇 1.29 7.41 4.35 7.43 🥇 2.40 11.74 7.23 6.07 8.59 9.84 7.61
vinai/PhoWhisper-small 14.37 23.58 18.97 18.72 9.15 22.13 17.48 22.12 19.53 14.96 17.73 1.64 10.00 5.82 9.08 3.16 14.56 8.12 🥇 6.01 8.83 11.65 8.77
Qwen/Qwen3-ASR-1.7B 15.72 25.18 20.45 🥇 13.91 19.17 19.53 15.91 29.94 17.13 15.76 18.76 6.11 8.77 7.44 5.58 9.88 11.97 7.66 20.65 4.89 8.24 9.84
Qwen/Qwen3-ASR-0.6B 24.36 22.84 23.60 21.83 20.91 23.37 18.92 33.27 10.84 17.31 20.92 8.54 13.66 11.10 7.11 12.61 14.65 8.61 26.69 6.78 10.60 12.44
nguyenvulebinh/wav2vec2-base-vietnamese-250h 26.75 38.90 32.83 27.26 22.40 27.50 21.63 39.65 8.82 22.58 24.26 9.62 17.40 13.51 13.84 14.08 19.27 10.50 22.53 7.56 16.45 14.89
nvidia/parakeet-ctc-0.6b-Vietnamese 🥇 11.74 🥇 15.05 🥇 13.39 27.82 18.03 19.79 🥇 14.12 🥇 16.47 30.90 18.36 20.78 6.60 11.04 8.82 27.13 8.42 13.94 7.56 12.82 7.95 11.85 12.81
hynt/Zipformer-30M-RNNT-6000h 21.85 29.29 25.57 20.78 16.40 🥇 18.00 18.19 32.36 🥇 3.29 13.52 17.51 3.61 🥇 3.97 🥇 3.79 🥇 4.71 7.26 🥇 7.91 🥇 6.23 12.84 🥇 1.93 🥇 6.54 🥇 6.77