Show HN: I trained a 9M speech model to fix my Mandarin tones
A developer has engineered a specialized 9-million parameter speech model aimed at assisting Mandarin learners in precisely correcting their pronunciation tones. This innovative project directly tackles the pervasive challenge non-native speakers face in reliably identifying and self-correcting tonal errors in spoken Mandarin. The underlying architecture is a Conformer-CTC model, rigorously trained on an extensive dataset of approximately 300 hours, incorporating data from both AISHELL and Primewords corpora. To ensure broad accessibility and efficient operation, the model has been highly optimized through INT8 quantization, reducing its size to a mere 11 MB, which allows it to execute entirely within a web browser utilizing ONNX Runtime Web. Functionally, it excels at grading per-syllable pronunciation and tones by employing Viterbi forced alignment, delivering immediate and objective feedback. This accessible, in-browser artificial intelligence solution represents a significant advancement for individuals dedicated to mastering Mandarin's complex tonal system, offering a robust tool for self-improvement.