kodama-ja-streaming-small
A Japanese streaming ASR model adapted from the English streaming model moonshine-streaming-small.
Overview
I adapted the English streaming speech recognition model moonshine-streaming-small (about 140M parameters) to Japanese and published it on Hugging Face as the Japanese streaming ASR model kodama-ja-streaming-small. It ships as ONNX / ORT split into three graphs (encoder, cross_kv_prefill and decoder) under the Apache-2.0 license.
What I did
I trained the Japanese adaptation on the full ReazonSpeech v2 dataset (about 35,000 hours) for 2 epochs, then handled the ONNX conversion, evaluation and release.
Technical notes
The write-up records an offline RTF of 0.19 on CPU and about 1.1 seconds to the first partial result. See the article for the full evaluation.

