yousan

← Back to Works

kodama-ja-streaming-small

A Japanese streaming ASR model adapted from the English streaming model moonshine-streaming-small.

Date
Tech
ASR, Streaming, ONNX, ReazonSpeech

Overview

I adapted the English streaming speech recognition model moonshine-streaming-small (about 140M parameters) to Japanese and published it on Hugging Face as the Japanese streaming ASR model kodama-ja-streaming-small. It ships as ONNX / ORT split into three graphs (encoder, cross_kv_prefill and decoder) under the Apache-2.0 license.

What I did

I trained the Japanese adaptation on the full ReazonSpeech v2 dataset (about 35,000 hours) for 2 epochs, then handled the ONNX conversion, evaluation and release.

Technical notes

The write-up records an offline RTF of 0.19 on CPU and about 1.1 seconds to the first partial result. See the article for the full evaluation.

Links