Fine-tuning J-Moshi-ext
Adding a specific voice and speaking style to J-Moshi-ext, a Japanese full-duplex dialogue model.
Overview
An experiment adding a voice and a speaking style to J-Moshi-ext (7.5B), a Japanese full-duplex dialogue model. I first fine-tuned a TTS model to generate synthetic dialogues, then trained J-Moshi-ext on them. After the voice stage, a second stage added an “ojousama” (young lady) speaking style.
What I did
I generated the synthetic dialogue data, ran the training with DeepSpeed ZeRO-3, evaluated the results and wrote them up.
Technical notes
The voice stage only partly matched the target voice and reduced pronunciation clarity; the article records this candidly.


