yousan

← Back to Works

Fine-tuning J-Moshi-ext

Adding a specific voice and speaking style to J-Moshi-ext, a Japanese full-duplex dialogue model.

Date
Tech
Full-duplex dialogue, J-Moshi-ext, DeepSpeed, Fine-tuning

Overview

An experiment adding a voice and a speaking style to J-Moshi-ext (7.5B), a Japanese full-duplex dialogue model. I first fine-tuned a TTS model to generate synthetic dialogues, then trained J-Moshi-ext on them. After the voice stage, a second stage added an “ojousama” (young lady) speaking style.

What I did

I generated the synthetic dialogue data, ran the training with DeepSpeed ZeRO-3, evaluated the results and wrote them up.

Technical notes

The voice stage only partly matched the target voice and reduced pronunciation clarity; the article records this candidly.

Links