Home/Works/From-scratch Japanese TTS
Text-to-Speech · Former CTO

From-scratch
Japanese TTS

As Livetoon's CTO, I designed and trained Japanese speech synthesis from the acoustic model down to the vocoder, from scratch. Not stitching existing synthesizers together — building it from the root ourselves. It's a field only a handful of companies in Japan, the big names included, attempt in-house, and I led it as the lead developer of a four-person team.

CTO · 2024–2025 Lead of a 4-person team Acoustic model → vocoder 120ms on a T4
Livetoon TTS
Japanese speech synthesis designed from the root, from the acoustic model down to the vocoder.
120ms
Short-form synthesis (measured on a T4)
a few
Companies in Japan that do this in-house
4people
Team I led as lead developer
'24–'25
As Livetoon's CTO
What

What I built

Turning text into speech is roughly two stages: an acoustic model that produces sound features from the text, and a vocoder that turns those features into an actual waveform. The usual setup uses an off-the-shelf part for one of them. At Livetoon we designed and trained both from scratch — because I didn't want the naturalness, or the latency, tied to someone else's choices.

Only a handful of companies in Japan, the big names included, can hold their own from-scratch speech synthesis. We took it on with a small four-person team. I worked as the lead developer, hands on everything from the architecture to how training was run. Image, speech, language — getting to experience the major generative modalities from the "training" side owes a lot to this period.

Used by

Used in the field

In October 2025, this TTS was used for the read-aloud in Fujitsu's healthcare AI-agent proof-of-concept. Someone on their side said the read-aloud "had surprisingly little of the awkwardness you usually get from synthesized speech." More than any self-reported number, that one line from someone who actually used it made the work feel worth it.

Coverage

Where it was covered

  • AWS Startup Blog. A technical interview about building Japanese speech synthesis from scratch.
  • NVIDIA-hosted webinar. I spoke on "Realizing digital health with agentic AI" (November 2025).
Related

Related

← Back to Works