Tutorial · 15 min · intermediate
On-device text-to-speech
Neural TTS running entirely on the phone — SuperTonic-3 over ONNX, 31 languages and 10 voices, streamed chunk-by-chunk into the platform's audio engine.
What you'll build — running on device.
No cloud, no API keys, no per-character billing: dartnative_supertonic_tts
runs neural speech synthesis entirely on the device — a pure-Dart, all-ONNX
four-model pipeline on dartnative_onnxruntime, with dartnative_audio
playing the raw PCM. You’ll build a “read it aloud” screen with a
full-screen 31-language picker, ten voices, a quality dial, and streaming
synthesis that starts speaking before it’s finished thinking. It’s the
DartNative playground’s SuperTonic demo, carried over byte-identical.
What you need
- A project from Your first DartNative app
dartnative_supertonic_tts,dartnative_audio,dartnative_onnxruntime, anddartnative_path_providerin your pubspec- iOS deployment target 16.0+ (
platform :ios, '16.0'in the Podfile + the Xcode project) —dartnative_onnxruntime’s minimum - Android minSdk 26 (
android/app/build.gradle.kts) —dartnative_path_provider’s minimum
Step 1 — The model, on demand
SuperTonic’s model is ~144 MB — you don’t ship that in your binary; you
download it on first use and cache it. ModelManager owns the files, the
engine runs them:
// bundledDir points at model files shipped in flutter_assets, if any;
// null (nothing bundled) → everything downloads on first use.
final ModelManager _manager = ModelManager(bundledDir: _resolveBundledDir());
_installed = _manager.isReady(); // already on disk?
// _manager.pendingDownloadMB() // for the "Download (~144 MB)" label
final tts = SuperTonicTTS.withManager(_manager);
await tts.initialize(
intraOpNumThreads: 2, // leave CPU headroom for the UI thread
onProgress: (p, msg) {
if (mounted) setState(() => _progress = p);
},
);
initialize() downloads whatever is missing (from HuggingFace, with
onProgress feeding your progress UI), then loads the four ONNX sessions.
Every run after that is fully offline. (The demo’s _resolveBundledDir
also lets a maintainer pre-bundle model files under assets/supertonic —
when nothing is there, everything simply downloads.) The tutorial’s action
button does double duty: Download until the model is on disk, Speak
after — and a delete control (_manager.delete(), behind a native
showAlert confirm) reclaims the space.
Step 2 — Stream, don’t wait
final _player = PcmStreamPlayer(); // configured lazily, on the first chunk:
// _player.configure(sampleRate: tts.sampleRate, channels: 1, bitsPerSample: 16)
await for (final Float32List audio in tts.generateStream(
text,
voice: _voice, // F1–F5 female, M1–M5 male
lang: _lang, // 31 languages — see TTSLanguage
speed: _speed, // 0.5–2.0×
steps: _steps, // quality dial: denoising steps, 2–16
)) {
_player.feedChunk(_float32ToPcm16(audio));
}
This is the key design decision: generate() would return one clip after
the whole text is synthesized, so the silence before the first word grows
with length. generateStream() splits the text into sentence-aligned
chunks and yields each one’s PCM the moment it’s ready — chunk 0 is already
sounding while the rest is still being generated. Time-to-first-audio stays
flat no matter how long the text is. (The tutorial also feeds a 0.12 s
silent gap between chunks — each chunk is trimmed to its exact duration, so
without it they’d butt together.)
The model emits float32 samples; the player takes 16-bit PCM — the conversion is ten lines:
Uint8List _float32ToPcm16(Float32List s) {
final out = Uint8List(s.length * 2);
final view = ByteData.view(out.buffer);
for (var i = 0; i < s.length; i++) {
final v = s[i] < -1 ? -1.0 : (s[i] > 1 ? 1.0 : s[i]);
view.setInt16(i * 2, (v * 32767).round(), Endian.little);
}
return out;
}
Call _player.flush() before a new utterance — it cuts anything still
sounding from the previous one.
Step 3 — Languages, voices, quality
Thirty-one languages are too many for chips, so the Language row is a
disclosure that pushes a full-screen picker — LanguagePickerScreen lists
every language SuperTonic speaks (TTSLanguage.all), native name with the
English name underneath and the current one check-marked, and pops with the
chosen code:
final code = await Navigator.push<String>(
context,
PageRoute(builder: (_) => LanguagePickerScreen(selectedCode: _lang)),
);
if (code != null && code != _lang) setState(() => _lang = code);
Voices stay chips, and every language ships a built-in sample string:
for (final v in supertonicVoices) ... // F1…F5, M1…M5 as chips
TTSLanguage.fromCode(_lang).nativeName // 'it' → 'Italiano'
TTSTestStrings.shortForLanguage(_lang) // a built-in sample per language
Everything is a per-utterance parameter — no re-initialization when the
user switches language, voice, speed, or the steps quality dial (2–16;
8 is the medium preset; higher is slower and cleaner). The tokenizer is
character-level — no per-language phonemizer or data files — and it
auto-inserts lexical-stress accents for words a character-level model would
otherwise mis-stress (Italian sàbato, perdòno).
Why this is native
The inference runs through
dartnative_onnxruntimeon the platform’s ML stack, and the audio path is the platform’s own engine fed raw PCM over FFI. Nothing crosses a network; nothing crosses a MethodChannel. A phone synthesizing natural speech in 31 languages by itself was science fiction a few years ago — now it’s a pubspec entry.
The finished code
In the public repo — dn create ., dn run. The demo screen and its
language picker are byte-identical copies of the playground’s files; only
the thin main.dart entry is tutorial-specific. When the playground screen
improves, this tutorial inherits it verbatim.