Tutorial · 15 min · intermediate

On-device text-to-speech

Neural TTS running entirely on the phone — SuperTonic-3 over ONNX, 31 languages and 10 voices, streamed chunk-by-chunk into the platform's audio engine.

What you'll build — running on device.

No cloud, no API keys, no per-character billing: dartnative_supertonic_tts runs neural speech synthesis entirely on the device — a pure-Dart, all-ONNX four-model pipeline on dartnative_onnxruntime, with dartnative_audio playing the raw PCM. You’ll build a “read it aloud” screen with a full-screen 31-language picker, ten voices, a quality dial, and streaming synthesis that starts speaking before it’s finished thinking. It’s the DartNative playground’s SuperTonic demo, carried over byte-identical.

What you need

  • A project from Your first DartNative app
  • dartnative_supertonic_tts, dartnative_audio, dartnative_onnxruntime, and dartnative_path_provider in your pubspec
  • iOS deployment target 16.0+ (platform :ios, '16.0' in the Podfile + the Xcode project) — dartnative_onnxruntime’s minimum
  • Android minSdk 26 (android/app/build.gradle.kts) — dartnative_path_provider’s minimum

Step 1 — The model, on demand

SuperTonic’s model is ~144 MB — you don’t ship that in your binary; you download it on first use and cache it. ModelManager owns the files, the engine runs them:

// bundledDir points at model files shipped in flutter_assets, if any;
// null (nothing bundled) → everything downloads on first use.
final ModelManager _manager = ModelManager(bundledDir: _resolveBundledDir());
_installed = _manager.isReady();             // already on disk?
// _manager.pendingDownloadMB()              // for the "Download (~144 MB)" label

final tts = SuperTonicTTS.withManager(_manager);
await tts.initialize(
  intraOpNumThreads: 2, // leave CPU headroom for the UI thread
  onProgress: (p, msg) {
    if (mounted) setState(() => _progress = p);
  },
);

initialize() downloads whatever is missing (from HuggingFace, with onProgress feeding your progress UI), then loads the four ONNX sessions. Every run after that is fully offline. (The demo’s _resolveBundledDir also lets a maintainer pre-bundle model files under assets/supertonic — when nothing is there, everything simply downloads.) The tutorial’s action button does double duty: Download until the model is on disk, Speak after — and a delete control (_manager.delete(), behind a native showAlert confirm) reclaims the space.

Step 2 — Stream, don’t wait

final _player = PcmStreamPlayer(); // configured lazily, on the first chunk:
// _player.configure(sampleRate: tts.sampleRate, channels: 1, bitsPerSample: 16)

await for (final Float32List audio in tts.generateStream(
  text,
  voice: _voice,  // F1–F5 female, M1–M5 male
  lang: _lang,    // 31 languages — see TTSLanguage
  speed: _speed,  // 0.5–2.0×
  steps: _steps,  // quality dial: denoising steps, 2–16
)) {
  _player.feedChunk(_float32ToPcm16(audio));
}

This is the key design decision: generate() would return one clip after the whole text is synthesized, so the silence before the first word grows with length. generateStream() splits the text into sentence-aligned chunks and yields each one’s PCM the moment it’s ready — chunk 0 is already sounding while the rest is still being generated. Time-to-first-audio stays flat no matter how long the text is. (The tutorial also feeds a 0.12 s silent gap between chunks — each chunk is trimmed to its exact duration, so without it they’d butt together.)

The model emits float32 samples; the player takes 16-bit PCM — the conversion is ten lines:

Uint8List _float32ToPcm16(Float32List s) {
  final out = Uint8List(s.length * 2);
  final view = ByteData.view(out.buffer);
  for (var i = 0; i < s.length; i++) {
    final v = s[i] < -1 ? -1.0 : (s[i] > 1 ? 1.0 : s[i]);
    view.setInt16(i * 2, (v * 32767).round(), Endian.little);
  }
  return out;
}

Call _player.flush() before a new utterance — it cuts anything still sounding from the previous one.

Step 3 — Languages, voices, quality

Thirty-one languages are too many for chips, so the Language row is a disclosure that pushes a full-screen picker — LanguagePickerScreen lists every language SuperTonic speaks (TTSLanguage.all), native name with the English name underneath and the current one check-marked, and pops with the chosen code:

final code = await Navigator.push<String>(
  context,
  PageRoute(builder: (_) => LanguagePickerScreen(selectedCode: _lang)),
);
if (code != null && code != _lang) setState(() => _lang = code);

Voices stay chips, and every language ships a built-in sample string:

for (final v in supertonicVoices) ...        // F1…F5, M1…M5 as chips
TTSLanguage.fromCode(_lang).nativeName       // 'it' → 'Italiano'
TTSTestStrings.shortForLanguage(_lang)       // a built-in sample per language

Everything is a per-utterance parameter — no re-initialization when the user switches language, voice, speed, or the steps quality dial (2–16; 8 is the medium preset; higher is slower and cleaner). The tokenizer is character-level — no per-language phonemizer or data files — and it auto-inserts lexical-stress accents for words a character-level model would otherwise mis-stress (Italian sàbato, perdòno).

Why this is native

The inference runs through dartnative_onnxruntime on the platform’s ML stack, and the audio path is the platform’s own engine fed raw PCM over FFI. Nothing crosses a network; nothing crosses a MethodChannel. A phone synthesizing natural speech in 31 languages by itself was science fiction a few years ago — now it’s a pubspec entry.

The finished code

In the public repo — dn create ., dn run. The demo screen and its language picker are byte-identical copies of the playground’s files; only the thin main.dart entry is tutorial-specific. When the playground screen improves, this tutorial inherits it verbatim.

Open the finished code →