ChorusReader
Every character gets a voice
Takes your books and gives every character a distinct voice, entirely on-device. Three engine tiers: Standard (50+ voices, 8 languages, free), Premium (per-character multi-voice), and Personal — clone any voice from 15 seconds of audio. Version 2.2 added a built-in store: 100+ hand-picked free classics, a daily shelf of open-access research papers, and your RSS feeds, one tap from your library.
Why multi-voice
Audiobook narration is a solved problem if you're fine with one voice reading everything. But characters have voices. Dialogue has emotion. ChorusReader assigns distinct voices to characters and modulates emotion per scene — happy, sad, angry, contemplative. All running on your iPhone, no cloud API calls, no per-character fees.
Capabilities
- Three engines — Standard: 50+ voices across 8 languages, free. Premium: per-character multi-voice narration. Personal: clone any voice from 15 seconds of audio, entirely on-device
- Built-in bookstore — 100+ hand-picked free classics with narrated samples, a daily list of open-access research papers with the reason each was chosen, and your RSS feeds
- A library that looks like one — redesigned bookcase with covers, shelves for Books, Papers, Imports, and Feeds
- 12 emotion categories for expressive narration; ePub and PDF support with chapter detection
- Real-time streaming synthesis — playback position saved continuously, interrupted narrations resume exactly where they stopped
- Powered by ChorusTTS — a custom accelerated pipeline built on Chatterbox weights (thanks Resemble AI)
Demo
The TTS pipeline
ChorusTTS is a custom Swift pipeline built on Chatterbox weights by Resemble AI. Three stages turn text into speech — each optimized for Apple Silicon and quantized to fit on a phone.
book analysis: Quote detection (heuristics + regex), then Character attribution (on-device LLM + turn tracking), then Emotion enrichment (per-segment profiles). synthesis: T3 transformer (text to speech tokens), then S3Gen (flow matching, tokens to mel), then HiFT vocoder (mel to 24 kHz audio), then Streaming playback (plays as chunks arrive, resumes from checkpoints)
- iOS: 4-bit quantized, compiled to CoreML on first launch.
- macOS: 8-bit quantized, native MLX.
Key decisions
-
Hybrid book analysis
regex-based quote detection feeds into on-device LLM (Foundation Models) for speaker attribution and emotion tagging
-
Each voice is a set of speaker embeddings, not a separate model
switching voices is instant, costs no additional memory
-
12 emotion profiles map to concrete TTS parameters (temperature, guidance weight, pacing)
the same text sounds different when a character whispers vs. shouts
-
Streaming synthesis with checkpoint resume
if the app is interrupted mid-generation, it picks up exactly where it left off
-
Voice cloning from a few seconds of reference audio via an LSTM encoder
bring your own narrator