projects/chorusreader

ChorusReader

Every character gets a voice

Takes your books and gives every character a distinct voice, entirely on-device. Three engine tiers: Standard (50+ voices, 8 languages, free), Premium (per-character multi-voice), and Personal — clone any voice from 15 seconds of audio. Version 2.2 added a built-in store: 100+ hand-picked free classics, a daily shelf of open-access research papers, and your RSS feeds, one tap from your library.

Why multi-voice

Audiobook narration is a solved problem if you're fine with one voice reading everything. But characters have voices. Dialogue has emotion. ChorusReader assigns distinct voices to characters and modulates emotion per scene — happy, sad, angry, contemplative. All running on your iPhone, no cloud API calls, no per-character fees.

Capabilities

  • Three engines — Standard: 50+ voices across 8 languages, free. Premium: per-character multi-voice narration. Personal: clone any voice from 15 seconds of audio, entirely on-device
  • Built-in bookstore — 100+ hand-picked free classics with narrated samples, a daily list of open-access research papers with the reason each was chosen, and your RSS feeds
  • A library that looks like one — redesigned bookcase with covers, shelves for Books, Papers, Imports, and Feeds
  • 12 emotion categories for expressive narration; ePub and PDF support with chapter detection
  • Real-time streaming synthesis — playback position saved continuously, interrupted narrations resume exactly where they stopped
  • Powered by ChorusTTS — a custom accelerated pipeline built on Chatterbox weights (thanks Resemble AI)

Demo

Multi-voice narration — each character gets a distinct voice and emotion.

The TTS pipeline

ChorusTTS is a custom Swift pipeline built on Chatterbox weights by Resemble AI. Three stages turn text into speech — each optimized for Apple Silicon and quantized to fit on a phone.

book analysis: Quote detection (heuristics + regex), then Character attribution (on-device LLM + turn tracking), then Emotion enrichment (per-segment profiles). synthesis: T3 transformer (text to speech tokens), then S3Gen (flow matching, tokens to mel), then HiFT vocoder (mel to 24 kHz audio), then Streaming playback (plays as chunks arrive, resumes from checkpoints)

  • iOS: 4-bit quantized, compiled to CoreML on first launch.
  • macOS: 8-bit quantized, native MLX.
End-to-end pipeline from book analysis to streaming playback. Thanks to Resemble AI for open-sourcing the Chatterbox weights.

Key decisions

  1. Hybrid book analysis

    regex-based quote detection feeds into on-device LLM (Foundation Models) for speaker attribution and emotion tagging

  2. Each voice is a set of speaker embeddings, not a separate model

    switching voices is instant, costs no additional memory

  3. 12 emotion profiles map to concrete TTS parameters (temperature, guidance weight, pacing)

    the same text sounds different when a character whispers vs. shouts

  4. Streaming synthesis with checkpoint resume

    if the app is interrupted mid-generation, it picks up exactly where it left off

  5. Voice cloning from a few seconds of reference audio via an LSTM encoder

    bring your own narrator