Type
Personal experiment
Year
2026
Built with
Pluggable TTS · Whisper · Image gen · Python/Flask
Read time
7 minutes

Spine. Turning a written story into voice, image, and conversation through AI.

Spine reading a sci-fi novel: script-style transcript with per-character speaker labels, an audio timeline, and AI-generated scene images

About this project

This is a personal experiment built for my own reading. This case study focuses on the design idea, the multi-voice approach, and the build, not on specific book content.

Challenge

Single-voice audiobooks flatten the experience. Every character speaks in the same narrator's voice, illustrations live in a different app, and looking up an unfamiliar phrase means leaving the reader entirely. The challenge was building something closer to how I want to experience a book: multiple voices, scene-level imagery, instant explainers, all in one place, and cheap enough to run on any book I owned.

Strategy

Treat the book as a structured script, not a wall of text. Each chapter is markdown with **[SPEAKER]** tags; each speaker maps to a TTS voice plus acting direction. The voice provider is pluggable: OpenAI's hosted catalog for the widest range of voices, or a local model like Chatterbox for cost-free, fully offline rendering. Run Whisper across the rendered audio for word-level karaoke sync. Use an LLM pass to detect scenes and generate illustrations on demand. Add a selection-based "explain" layer that opens an LLM chat about any phrase, and persist what was explained or imagined so it can be revisited with a click. Content-address every artifact so re-runs cost nothing.

How a chapter becomes an audiobook

  1. 01

    Parse

    Read the markdown chapter and split it into speaker-tagged lines.

    Cached
  2. 02

    Cues

    Map each speaker to a voice and its acting direction from voices.json.

    Cached
  3. 03

    Render

    Send each cue to the TTS provider, hosted or local, and stitch the audio.

    Cached
  4. 04

    Word sync

    Run Whisper over the audio to get a timestamp for every word.

    Cached
  5. 05

    Scenes

    An LLM pass finds scene boundaries and writes one image prompt per scene.

    Cached
  6. 06

    Images

    Illustrations render when the reader reaches them, with a shared character sheet.

    On demand

Every step writes to a content-addressed cache, so changing one voice re-renders one character's lines and nothing else.

Results

A working local audiobook reader with multi-voice TTS, word-level seek-on-click, per-scene illustrations, persistent highlights, and an LLM explain layer. A full novel (~120k words, ~80 scenes) renders end to end, with every step cached so iteration is free. The pipeline is generic enough to work on any book: swap the markdown script, the voice mappings, and the character sheets, and it produces a new audiobook.

The Spine reader mid-playback: a chapter timeline at the top, speaker labels in the left margin, the word being spoken highlighted in yellow, a passage underlined in blue from an earlier image request, and an AI illustration titled Jack in the Magnetic Forest in the right column
The reader, mid-playback. The word being spoken is highlighted, speaker labels sit in the margin, and the illustration for the current scene stays in the right column. The blue underline is a passage a previous session turned into an image.

The story

Spine started from a small frustration: I wanted to listen to a book and have characters sound like characters, see what a scene actually looked like, and pull up an explanation of a strange phrase without losing my place. The pieces existed separately (TTS, illustration models, LLM chat) but nothing wove them together into a reader.

The most interesting design move was treating voice as a per-character prompt, not just a voice selection. Each speaker in voices.json gets a voice plus acting direction: "weary, introspective, slightly raspy, slow deliberate pacing"; "practiced friendliness with subtly wrong rhythm, human words but alien prosody." Same TTS model, but the same voice can sound like a different character in a different mood depending on the instruction. The whole pipeline becomes a director's chair, not a button.

The audition screen for the Narrator: a character list on the left, a grid of twelve voices with short descriptions, a long acting-direction prompt, a sample paragraph, and a list of rendered takes with play controls
Casting a character. Each speaker gets a voice from the catalog and an acting direction written like a note to an actor. Takes are rendered against a sample passage and kept, so a voice can be tuned before the full chapter is committed.

The second move was making illustrations feel like part of the prose rather than decoration. The LLM detects scene boundaries from the script, generates a single image prompt per scene, and the reader's sticky right column tracks the viewport so each scene's image surfaces when the corresponding text plays. Images are spoiler-locked behind audio playback. A consistent character sheet runs through every prompt so the same character is recognizable across scenes.

The opening of chapter two, Theseus, with the epigraph, speaker labels, and a generated illustration titled Resurrection in the Coffin Row showing crew members waking in a ship corridor
A scene with its image. The illustration for the opening scene of a chapter, generated from the LLM's scene prompt and the shared character sheet.
A scene titled Autopsy of a Scrambler whose image has not been generated yet: the right column shows a striped placeholder with a Generate image button and the note click to choose provider
Before it exists. Scenes hold a placeholder until the reader asks for the image, so nothing renders, or costs anything, ahead of the reading.

The third move was layering interactive reading on top. Select any prose to either get an LLM explainer (in-context for that passage, with follow-up chat) or to generate a custom image. Both leave a colored underline behind, so the book accumulates a quiet trail of what's been explored. Click a marked passage later and the cached explanation or image comes back instantly.

A phrase selected in the prose with a small floating menu offering two actions, Explain and Generate image, beside a generated illustration titled Sascha Breaks the Conversation
Select anything. Highlighting a phrase offers two actions: explain it, or turn it into an image.
The Explain panel open over the reader: the selected phrase at the top, toggles for in this passage and general, a short explanation that ties the phrase to the scene, and a follow-up question box
The explain layer. The answer is written for this passage, not the phrase in general, and a follow-up box keeps the conversation in place. The phrase keeps an underline so the answer can be reopened later.

Spine is less about the specific output and more about what the build itself proves. Take a long-form medium that's been mostly static, treat it as a structured material, and you can produce something genuinely multimodal with a small set of generative AI APIs and a careful caching strategy. The hard part isn't the AI calls. It's deciding what reading should feel like.

Select highlights

  1. Defined and built Spine end to end as a personal experiment in multimodal reading.
  2. Designed per-character voice direction as prompts, not just voice picks, so one TTS catalog produces distinct characters.
  3. Built a multi-pass pipeline: parse → cues → render → word-sync → scenes → on-demand image gen, all content-addressed cached.
  4. Designed a reader with karaoke-style word sync, sticky scene illustrations, and selection-based Explain and Generate-image actions.
  5. Built a per-voice audition UI so character voices and acting direction could be tuned iteratively before committing to a full render.
  6. Architected the TTS layer as a pluggable provider so the same script renders through OpenAI's hosted voices or a local model like Chatterbox, depending on cost, privacy, and voice-catalog needs.