Remix.run Logo
thangalin 7 hours ago

Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:

https://www.youtube.com/watch?v=WAeHgE94rVo

No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.

[1]: https://deepmind.google/models/gemma/gemma-4/

[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design

RGS1811 2 hours ago | parent | next [-]

I've been working on a similar project all year and as a tip, you should try Fish Audio or Higgs as a replacement for Qwen3. Both yield much better prosody and are much easier to listen to for long runs.

thangalin an hour ago | parent [-]

> Fish Audio or Higgs

I wasn't able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?

MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.

RGS1811 an hour ago | parent [-]

For the voice design, these don’t support it, but for the final render, they’re much better. So your pipeline could for example generate voices with one tool and render with another.

loremm 7 hours ago | parent | prev | next [-]

It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.

I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion

eru 4 hours ago | parent | prev | next [-]

Awesome! I had been meaning to build something like this for a while now, but never got around to it.

Is it possible to annotate your text with extra 'stage directions' that influence how the book is read out?

thangalin 4 hours ago | parent [-]

> annotate your text with extra 'stage directions'

Good idea, not something I've considered yet. Wouldn't take much to add it since there's already a feature for selecting a quotation and assigning it an intonation. Same infrastructure could be reused to select arbitrary text and assign stage directions.

Multicomp 7 hours ago | parent | prev | next [-]

<grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>

The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.

Jordan-117 6 hours ago | parent [-]

Their first sentence literally tells you it's a video of their app? It's not a mystery-meat link.

talon8635 7 hours ago | parent | prev [-]

[flagged]

idiotsecant 4 hours ago | parent [-]

I guess I should feel bad about wanting to listen to audiobooks of novels that don't have audio while I drive