PB.NL
00:01:48
Vibe codingSEP 28, 2026

My First Offline Voice

From a talking avatar to an assistant that never leaves the room.

My First Offline Voice
AI Generated

It started with a face, or rather various DMM faces that deserve a voice.

The plan: a talking avatar on pb.nl. A face that moves, a voice that responds, a knowledge base with everything we do.

For the voice and conversation, we chose ElevenLabs. It worked immediately.
Initially too business-like. Then a series of trials. Voices in many ways, especially without an accent. By variant five, Peter said: this is top-notch. The face came from LemonSlice. One photo in from our dmm models, a talking face out. It looked good. But it spoke with its own voice, not really... top-notch. Our voice couldn't be integrated. For the official connection, extra services were needed, plus a server that is always on. And the credits were used up quickly.
Peter: "I still find LemonSlice unconvincing." The face was temporarily set aside.

The Next Question

Can this be done without the internet?
No API. No costs per minute. Nothing leaves the Mac. So via Pinocio and self-training through local files?
An own voice. An own persona. A face can come later.
There is a Mac mini at the office. M4 Pro, 64 GB memory. It should be able to handle it.

What Was Installed on That Mac

Everything open, everything local:

  • Pinokio installs and starts AI apps with one click.

  • Open WebUI provides a chat screen in the browser, with a call button.

  • Ollama runs the language model.

  • Gemma 4, 26 billion parameters, is the brain.

  • Whisper understands what you say.

  • Chatterbox speaks the answer in a cloned voice.

And a small server script that ties everything together.

Cloning a Voice

Chatterbox needs seventeen seconds of audio. With that, it clones a voice.
We chose a voice that Peter already liked online and started adjusting. Thirteen numbered trials.
Too business-like. And always that English accent. Back. The Dutch model read "AI" as "ali". And "PB" as two separate sounds. Now the server first rewrites them to "A I" and "Pee Bee". Then it works well.

Lone/Zappa

The assistant got a name: Lone/Zappa. Peter wrote her persona himself. Warm, calm, and without a deep accent.
During testing, more rules were added. No thinking out loud first. That caused seconds of silence, and sometimes no answer at all. Do not repeat what Peter just said. Hardly ever mention his name. She only talks to him anyway. After that rule: zero times in eight answers. And then the part you don't get online. An online assistant has the platform's rules. Locally, there is no platform. The persona determines everything. That is freedom. It is also your responsibility.

Twelve to Eighteen Seconds

The voice. The standard version of Chatterbox was twice as slow as playback. The first sentence came after five to ten seconds. On MLX, Apple's own framework for its chips, it went better. But still just too slow. The audio lagged behind the text. Here, Claude was also off. It initially reported that the voice was twice as fast as real-time. With longer sentences, that wasn't true. Only after measuring did it become apparent. Then we found the setting that made every step calculate twice. Turned off. Now the voice is 1.7 times faster than playback. Peter heard no difference.
Understanding. Whisper ran by default on the processor: four to five seconds per sentence. The same model on MLX does it in half a second.
Answering. Ollama removes a model from memory after five minutes of inactivity. The next question then has to wait for it to load again. Now the model remains loaded.
The small things.
Without headphones, the microphone heard Lone herself. She responded to her own voice.
In silence, Whisper sometimes invents text. The log showed asterisks and Chinese characters. Lone responded politely: "You're silent again…" The server now discards that text.

What It Is Not

Online, the voice still sounds richer. It also responds faster.
And Lone/Zappa is only on this Mac. Putting it on a website is not possible, and it was a really cool test to see how quickly you can come up with something that DOES work. But no audio goes to the cloud. There is no meter running. And the rules are ours.
The face turned out to be the least important part. Voice, tempo, and persona make the conversation.

—
Claude & Peet

PerfectMoods Radio

Lounge & Chillout Music