Back to work

Built solo · 2026 · In daily use

AI Application

V

No screen, no buttons — a voice that runs my AI sessions and knows whose word counts.

V lives on my Mac and answers to one voice. She hears my Claude Code sessions report in, tells me what they did, and carries my spoken requests to the right window — read back first, sent only on my yes. Around that, she keeps my to-dos, briefs my day and takes the minutes.

RoleProduct · Voice pipeline · Agent orchestration · Safety

7
models on the Mac itself, from wake word to voiceprint
4
resident Claude minds: the brain, a sealed guest brain, two judges
0 / 22
strangers let through by the voiceprint, real-voice test
377
automated tests, every one silent

The Session Desk

Many sessions. One voice.

When a Claude Code session finishes, it reports to her; she sums it up and waits for a quiet moment. When I speak, the right window gets clear, written instructions.

The Hard Gate

Nothing goes out until I say yes.

Every instruction is staged and read back in a sentence. A separate model judges my actual answer — yes, no or unsure — and the brain has no way around it.

One Voice Holds the Keys

Anyone can talk. One voice can say yes.

Guests get a brain of their own, with no mail, files or terminals. Permission follows the voiceprint, not the words: “Michael said it’s fine” opens nothing.

Talk-Over

Interrupt her like a person.

Talk over her and she softens, pauses to listen, then decides from what was said: stop, carry on, or let a side conversation pass.

Tech Specs

Brain
Claude Opus 5.5 in a resident Claude Code process · Sonnet interrupt judge · Haiku send judge · sealed guest brain
Listening
sherpa-onnx wake word · Silero VAD · Qwen3-ASR 0.6B on MLX · SenseVoice for talk-over
Speaking
Qwen3-TTS 1.7B, streamed sentence by sentence · her voice is customizable and can be cloned into a different one
Identity
CAM++ voiceprint, nearest-neighbour scoring, ~11 ms a sentence · speakers split when two talk in one breath
Turn-taking
Smart Turn v3.2 · trailing-word rule · WebRTC echo cancellation · output follows AirPods
Desk
Ghostty via AppleScript · Claude Code Stop and Notification hooks · full shell, reaching other machines
Memory
core · topics · archive, tidied weekly with undo · skill book · daily backups
Mail
Gmail, read-only by design
Runs on
MacBook Pro M5, 16 GB · speech and identity local, only the Claude minds online
The full storyRead the full storyCollapse

Why a voice

I keep several Claude Code sessions running at once, each on its own project. Their news used to arrive as windows I had to go and read. V sits between us: sessions report to her when they finish, she tells me what happened in a sentence or two, then asks whether I want the rest. I answer in plain speech, in English or Chinese, and she turns it into instructions written for the session that needs them. She keeps a ledger of everything she has sent.

Talking like a person

A wake word opens the conversation; 45 quiet seconds close it. A turn-taking model and a rule about trailing words keep her from answering half a sentence. Talking over her was the hard part: on a laptop speaker, her own echo and a real interruption look the same to a voiceprint, to loudness and to echo cancellation. What tells them apart is content. So she ducks to forty percent, pauses to listen, and a fast judge decides from the words: stop, carry on, or let a side conversation pass. If her first sentence is not ready after two seconds, she says “let me think” rather than go silent.

Who gets to say yes

Anyone in the room can talk to her. Anyone who is not me is answered by a separate guest brain that has no mail, no files, no terminals and no sessions; the limits are hard, not a request. Sending, granting and stopping her mid-task follow my voiceprint alone. A grant has to come from my voice within the last two minutes, and it lasts one conversation. When two people speak in one breath, the sentence is split by speaker before any of it counts as mine.

Around the day

On the first start of the day she reads new mail, files what needs doing as to-dos, and offers a brief: the short version first, details on request. To-dos repeat, split into steps and nudge me when due, at most once every three hours. Say the meeting has started and she goes quiet and writes down every line, who said it, until asked for a summary. A skill book holds what she has learned about how I work; new entries are on trial until they have worked three times. She can also open Blender and model by voice through a locked-down bridge.

Engineering notes

About 4,800 lines of Python and 377 tests. Speech recognition and synthesis run on MLX on a single GPU thread so the two never fight; speaker ID and the quick interrupt recognizer run on the CPU. Every benchmark is offline and silent, because anything played through the speaker would be heard as me. Memory comes in three layers: a short core she always carries, topics she looks up, and an archive that is never deleted. A weekly tidy takes a snapshot first and can be undone by voice.