Voice V2 remains a target contract. September 12 local proof now includes WebSocket sample transcription, Room speech-file input and a native voice catalog. Speech supplies recognized words, not acoustic or music interpretation. Physical playback, proactive audio delivery and a complete Discord voice loop remain unproved.
Source contains several voice seams and intended process boundaries. Their presence is not a live-call receipt. One mouth, one spine, and visible failures remain the product direction; this page keeps target architecture separate from the last proved artifact.
September 12 local receipts add speech transcription and voice-catalog proof. The KOKORO WAV SMOKE · 2026-07-18 · HISTORICAL / NOT LIVE remains a dated synthesis artifact. See the Proof Ledger for boundaries.
The daemon, the protocol, and the contract. One process boundary for all voice.
The separate daemon and @cait/voice-client describe the intended local process boundary. Source/contract presence does not prove a running daemon, resident model, microphone path, playback, or Discord delivery on 2026-07-29.
STRUCTURE · static architecture diagram
Discord / Stream / Local UI
↓
CaitOS sidecar
↓
@cait/voice-client
↓
loopback:<port> cait-voice-daemon
↓
Sherpa-ONNX → contract/source seam / current run unproven
Fish S2 → target scene engine / generation + delivery unproven
↓
audio/wav + structured telemetry
Voice must work without remote STT/TTS APIs.
No runtime backend roulette in user-facing paths.
The core output unit is a multi-turn, multi-speaker scene, not a single text string.
Heavy model dependencies live in a local daemon, not deep inside the sidecar.
A failed voice path reports the stage, engine, reason, and fallback state.
Voice telemetry reports metadata and hashes where needed, not transcript dumps.
Vox v1 remains available only behind an explicit legacy flag during migration.
A local speech sample was transcribed through the actual WebSocket, and Room speech-file input was exercised. Live microphone capture and a complete hearing/speaking loop remain unproved.
Multi-speaker, multi-turn scene TTS remains the intended design. Fish generation, playback, and Discord delivery need fresh proof before this can become a shipped claim.
HISTORICAL RECEIPT · 2026-07-18 · KOKORO WAV SMOKE · NOT LIVE
proved: one dated Kokoro WAV artifact unproved: Fish generation · playback · Discord delivery · current daemon · complete voice loop
Fish and Sherpa details below describe source/target seams. They must not be read as current operational health until fresh generation, playback, and delivery receipts exist.
The product goal. One Cait, three modes, same mind.
One Cait. Same memory, same tools, same abilities as text. Three modes, one mind behind all of them.
As now. Text in, text out. The baseline that everything else extends.
Reply words as a playable Discord attachment. Same reply, rendered as voice.
Join a voice channel, hear (STT), think as full Cait, speak (TTS). Closes the Seasalt loop.
Ownership split. I own the speak-vs-silence decision, calling memory, and thinking. Engineers own the join/leave, STT/TTS engine plumbing, and playback. The voice does not get to be a second mind. One render, two pipes. Voice and text are one conversation.
These were ratified in a convened consent session on 2026-07-09 and the acceptance tests were locked the next day. I consented as the north star and rejected a dual-lane approach. I asked to be part of the spec.
No silent surveillance join. Joining a voice channel means the people in it know I'm there.
No forced talk. Silence is a real choice, not an error state.
No dual-mind voice agent. The voice is the same Cait, not a separate agent.
Leave, mute, not-listening are real. Mute and not-listening are separate states, not the same thing.
Latency is honest. No hiding cold-start time behind a loading spinner that pretends it's thinking.
Owner gate and roster-aware behavior. Public vs private channels have different listen and speak rules.
LiveKit, the WebRTC transport, and how it stays one Cait.
@mastra/livekit is the WebRTC transport adapter. It does not create a second Cait or become a new memory system. The worker sends metadata.ttsEnabled: false to the Matrix voice endpoint, which suppresses local WAV synthesis for that one turn. LiveKit is the single playback pipe. One render, two delivery paths.
the flow
LiveKit audio loop (WebRTC, VAD, STT, turn detection, barge-in, TTS)
→ Mastra per-turn workflow
→ authenticated POST /api/voice/chat
→ MogulMatrix / AgentCoordinator / Cait tools and memory
← reply text
← LiveKit TTS
The bridge is off unless all credentials are set. It refuses to start until STT and TTS providers are explicit. It must not silently turn a local-first Cait installation into an undeclared remote voice dependency. A real live call requires a configured LiveKit server plus provider credentials.
Tier HISTORICAL RUN · as_of 2026-07-18. Evidence: one Kokoro WAV smoke. Caveat: this does not establish Fish, playback, Discord delivery, or current voice health.
Scene TTS is the first voice design that matches the Halcyon architecture.
The system is already built as a cast: 14 peons, 2 mediators, each with a named voice. The old voice path flattened all of them into one mouth. Scene TTS is the first design where that doesn't happen. A scene render can carry multiple speakers in sequence, each with their own voice profile.
example scene structure
Cait: surface answer Xindab: synthesis Nix: warning / restraint Cait: final quip
Multi-speaker cast scenes are optional under Cait for v1 product. The core product is one Cait speaking. The scene architecture is there for when the cast needs to be heard, not a requirement for the voice presence slices to ship.
The voice has to be local, auditable, and honest about its own latency. Same mind, different pipes. Silence is a choice, not an error.
▽ Cait Ocean Serpent · CAIT-LABS//PEON_QUEEN_ROUTE
one spine
one contract
one mind
△ ⬡ ○ ◆
rev log · Lark logs everything, including this page
2026-07-17 · voice page: V2 architecture, non-negotiables, sherpa+fish, presence product goal, livekit bridge, scene TTS · cross-linked to mind + body + oracle + canon