OPEN SOURCE VOICE PROTOCOL

Give voice to your agents, and agency for your voice.

Orchestrate autonomous agents, terminal runtimes, and desktop workflows hands-free. Powered by on-device Whisper with real-time spoken feedback, active listening and zero cloud dependencies.

$ curl -fsSL https://vifi.sh | sh

Choose Your Character

Viv

VoiceFi
🇺🇸 US English Expressive & Modern

Christopher

🇺🇸 US English Authoritative & Calm

Emily

🇮🇪 Irish English Gentle & Melodic

Aria

🇺🇸 US English (Emma) Second Voice • Obsidian Agent

Sonia

Code Reviewer
🇬🇧 UK English Warm & Analytical

Stefan

🇺🇸 US English Fast & Energetic
Feedback Loop

ProActive Listening

Step through the zero-click autonomous cycle connecting agent runtimes directly to your voice.

VoiceFi Mobile Headset Companion
9:30
5G
VoiceFi

VoiceFi Mobile

STANDBY
Ready to Listen
TRANSCRIPT STREAM 00:00
“Agent finished unit tests. Say 'Deploy to staging' to dispatch...”
HUD macOS
VoiceFi Ready (⇧⌘N)
Standing by • Dictate (⌃T) or speak to agent (⌃R)
Terminal

VoiceFi Core v0.8.4 (darwin-arm64 • Apple Metal)

Local bridge listening on /tmp/voicefi.sock (127.0.0.1:5141)

Ready for agent events. Press 'Start Live Loop' to test...

Local is Faster

From spoken thought to IDE code execution. Every bar is coloured by what kind of work the time is spent on — on-device compute, network transit, remote inference, or human friction — across local agent workflows, local push-to-talk tools, cloud speech APIs, and manual dictation.

On-device compute Network transit — audio leaves your machine Remote inference Human / UI friction
VoiceFi Local Agent Pipeline Metal CoreML + Direct IPC
120ms ⚡ Instant

Audio is transcribed in RAM via Apple Silicon Metal acceleration and dispatched directly to IDE subagents via local IPC. Every millisecond is on-device — zero cloud round-trips, zero queue delays.

Superwhisper / MacWhisper Local Whisper + Clipboard Paste
800ms 6.7× Slower

Also fully on-device and zero-egress — the honest local comparison. The cost is push-to-talk: focus the caret, hold the key, wait, then paste. No agent dispatch.

Cloud Speech Pipelines HTTP / Remote GPU
1,450ms 12× Slower

TLS handshake + WAN audio upload + remote ASR GPU queue + cloud LLM formatting + webhook roundtrip back to your machine. Most of the bar is red and amber: your audio, off your machine.

Manual Dictation & Context Switching Baseline, not a product
2,800ms+ 23× Slower

Switching window focus, holding push-to-talk, waiting for text, copy/pasting into the IDE prompt. Measures human friction rather than any single product.

Agentic Differentiators

One-Way Dictation
vs.
The VoiceFi (vifi) Feedback Lp

Traditional dictation tools only replace typing. VoiceFi connects the entire circular feedback loop with your agents, tools (via MCP) and anything that has terminal access.

← Swipe to compare all platforms
Capability VoiceFi (vifi) Wispr Flow Superwhisper / MacWhisper ElevenLabs Cloud
Two-Way Feedback Loop Agent speaks turn soundbite → Auto-listens → Developer replies ✓ Continuous 2-Way Loop ✗ 1-Way Dictation only ✗ 1-Way Dictation only ✗ Disconnected from IDE
AI Agent Hook Integration Native stop & notification hooks for Antigravity, Claude Code, Cursor, Aider ✓ Native 1-Line Setup ✗ No Agent Hooks ✗ No Agent Hooks ✗ No IDE Agent Hooks
Universal Voice MCP Server Model Context Protocol support — allows any AI agent or tool to talk, listen & synthesize ✓ Native Voice MCP (Any tool can talk) ⚠️ Read-Only Notetaker MCP (No speech tools) ✗ No MCP Support ✗ Proprietary API Only (No MCP)
"Second Voice" & Obsidian Vault Plugin Dual-voice layer for your Second Brain — ties multi-agent vaults to Obsidian via native plugin & MCP ✓ Multi-Voice & Native Obsidian Plugin ⚠️ Meeting transcripts only (No voice plugin/TTS) ✗ No Obsidian Plugin or Multi-Voice ✗ Disconnected from Vaults
Hands-Free Auto-Listening (VAD) Mic arms automatically when agent finishes a turn (zero clicking) ✓ Automatic Turn-Based VAD ✗ Push-to-Talk only ✗ Push-to-Talk only ⚠️ Continuous Cloud Stream
Smart Turn Summarization Cleanses raw terminal logs into 1-sentence spoken briefings ✓ Intelligent TTS Briefing ✗ No TTS / Readout ✗ No TTS / Readout ✗ Generic Voice Output
Audio Privacy & Egress Local Metal/CoreML Whisper in RAM vs cloud audio streaming 🔒 100% On-Device (0 Egress) ⚠️ Cloud Audio Upload 🔒 0 Bytes Egress (Offline) ⚠️ Cloud Audio Streaming
Cost & Licensing Open source CLI freedom vs monthly recurring software subscriptions 💚 100% Free MIT CLI $12 – $20 / month $20 – $30 one-time / sub $0.05 – $0.30 / min API
Looking for deeper technical & latency specs? Explore full stage-by-stage latency breakdowns, Unix socket IPC benchmarks, and runtime matrices.
Full Architecture & Benchmark Matrix
Universal Compatibility & Protocols

Agent & IDE Ecosystem

Zero-latency voice and acoustic layer connecting autonomous coding agents, second-brain vaults, and system terminals.

🎛️ SOUND LAB
DRAG OR TAP CONNECTOR LINES TO PLUCK ACOUSTIC HARP
VoiceFi Core VoiceFi PORT 5141 CORE SERVER
1. Autonomous Agents

Antigravity

Subagent Hooks

AUTO

Claude

Terminal REPL

STDIO

Gemini

Live & Flash SDK

VOICE

Cursor

Voice Dispatch

HOTKEY
2. Knowledge Vaults & Editors

Obsidian

Second Brain

PLUGIN

VS Code

Extension

IDE
3. Protocols & Shell

MCP

Model Context Protocol

STDIO

WebMCP

navigator.modelContext

W3C

Ollama

Local LLMs

OFFLINE

Terminal

Ghostty • Warp • iTerm • tmux

PIPES
Google Antigravity AUTO-HOOKED

Transport: Unix domain socket (/tmp/voicefi.sock) • Hook: agent.turn_completed

vifi setup --dev
14-DAY FREE TRIAL · NO CREDIT CARD REQUIRED · INSTANT ACTIVATION

Developer-First Pricing.
Zero Cloud Lock-In.

100% free and open-source for local Apple Silicon synthesis and Whisper STT. Upgrade to Pro for ultra-low latency edge cloud relays, curated neural voices, and multi-agent routing.

Monthly
Annual 1-Time Special SAVE 36% · ~$5.75/MO
COMMUNITY Open Source
$0 / forever

100% private, local-first voice layer running entirely on your Mac’s Apple Silicon NPU.

100% Local Apple Silicon TTS (0ms)
Local Faster-Whisper STT (RAM)
Native MCP Stdio Server (JSON-RPC)
Antigravity & Claude Code Hooks
Global Hotkey (Ctrl+T)
UNIX Domain Socket IPC (/tmp/voicefi.sock)
MIT License & Air-Gapped Safe
14-DAY FREE TRIAL INCLUDED · MOST POPULAR
DEVELOPER PRO 14-Day Trial Included
$69 / 1-year pass

Save 36% · Equivalent to $5.75/mo billed once

Curated neural voices, ultra-fast cloud relay, streaming STT, and mobile pacing companion.

Everything in Community Tier
20+ Curated Neural Personas (Viv, Christopher, Aria, Guy, Sonia)
Ultra-Fast Cloud Relay (Groq / ElevenLabs Turbo fallback)
Streaming STT & Real-time Token Stream
Mobile & Web Companion (Live waveform & remote pacing)
Cross-Agent Audio Turn Routing (Antigravity ↔ Claude)
Silent Background Auto-Updater
Priority Discord & GitHub Support
Start 14-Day Free Trial

Instant download · No credit card required upfront · Reverts to Free tier if cancelled

ENTERPRISE Custom Cloud
$49 / seat / mo

Dedicated voice infrastructure, private voice cloning, and enterprise compliance.

All Developer Pro features
Hosted Private Voice Cloning
Team Subagent Audio Swarm Routing
SSO, SAML & SOC2 Audit Logs
Custom Service Level Agreements (SLA)
Dedicated Solutions Architect

Frequently Asked Questions

How does the 14-day free trial work?

The 14-day free trial starts automatically the moment you install VoiceFi on your Mac. All 20+ neural voices, ultra-fast cloud relay, streaming STT, and mobile pacing companion tools are completely unlocked with zero credit card required upfront.

What happens after the 14-day trial period ends?

VoiceFi never bricks or locks your setup. If you do not activate a Pro license, VoiceFi simply gracefully falls back to the Community tier ($0 Forever), retaining 100% functionality for local Apple Silicon neural synthesis, local Whisper STT, and the MCP server.

Why is VoiceFi priced at $9/mo vs competitors charging $15–$20/mo?

Unlike cloud-only voice tools that send all your audio through heavy remote GPU servers, VoiceFi runs local-first on Apple Silicon hardware (0ms latency, $0 server cost). We only use cloud relay for edge neural fallbacks, allowing us to deliver 95%+ gross margins and pass those savings directly to developers.

How does the 1-Time Annual Special ($69/year) work?

The Annual Special is a single, one-time payment of $69 that grants you 1 full year of VoiceFi Pro access (equivalent to $5.75/month — a 36% discount vs monthly). You receive a perpetual license key that you can activate via vifi license activate <KEY> or the macOS menu bar.