Voice dock — orb, modes, and Jarvis
The voice orb sits at the bottom of the left sidebar. Hold it to talk to Jarvis, click it in Note-Taking, or hold it (or Fn) in Transcribe to paste into the last widget you clicked. Speech-to-text is local Whisper; spoken replies are local Piper.
The orb at the bottom of the sidebar
The dock is pinned in the sidebar under Agents, Widgets, and Skills — not in the title bar. Idle, you see the orb and the mode label under it. While you record or Piper plays, an equalizer animates beside the button. After a take, a short You / Assistant transcript appears next to the orb and fades (hover to keep it open).
Jarvis is both the spoken assistant (wake word or hold-to-talk) and the canvas chat widget. The same Settings → Agent provider powers both. Configure X.ai, DeepSeek, Kimi, Z.ai, Anthropic, OpenAI, Gemini, Ollama, or a custom OpenAI-compatible endpoint there — then talk or type.
Microphone and speaker devices, Piper voice/speed/volume, the Whisper model, the launch greeting, and the wake word all live in Settings → Voice. Endpoints and start/stop live on Settings → Local Servers.
Three modes
Switch from the menu bar Voice → Mode, or click the label under the orb. You cannot change mode mid-recording — finish or cancel the current take first. Menu labels and dock labels differ slightly; the hints below are what the app prints.
| Menu Voice → Mode | Dock label | Hint | How you record |
|---|---|---|---|
| Normal | Normal Mode | Hold to speak · optional wake word | Hold the orb. Optional wake word stays armed. |
| Transcription | Transcribe | Hold to record · paste into last clicked widget | Hold the orb, or hold Fn for 1.5s. Wake word is paused. |
| Note-Taking | Note-Taking | Click to record · summarize · upload (wake word paused) | Click the orb to start, click again to stop. |
The dock menu restates each path: Normal shows whether the wake word is on; Transcribe adds hold Fn 1.5s; Note-Taking reminds you the wake listener is paused.
Activity pills
A small label beside the orb tracks the current step. These are the exact strings:
| Pill | When |
|---|---|
| Listening | Microphone is open (hold, click-to-record, or post-wake command). |
| Transcribing | Local whisper-cli (or the Whisper fallback) is turning audio into text. |
| Thinking | The Settings → Agent provider is writing a reply or note summary. |
| Tool calling | Jarvis is running a canvas or MCP tool (widgets, search, Extensions, HyperFrames…). |
| Generating voice | Piper is synthesizing the spoken line. |
| Speaking | Piper audio is playing on the chosen output device. |
| Telegram | An incoming Telegram message is driving the orb. |
| Summarizing | Note-Taking only — provider is turning the transcript into bullets + title. |
| Saving note | Note-Taking only — upload to your Fuse Intelligence profile. |
| Pasting | Transcribe only — text is going into the last clicked widget. |
Idle hides the pill. Thinking has no equalizer; Listening and Speaking do.
Jarvis wake word
Hands-free listening lives in Settings → Voice under the Wake word section. It is armed in Normal Mode only.
- Listen for a wake word — toggle, default on. Off = hold-to-talk only.
- Wake word — text field, default Jarvis (one or two spoken words, max 32 characters).
- Sensitivity — Low / Medium / High. Default Medium at
0.4. Lower if background speech wakes it; raise if it misses your name.
- Stay in Normal Mode with Listen for a wake word on.
- Say the name. VibeFuse plays the bundled “Yes, sir?” clip (no Piper round-trip) and starts recording.
- Speak the command. The take ends on silence, or after 12 seconds with no speech.
- The same pipeline as hold-to-talk runs: transcribe → provider (and tools) → Piper.
| Behavior | Detail |
|---|---|
| Works while Piper is speaking | A new wake interrupts playback and starts a fresh command — including while Jarvis is using tools. |
| Note-Taking / Transcribe | The listener pauses. Switch back to Normal Mode to arm it again. |
| Dismiss phrases | After the ack, say never mind, go away, just kidding, sleep, go to sleep, or no — no provider call, no spoken reply. |
| Privacy | Listening and the “Yes, sir?” clip are on-device. The provider is called only after a real command. |
Note-Taking
Click-to-record dictation that transcribes locally, summarizes with your Agent provider, and uploads to your profile. The wake word stays paused for the whole mode.
- Set a provider in Settings → Agent (X.ai, DeepSeek, Kimi, Z.ai, Anthropic, and the other cards). The summary step uses it.
- Switch the dock to Note-Taking.
- Click the orb to start (do not hold). Click again to stop. Very short takes are ignored.
- Pills walk Transcribing → Summarizing → Saving note.
- Review the note at Account → Notes and in Settings → Notes.
What is saved
- Default title: Voice note (replaced by the provider title when summarization works).
- Body: bullet summary, then a divider, then the full transcript:
If summarization fails, the raw transcript still uploads under a fallback title — nothing is dropped.
Normal Mode and Transcribe never write to your profile. Note-Taking is the only upload path, and it needs a valid VF- key.
Transcribe
Dictation into the canvas — not a cloud note. Hold the orb, or hold Fn for 1.5 seconds, speak, then release. The pill shows Pasting while text is injected into the last clicked widget (Terminal, Jarvis, a CLI pane, and other paste-aware widgets).
Settings → Voice
Assistants group, subtitle Microphone, speakers, and wake word. Five sections, in this order:
| Section | Controls |
|---|---|
| Audio devices | Output device (Piper, wake-word audio, video, browser webviews), Input device (hold-to-talk, notes, wake word), Refresh devices. |
| Spoken output | Piper Voice (Jarvis, Lewis, Michael, Adam, Heart, Bella, Emma), Speed 0.50–2.00, Volume 0–150%. Full engine notes: Piper TTS. |
| Launch greeting | Greet me on launch — bundled “Good morning” (5:00–11:59), “Good afternoon” (12:00–16:59), or “Good evening”. Uses the output device and volume above. If autoplay is blocked, it plays on the first click or key press. |
| Whisper model | Local ggml model for the next listen (tiny is fastest). Model picker lives here — not on Local Servers. See Whisper STT. |
| Wake word | Listen for a wake word, Wake word field, Sensitivity. See above. |
A help line under Whisper points at Settings → Local Servers for the Piper and Whisper URLs.
Status bar
The footer chips that matter for voice:
| Chip | Shows | Click opens |
|---|---|---|
| Mode | Normal, Transcription, or Note-Taking | Settings → Voice |
| Voice | Selected Piper voice (Jarvis, Lewis, …) | Settings → Local Servers |
| Speech | Whisper Ready / Down | Settings → Local Servers |
| Piper | Piper Ready / Down | Settings → Local Servers |
Provider and Model chips sit beside them and open Settings → Agent. Discord opens Settings → Extensions.
Troubleshooting
| Symptom | Fix |
|---|---|
| Orb does nothing | Allow the microphone in Windows Privacy. Use Refresh devices, then pick Input / Output. |
| Wake word never fires | Normal Mode + Listen for a wake word on. Say the name clearly. Raise Sensitivity if it misses; lower it if TV speech wakes it. |
| Wake word while Piper talks | Supported — interrupt and give a new command. If it never interrupts, confirm you are still in Normal Mode. |
| No spoken reply | Turn on Use local Piper in Settings → Local Servers and click Test Piper. Voice/speed/volume stay on Settings → Voice. See Piper TTS. |
| Transcription empty or slow | Need ffmpeg on PATH and a Whisper model on disk. Pick a smaller model in Settings → Voice. See Whisper STT. |
| Notes not uploading | Note-Taking needs a Settings → Agent provider, a valid VF- key, and network to fuseintelligence.org. Check Account → Notes and Settings → Notes. |
| Transcribe does not paste | Click a widget first. Hold the orb or hold Fn 1.5s, then release. |
| Cannot switch modes | Finish or cancel the current take. The picker is locked while Listening / Transcribing / Thinking. |