Voice dock — orb, modes, and Jarvis

The voice orb sits at the bottom of the left sidebar. Hold it to talk to Jarvis, click it in Note-Taking, or hold it (or Fn) in Transcribe to paste into the last widget you clicked. Speech-to-text is local Whisper; spoken replies are local Piper.

Normal · Transcribe · Note-Taking Wake word Jarvis Local Whisper + Piper

The orb at the bottom of the sidebar

The dock is pinned in the sidebar under Agents, Widgets, and Skills — not in the title bar. Idle, you see the orb and the mode label under it. While you record or Piper plays, an equalizer animates beside the button. After a take, a short You / Assistant transcript appears next to the orb and fades (hover to keep it open).

Jarvis is both the spoken assistant (wake word or hold-to-talk) and the canvas chat widget. The same Settings → Agent provider powers both. Configure X.ai, DeepSeek, Kimi, Z.ai, Anthropic, OpenAI, Gemini, Ollama, or a custom OpenAI-compatible endpoint there — then talk or type.

Microphone and speaker devices, Piper voice/speed/volume, the Whisper model, the launch greeting, and the wake word all live in Settings → Voice. Endpoints and start/stop live on Settings → Local Servers.

Three modes

Switch from the menu bar Voice → Mode, or click the label under the orb. You cannot change mode mid-recording — finish or cancel the current take first. Menu labels and dock labels differ slightly; the hints below are what the app prints.

Menu Voice → ModeDock labelHintHow you record
Normal Normal Mode Hold to speak · optional wake word Hold the orb. Optional wake word stays armed.
Transcription Transcribe Hold to record · paste into last clicked widget Hold the orb, or hold Fn for 1.5s. Wake word is paused.
Note-Taking Note-Taking Click to record · summarize · upload (wake word paused) Click the orb to start, click again to stop.

The dock menu restates each path: Normal shows whether the wake word is on; Transcribe adds hold Fn 1.5s; Note-Taking reminds you the wake listener is paused.

Activity pills

A small label beside the orb tracks the current step. These are the exact strings:

PillWhen
ListeningMicrophone is open (hold, click-to-record, or post-wake command).
TranscribingLocal whisper-cli (or the Whisper fallback) is turning audio into text.
ThinkingThe Settings → Agent provider is writing a reply or note summary.
Tool callingJarvis is running a canvas or MCP tool (widgets, search, Extensions, HyperFrames…).
Generating voicePiper is synthesizing the spoken line.
SpeakingPiper audio is playing on the chosen output device.
TelegramAn incoming Telegram message is driving the orb.
SummarizingNote-Taking only — provider is turning the transcript into bullets + title.
Saving noteNote-Taking only — upload to your Fuse Intelligence profile.
PastingTranscribe only — text is going into the last clicked widget.

Idle hides the pill. Thinking has no equalizer; Listening and Speaking do.

Jarvis wake word

Hands-free listening lives in Settings → Voice under the Wake word section. It is armed in Normal Mode only.

  • Listen for a wake word — toggle, default on. Off = hold-to-talk only.
  • Wake word — text field, default Jarvis (one or two spoken words, max 32 characters).
  • Sensitivity — Low / Medium / High. Default Medium at 0.4. Lower if background speech wakes it; raise if it misses your name.
  1. Stay in Normal Mode with Listen for a wake word on.
  2. Say the name. VibeFuse plays the bundled “Yes, sir?” clip (no Piper round-trip) and starts recording.
  3. Speak the command. The take ends on silence, or after 12 seconds with no speech.
  4. The same pipeline as hold-to-talk runs: transcribe → provider (and tools) → Piper.
BehaviorDetail
Works while Piper is speakingA new wake interrupts playback and starts a fresh command — including while Jarvis is using tools.
Note-Taking / TranscribeThe listener pauses. Switch back to Normal Mode to arm it again.
Dismiss phrasesAfter the ack, say never mind, go away, just kidding, sleep, go to sleep, or no — no provider call, no spoken reply.
PrivacyListening and the “Yes, sir?” clip are on-device. The provider is called only after a real command.
Wake word not firing? Confirm Normal Mode, grant Windows Privacy → Microphone, and keep Listen for a wake word on. Very short or common custom names false-trigger — stick with Jarvis unless you need another name.

Note-Taking

Click-to-record dictation that transcribes locally, summarizes with your Agent provider, and uploads to your profile. The wake word stays paused for the whole mode.

  1. Set a provider in Settings → Agent (X.ai, DeepSeek, Kimi, Z.ai, Anthropic, and the other cards). The summary step uses it.
  2. Switch the dock to Note-Taking.
  3. Click the orb to start (do not hold). Click again to stop. Very short takes are ignored.
  4. Pills walk TranscribingSummarizingSaving note.
  5. Review the note at Account → Notes and in Settings → Notes.

What is saved

  • Default title: Voice note (replaced by the provider title when summarization works).
  • Body: bullet summary, then a divider, then the full transcript:
- first takeaway - second takeaway --- Transcript the raw words you spoke

If summarization fails, the raw transcript still uploads under a fallback title — nothing is dropped.

Normal Mode and Transcribe never write to your profile. Note-Taking is the only upload path, and it needs a valid VF- key.

Transcribe

Dictation into the canvas — not a cloud note. Hold the orb, or hold Fn for 1.5 seconds, speak, then release. The pill shows Pasting while text is injected into the last clicked widget (Terminal, Jarvis, a CLI pane, and other paste-aware widgets).

Click the destination widget first. If nothing was focused, there is nowhere to paste.

Settings → Voice

Assistants group, subtitle Microphone, speakers, and wake word. Five sections, in this order:

SectionControls
Audio devices Output device (Piper, wake-word audio, video, browser webviews), Input device (hold-to-talk, notes, wake word), Refresh devices.
Spoken output Piper Voice (Jarvis, Lewis, Michael, Adam, Heart, Bella, Emma), Speed 0.50–2.00, Volume 0–150%. Full engine notes: Piper TTS.
Launch greeting Greet me on launch — bundled “Good morning” (5:00–11:59), “Good afternoon” (12:00–16:59), or “Good evening”. Uses the output device and volume above. If autoplay is blocked, it plays on the first click or key press.
Whisper model Local ggml model for the next listen (tiny is fastest). Model picker lives here — not on Local Servers. See Whisper STT.
Wake word Listen for a wake word, Wake word field, Sensitivity. See above.

A help line under Whisper points at Settings → Local Servers for the Piper and Whisper URLs.

Status bar

The footer chips that matter for voice:

ChipShowsClick opens
ModeNormal, Transcription, or Note-TakingSettings → Voice
VoiceSelected Piper voice (Jarvis, Lewis, …)Settings → Local Servers
SpeechWhisper Ready / DownSettings → Local Servers
PiperPiper Ready / DownSettings → Local Servers

Provider and Model chips sit beside them and open Settings → Agent. Discord opens Settings → Extensions.

Troubleshooting

SymptomFix
Orb does nothingAllow the microphone in Windows Privacy. Use Refresh devices, then pick Input / Output.
Wake word never firesNormal Mode + Listen for a wake word on. Say the name clearly. Raise Sensitivity if it misses; lower it if TV speech wakes it.
Wake word while Piper talksSupported — interrupt and give a new command. If it never interrupts, confirm you are still in Normal Mode.
No spoken replyTurn on Use local Piper in Settings → Local Servers and click Test Piper. Voice/speed/volume stay on Settings → Voice. See Piper TTS.
Transcription empty or slowNeed ffmpeg on PATH and a Whisper model on disk. Pick a smaller model in Settings → Voice. See Whisper STT.
Notes not uploadingNote-Taking needs a Settings → Agent provider, a valid VF- key, and network to fuseintelligence.org. Check Account → Notes and Settings → Notes.
Transcribe does not pasteClick a widget first. Hold the orb or hold Fn 1.5s, then release.
Cannot switch modesFinish or cancel the current take. The picker is locked while Listening / Transcribing / Thinking.