Reliability fixes — resolved "Transcription failed" errors and doubled audio caused by a stray second microphone capture after rapid dictation starts, and a rare timing issue that could garble the last word of a dictation.
Changelog
Major product updates across Windows and macOS.
This page tracks the bigger FluidVox milestones pulled from the git history so we have a durable release record we can keep extending over time.
Better AI-cleanup diagnostics — support logs can now separate connection setup, provider response waiting, and download delays without recording credentials, addresses, or dictated content.
July 14, 2026
Fix: last words could occasionally garble at the end of a dictation (v2.12.4)
- Fixed a rare timing issue where the final fraction of a second of audio could be duplicated or corrupted as a recording stopped, occasionally garbling the last word of a cloud-transcribed dictation. The same fix also prevents rare truncation at the tail of batch recordings.
July 13, 2026
Fix: "Transcription failed" errors and doubled audio after rapid dictation starts (v2.12.3)
- Fixed a bug where pressing the dictation hotkey twice in quick succession while the microphone was still waking up could leave a second, invisible audio capture running. From then on most dictations failed with "Transcription failed: object reference not set…" and the rest transcribed poorly because every recording contained doubled audio. Bluetooth headsets were especially affected because they wake slowly. The stray capture also kept the microphone active until the app was restarted; that no longer happens.
- Fixed dictations being lost when text cleanup returned nothing (for example, a Polished-style dictation the cleanup model "polished away"). FluidVox now falls back to your raw words instead of showing an error, matching the macOS behavior.
June 20, 2026
Fix: dictation could silently fail to insert text in some fields (v2.12.2)
- Fixed a bug where dictation could finish and save to your history but never type into the target field, leaving nothing inserted. It happened when the text cursor sat far down in a long text box, or in apps that report their text state in an unusual way. Dictation now inserts reliably in these cases, and the pre-insert check can no longer abort a paste if it fails.
June 14, 2026
Faster, more reliable text cleanup and a stronger Vox Agent (v2.12.1)
- Dictation cleanup — used for non-natural styles and non-English transcripts — now runs on Google's GA Gemini 3.1 Flash-Lite instead of a preview model. This moves cleanup off a deprioritized path that occasionally caused 15–30 second pauses, so dictation finishes more consistently and with lower latency overall.
- The Vox Agent now runs on Gemini 3.5 Flash for stronger multi-step reasoning. Heavy generation (documents, slides, sheets, flashcards) is unchanged.
June 13, 2026
VM-safe hotkey detection, combo-release fix, and clearer Settings copy (v2.12.0)
- New "VM-safe hotkey detection" setting (on by default) adds a third hotkey detector that uses Windows Raw Input, so the dictation trigger fires inside VMware Workstation, VirtualBox, Remote Desktop, and Citrix — where the normal low-level hook gets swallowed by the virtual-machine keyboard grab. For full dictation into a VM, turn this on together with "Type out everywhere" — that combo handles both the trigger and the text delivery.
- Fixed a bug where combo hotkeys (e.g. Ctrl+Q) could start walkie-talkie recording but never stop it on key-up. The Raw Input release path now also stops the recording when Windows' built-in RegisterHotKey started it, and the combo is now re-press safe — pressing it again while recording is treated as a stop.
- Alt-modified combos (e.g. Alt+Space) now work in the VM-safe path: the Raw Input handler also matches WM_SYSKEYDOWN / WM_SYSKEYUP, the message types Windows uses for Alt-modified key events.
- Settings copy pass: the "Type out everywhere" and "VM-safe hotkey detection" cards now describe what each toggle does in plain language, and an informational warning (not a hide) tells you when your active recording hotkey is Alt-based and might misbehave inside VMs.
- New developer-facing repro harness (
tools/typeout-repro/TypeOutRepro8.cs+run-repro8.ps1) covers Raw Input registration, modifier+key matching, observe-only semantics, and the combo-release path — 12 scenarios, all green.
June 10, 2026
Type Out Everywhere: dictate into VMs, Remote Desktop, and locked-down apps (v2.11.0)
- New "Type Out Everywhere (VMs & Remote Desktop)" setting delivers dictation as real keystrokes wherever your cursor is — so text lands in places a simulated paste can't reach, like VMware/VirtualBox guests, Remote Desktop and Citrix sessions, or locked-down fields. Flip it in Settings, or quickly from the tray icon menu.
- The type-out engine was rebuilt to send true scan-code keystrokes — the same events a physical keyboard produces — fixing garbled type-out text (words collapsing into "....." or going missing) that could occur in apps with spellcheck, such as Notepad and Chromium-based editors.
- Multi-line dictations keep their line breaks in type-out mode: newlines are typed as Shift+Enter, so chat apps insert a new line instead of sending your message mid-dictation.
- Caps Lock no longer inverts typed-out text — it's temporarily parked while typing and restored afterwards, and typing stops automatically if you switch to a different app mid-dictation.
- The "Type Out in Browsers & Code Editors" setting now greys out while Type Out Everywhere is on, since the everywhere mode already covers those apps.
June 7, 2026
Local licenses go fully offline, with offline device transfer (v2.9.0–2.10.0)
- Local licenses now work completely offline using a securely signed license file — including launching the app and dictating with no internet connection at all, even before sign-in.
- Offline device transfer lets you move a Local license to another PC without contacting the server: export a transfer receipt on the old machine and import it on the new one.
- License file export/import, plus automatic license refresh whenever you do sign in online, keep the offline license current.
- A round of licensing reliability fixes: atomic license-file writes, smoother renewals, and an upgrade-safe database migration.
June 2, 2026
Instant paste by default, plus reliability fixes (v2.7.5)
- Dictated text once again pastes instantly in browsers and code editors by default. The previous version typed text out character-by-character in Chromium- and Electron-based apps (Chrome, VS Code, Cursor), which was noticeably slower; instant clipboard paste is the default everywhere again.
- Added a "Type out in browsers & code editors" setting for the occasional app that ignores a simulated paste (such as Cursor). Leave it off for instant paste; turn it on only where text won't otherwise land.
- Fixed a rare issue where pressing the dictation key together with another key could leave FluidVox stuck in a "listening" state and unresponsive until you restarted it.
- The on-device Parakeet model now loads once at startup instead of twice, roughly halving its memory use and speeding up launch.
June 1, 2026
Spacing fix and reliable pasting in Cursor and Electron apps (v2.7.4)
- Fixed an unwanted space that could appear before dictated text in apps like Microsoft Teams, OneNote, and Cursor. FluidVox now decides spacing from the live cursor position rather than your previous dictation, so a space is only added when you're genuinely continuing a sentence after punctuation.
- Transcriptions now paste reliably into Chromium- and Electron-based editors such as Cursor and VS Code. Previously a simulated paste could be silently dropped in those apps; text is now typed directly into the focused editor.
May 19, 2026
Customizable recording sounds (v2.7.3)
- Recording start and stop cue sounds are now fully customizable in Settings — pick from 16 curated Windows system sounds for each cue, or set either one to "None (silent)" to mute it independently.
- A master "Recording Sounds" toggle lets you turn audio feedback off entirely.
- Each picker has a preview button so you can hear the sound before committing to it.
- New installs now default to the friendlier Speech On / Speech Off cues, better suited to a voice app than the previous hardware-insert/remove chime. Existing users keep their current sounds until they change the picker.
April 2, 2026
Natural style zero-network mode, hands-free hotkey customization, and injection fix (v2.6.1–2.6.3)
- Natural transcription style now runs entirely on-device with zero network calls — no cloud cleanup, no quota usage, and no BYO key required.
- Hands-free recording hotkey is now fully customizable in Settings, with the same modifier+key picker used for the dictation and command hotkeys.
- Three-way hotkey conflict detection prevents assigning the same key combo to dictation, command, and hands-free hotkeys.
- Text injection spacing logic was simplified — a space is now only prepended when actual punctuation is detected immediately before the cursor, reducing false positives.
- The target window is now focused before reading the cursor position, fixing intermittent spacing failures in apps like Notepad.
April 1, 2026
WASAPI audio fixes, multi-language auto-detect, and default style improvements (v2.6.0)
- WASAPI microphone device now stays alive for 5 minutes between recordings, eliminating broken warm restarts that caused silent captures.
- The RMS silence threshold was lowered from 0.02 to 0.01, fixing false speech rejection from quiet wireless microphones.
- Multi-language auto-detect mode lets you select target languages for Whisper, using native probability scoring to pick the best candidate — especially helpful for similar languages like Urdu and Hindi.
- Default per-app style categories now seed with Natural (follow global setting) instead of specific overrides, giving new users a cleaner starting experience.
- Fixed a false "Performance restored" power notification that appeared when already on AC power.
March 27, 2026
Custom styles, system audio muting, and audio capture upgrade (v2.4.11–2.4.13)
- Custom Styles let you create your own transcription style with a free-text prompt — translate to another language, rewrite as a haiku, add annotations, or anything else you can describe. Up to 5 custom styles per user, assignable per-app.
- Custom styles now work correctly with personal Gemini API keys — translation and annotation prompts are no longer overridden by built-in language constraints.
- A new "Mute System Audio on Record" toggle silences other apps while you dictate, so background noise doesn't interfere with transcription.
- Microphone capture migrated from WaveIn to WASAPI, giving instant mic release after recording ends — other apps regain mic access immediately.
- English-only Whisper models (Tiny, Base, Small) no longer show Auto Detect or non-English languages in the language picker, preventing misconfiguration.
- Win key combo hotkey now stops cleanly on key release, fixing edge cases where the hotkey could get stuck.
- Device identification migrated to a more reliable hardware fingerprint.
March 26, 2026
Whisper language detection fix and text injection reliability (v2.4.10)
- Fixed Whisper auto-detect language not working correctly — the processor was configured with an empty language string instead of "auto", causing unreliable language detection for multilingual users.
- Text injection now releases any stuck modifier keys (Alt, Shift, Win) before pasting, preventing the destination app's menu bar from activating instead of receiving the paste.
- Alt key release uses dummy scancode padding to avoid triggering Windows menu bar activation when clearing stuck state.
March 26, 2026
Power state resilience and performance scaling (v2.4.9)
- The app now automatically recovers when switching between AC power and battery — no restart required.
- Audio devices, hotkeys, and efficiency mode settings are revalidated after every power transition.
- Local transcription models (Parakeet and Whisper) are automatically reinitialized when the power source changes.
- Whisper CUDA sessions gracefully fall back to CPU if the GPU becomes unavailable on battery, and restore GPU when AC power returns.
- Parakeet thread count now scales dynamically based on your CPU cores and power state for optimal performance.
- The system tray icon automatically recovers if Windows Explorer restarts.
- Fixed an issue where holding the Alt hotkey for long recordings could trigger the Windows menu bar.
March 25, 2026
Streaming local transcription and microphone improvements (v2.4.3–2.4.5)
- Local transcription now uses real-time streaming instead of file-based recording, significantly reducing latency.
- Recording starts instantly with no delay between pressing the hotkey and audio capture beginning.
- The selected microphone is now correctly used across all recording modes, including Help Mode voice input.
- Automatic gain control with fade-in improves audio quality, especially for quiet microphones.
- Empty audio with no speech detected is now skipped instead of producing a blank transcription.
- Alt-key combo hotkeys no longer trigger Word's KeyTips ribbon after recording stops.
March 24, 2026
CUDA auto-detection and transcription reliability (v2.4.0–2.4.2)
- The app now auto-detects your GPU and CUDA availability, with a built-in GPU setup guide for optimal local transcription performance.
- Cold-start model warmup eliminates the first-transcription delay after launching the app.
- Fixed garbled transcription caused by an incorrect sample rate mismatch detection.
- Fixed a crash that could occur when recording very short silence-only audio with Whisper.
- Improved spacing between consecutive dictations.
March 22, 2026
Local file transcription (v2.3.0)
- Imported audio and video files can now be transcribed using on-device Whisper or Parakeet models — no cloud access required.
- A new Local/Cloud toggle in the Audio Transcriptions tab lets you choose between on-device and cloud transcription for files.
- Notes now save correctly for Local tier users.
March 21, 2026
Local Monthly subscription plan (v2.2.0)
- A new Local Monthly plan is available for users who want local-only transcription on a monthly billing cycle.
- Device management support was added for multi-device licensing.
March 12, 2026
Help Mode: real-time screen-sharing guidance (v2.1.0)
- Help Mode lets you share your screen and get spoken step-by-step guidance in real time, powered by Gemini Live API.
- Local transcription quality improved with proper audio resampling and automatic gain control.
- Fixed an issue where the update dialog could appear in a loop on launch.
- Fixed a race condition where Alt-key hotkeys could leave the recording stuck.
March 11, 2026
One-click support diagnostics (v2.0.1)
- A new "Copy Support Info" button in Settings copies a full diagnostic snapshot (account, hardware, settings) to the clipboard for easy pasting into support messages.
- Internal analytics and telemetry improvements for better reliability.
March 10, 2026
Parakeet local transcription — fast, accurate, no GPU required (v2.0.0)
- Parakeet CPU-based transcription models are now available as a second local engine alongside Whisper, powered by sherpa-onnx.
- Two Parakeet models ship: a fast English-only model and a multilingual model supporting 25 languages — both run efficiently on CPU with no GPU needed.
- The onboarding hardware check now recommends Parakeet for machines without a capable GPU, so every user gets a good local transcription experience out of the box.
- The Manage Local Models dialog shows both Parakeet and Whisper models with hardware-aware "Recommended" badges and descriptions explaining when each engine works best.
- Active Parakeet models are preloaded in the background at app startup to eliminate cold-start latency on the first transcription.
- Model extraction was rebuilt to handle large archives reliably, fixing a rare issue where downloaded model files could be silently truncated.
March 10, 2026
Local Whisper language detection and Gemini response filtering (v1.3.3)
- Local Whisper transcription now detects the spoken language in auto-detect mode and passes it to Gemini for correct cleanup.
- The "Auto" transcription engine option was removed — users now choose Cloud or Local directly, simplifying settings.
- Gemini meta-responses (e.g. "Please provide the audio file...") are now filtered out instead of being injected as text.
March 9, 2026
Rich plan dialog, media pause, Microsoft OAuth, and hotkey improvements (v1.3.0)
- A new Choose Your Plan dialog matches the macOS design with rich plan cards and feature breakdowns.
- Media now automatically pauses when you start recording and resumes when you stop.
- Microsoft OAuth was added for Calendar and Tasks integration alongside Google.
- Configurable Win key hotkeys and additional hotkey combinations were added in Settings.
- Text injection was improved with streaming and progressive insertion for faster results.
March 8, 2026
Smoother Google Sign-In and smarter model recommendations (v1.2.9)
- Google Sign-In no longer triggers the "unverified app" warning — only basic profile scopes are requested at login.
- Calendar and Tasks permissions are now requested incrementally, only when the user first uses those features.
- The Whisper model management dialog now shows a GPU-aware "Recommended" badge next to the best model for your hardware.
March 7, 2026
Windows agent file handling got a major usability pass
- The file-management tool now supports fuzzy file search, better directory filtering, and optional search parameters.
- PowerShell output is being surfaced more clearly, including deferred display for larger results.
- The Vox result window can now render open-file and open-folder actions directly from agent results.
- Agent prompt rules were tightened so tool output is repeated back to the user instead of being treated as visible by default.
- Math and markdown rendering was updated to support the new file-action links inside results.
March 6, 2026
Windows moved into the 1.2.x release line
- Windows 1.2.4 through 1.2.6 shipped with Gemini key UX fixes and a Windows version bump.
- Large local Whisper model crashes were addressed with VRAM detection and recovery logic.
- Bring-your-own Gemini key routing was wired into the Windows product flow.
- RDP clipboard handling was fixed for remote desktop users.
March 5, 2026
Update prompts, verification flow, and toolbar behavior improved
- Windows 1.2.0 added an update dialog on launch so new builds are easier to discover.
- Email verification for password sign-up shipped, then moved onto the Resend-based verification flow.
- The floating toolbar became draggable.
- Speculative miss handling was optimized to keep the interaction loop tighter.
March 2-4, 2026
Vox Agent, auto-updates, and a broader Windows UI overhaul landed
- A major Windows feature update shipped with hotkey fixes, Vox Agent support, and UI overhaul work.
- Velopack auto-update support was integrated into the Windows app.
- Tray menu behavior and per-app style handling were improved.
- Cloud function prompt issues were fixed as the Windows agent stack matured.
February 23, 2026
Recording flow and hotkey behavior got much tighter
- Instant recording landed, with better start and stop sound timing.
- Alt-key recording issues were fixed, including menu activation on key release.
- Short silent recordings now skip unnecessary fallbacks instead of creating noisy failures.
- System output is muted during recording to reduce interference.
February 23, 2026
Update delivery and release infrastructure were added
- A visible Check for Updates button was added to the Windows app.
- The auto-update channel was corrected to the Windows release stream.
- Release pipeline work shipped with code signing and packaging improvements.
- License binding and auto-start behavior were hardened for real installs.
February 21-23, 2026
Vox actions expanded beyond plain dictation
- Hey Vox gained screenshot capture support for visual tasks.
- Notepad fallback handling improved text insertion reliability.
- Audio transcription UI work shipped so longer-form audio flows are visible in the app.
February 20-22, 2026
Security, onboarding, and subscription checks were upgraded
- Security hardening landed across the Windows app and release flow.
- Local Whisper support was added as part of the on-device transcription path.
- Email verification gating shipped for new accounts.
- Trial dates were added to new user profiles so account state is clearer.
February 10-11, 2026
The Windows app launched and quickly closed the first UX gaps
- The initial WinUI 3 Windows app foundation shipped.
- Dashboard UI parity work brought the Windows experience closer to macOS.
- Tray icon context menu, text injection, microphone selection, and latency issues were fixed early.
July 17, 2026
Detailed network diagnostics for AI cleanup (v2.11.7)
- Diagnostic exports now separate DNS, connection, TLS, request upload, wait-for-first-byte, and response download timings for managed cleanup, direct Gemini, and OpenRouter. They also record connection reuse, HTTP protocol, proxy state, and whether the route is constrained or expensive.
- FluidVox records a privacy-safe network-path snapshot at cleanup start and completion, making it possible to detect connectivity changes during a long dictation. Credentials, URLs, IP addresses, Wi-Fi names, and dictated text are never included.
July 17, 2026
Faithful numbers and more reliable AI polishing (v2.11.6)
- Polished dictation now preserves every spoken number, date, time, percentage, quantity, and amount exactly as transcribed. Cleanup no longer performs unspoken arithmetic or replaces a repeated value to make the message seem internally consistent.
- Gemini cleanup now has up to 15 seconds to return a valid result before FluidVox safely uses the original transcript. This recovers slow but successful responses that previously arrived just beyond the 12-second cutoff.
July 17, 2026
Fix: text could silently fail to appear in Chrome and other browsers (v2.11.5)
- Browsers now always receive dictated text through guarded typing into the verified editor. Chrome could acknowledge a direct Accessibility insertion without ever applying it to the page — FluidVox believed delivery succeeded while nothing appeared — so that insertion path has been retired for browsers.
- The exact recording-time field and window are re-verified immediately before the first character is typed, so a page script that moves focus while your transcript is being prepared can no longer pull the text into a different field.
July 16, 2026
Fix: Chrome’s editor could go undetected, blocking delivery (v2.11.4)
- Chrome sometimes puts its accessibility layer to sleep, leaving the focused editor invisible to FluidVox at recording start — dictations then ended with “Couldn’t paste transcription.” FluidVox now wakes Chrome’s accessibility engine as recording begins so the exact editor is found the first time.
- A final editor re-check runs right before delivery, recovering the field even when it was hidden at recording start.
July 16, 2026
Restored dependable pasting across Mac apps (v2.11.3)
- Returned to a simpler delivery flow: FluidVox restores and verifies the recording-time app and window once, then sends the transcript without repeatedly asking macOS Accessibility to approve every character.
- Chrome and Gmail now get a brief, spaced field check as recording begins when the focused editor is temporarily hidden. FluidVox keeps that exact editor and never sends text to a bare browser window.
- Codex and Claude no longer lose the remainder of a transcript because of one brief Accessibility timeout after typing has begun.
- The recording-time interaction guard now stays in force through restoration and delivery, so clicking or typing elsewhere cannot redirect a transcript into the wrong field.
- When direct delivery is needed, text is now sent in compact guarded chunks instead of a flood of single-character events, improving spaces, punctuation, and long-form reliability without borrowing the clipboard.
July 16, 2026
Reliable Chrome, Gmail, Codex, and Claude pasting (v2.11.2)
- Restored native paste delivery for Chrome, Gmail, and longer dictations after FluidVox revalidates the exact recording-time text field, avoiding the dropped spaces and punctuation that busy browser editors can produce with character-by-character typing.
- Made the final target check and paste one guarded operation. FluidVox now restores your previous clipboard only after the original field proves the transcript landed; if delivery cannot be verified, the transcript remains safely available instead of a delayed paste receiving stale clipboard content.
- Codex and Claude can now be brought back to the exact captured window before delivery. Held-key repeats and inertial trackpad scrolling also no longer cancel a valid hands-free dictation.
- Bring-your-own-key cleanup now calls and waits for only the Gemini or OpenRouter model you selected. FluidVox no longer races a hidden faster model and uses whichever reply arrives first.
July 14, 2026
More reliable dictation in Codex and other Mac editors (v2.10.9)
- Fixed dictations that completed and appeared in history but sometimes did not show up in Codex. When Codex does not expose its focused editor through macOS Accessibility, FluidVox now uses a tightly guarded Codex-only typing fallback.
- The fallback continuously verifies that the same Codex app and window are still active, and stops immediately if you type, click, or switch apps—so text cannot drift into the wrong place.
- Improved text-target capture in Electron and web-based editors by safely retrying temporary macOS Accessibility failures, while keeping the normal fast path unchanged.
July 14, 2026
Faster dictation, reliable wake-up, and safer text insertion (v2.10.8)
- Dictation cleanup on macOS now uses Gemini 3.5 Flash with a strict 2.5-second limit. If a safe cleaned version is not ready in time, FluidVox immediately uses your original transcript instead of making you wait.
- Fixed cases where dictation could feel slow or fail to appear after the app had been idle, the Mac had slept, or the user session had been unlocked. Recording now starts promptly, and microphone, model, and hotkey state recover cleanly.
- FluidVox now remembers the exact app, window, and text field where a dictation began. Finished text can no longer drift into a newly focused field, arrive through a delayed clipboard paste, or be inserted twice. If the original field is no longer available, the transcript stays safely in history and FluidVox tells you what happened.
- Rapid consecutive dictations and cancelled Vox actions are now isolated from one another, so an older result cannot overwrite or insert after a newer recording. Selection-based transforms also replace text only when the original selection is still unchanged.
June 24, 2026
More reliable, faster dictation cleanup (v2.10.4)
- Dictation cleanup — used for non-natural styles and non-English transcripts — now sticks to cleaning your words and no longer occasionally “answers” a request you dictated instead of tidying it up.
- Cleanup runs on a faster, generally-available model path, so your text appears with lower latency than before.
June 24, 2026
Cleanup never rewrites what you said when using your own Gemini key (v2.10.3)
- Fixed a bug where, with on-device transcription, your own Gemini key, and a non-natural style, the cleanup step could “answer” a long dictated request instead of cleaning it — replacing your spoken words with a reply.
- Upgraded the cleanup model and added a safety check that falls back to your raw words whenever cleanup diverges from what you actually said.
- Your original transcript is now preserved in history with a one-tap copy original button, so nothing you dictated is ever lost.
- Removed a cap that previously truncated very long dictations.
June 14, 2026
Faster cleanup, support log export, and diagnostics (v2.10.0–2.10.2)
- Text cleanup now runs on a faster model path, eliminating the occasional 15–30 second stalls some dictations hit before your text appeared.
- Added a private On-device cleanup option (macOS 26+) that polishes your dictation fully offline using Apple’s on-device models — no Gemini key and no cloud required. It stays strictly private and falls back to your raw words rather than ever sending text to the cloud.
- Added a Diagnostics section in Settings with an Export Logs for Support button, so you can send us a timeline of an issue without any setup. Logs never contain your dictated text.
June 8, 2026
Mouse-button hotkey fix and a Watch & Learn cleanup (v2.9.8–2.9.9)
- Fixed mouse-button triggers that could get stuck “on” — pressing a modifier key while a mouse button was your recording trigger could latch a recording indefinitely. Mouse-button hotkeys now ignore stray keyboard events.
- Switching your hotkey now safely resets any in-progress recording state, so a key held across a settings change can’t leave dictation stuck.
- Added an opt-in Keep model loaded setting that keeps the on-device model warm, removing the reload delay after idle periods.
- Added Outlook.com and Outlook 365 web compose options, and tidied up the Watch & Learn video section.
May 25, 2026
Independent left/right modifier sides for recording and commands (v2.9.7)
- The command (Vox Agent) hotkey can now use its own modifier side, separate from your recording hotkey.
- You can now set combinations like Left Command to record and Right Command for commands, with conflict detection that understands left vs. right.
May 5, 2026
Rock-solid long dictations on the on-device models (v2.9.3–2.9.5)
- Fixed crashes that could occur during long, multi-minute dictations when using the on-device transcription models, by dramatically reducing memory pressure during recording.
- Auto-learn correction monitoring is now skipped for very large pasted dictations, where it could stall under memory pressure, keeping long sessions smooth.
- The Dashboard window now remembers its size and position between launches.
- Long-form audio file and audio-note transcriptions now stay raw and unstyled end-to-end for cleaner, more faithful results.
April 20, 2026
Longer custom style instructions (v2.9.1–2.9.2)
- Custom style instructions can now be much longer (up to 3,000 characters), giving you room for more detailed formatting guidance.
- Improved Help Mode reliability on older versions of macOS.
April 1, 2026
Multilingual auto-detect and per-app styles that follow your default (v2.8.0–2.9.0)
- Added an Auto Detect mode for multilingual on-device models — pick the set of languages you speak and the app figures out which one you’re using on the fly, with a new toggle to switch between auto-detect and a single fixed language.
- Improved language detection accuracy, including better handling of Hindi and Urdu.
- Per-app styles now follow your global default until you deliberately change them, so setting a new default applies everywhere you haven’t customized.
- Fixed Help Mode audio playback and improved screen-recording permission handling, with a clear restart prompt when permission needs an app restart.
March 28, 2026
Watch & Learn tutorials and a smarter Help Mode (v2.6.1–2.7.0)
- Added a new Watch & Learn tab to the dashboard with featured and curated tutorial videos to help you get more out of FluidVox.
- Upgraded Help Mode to a newer, more capable model with a much larger context window for better answers.
- Fixed false agent triggers: normal dictation containing words like “walks,” “docs,” or “talks” no longer accidentally launches the Vox Agent. The agent is now only invoked by its explicit command hotkey.
March 27, 2026
Create your own custom dictation styles (v2.6.0)
- You can now create and save your own custom styles with your own instructions, beyond the built-in ones — available throughout the app in the Style view, per-app Style Matching, the apps list, and the floating toolbar’s quick-switch chips.
- Your custom styles appear at the top of the style pickers for quick access.
- The language list now reflects only the languages your active model actually supports.
- Improved spacing when text is inserted right after existing text, for cleaner consecutive dictation.
March 22, 2026
Local file transcription for all users (v2.5.0)
- Imported audio and video files can now be transcribed using on-device Whisper or Parakeet models — no cloud access or Pro subscription required.
- A new Local/Cloud toggle in the Audio Transcriptions tab lets you choose between on-device and cloud transcription.
- Free users with a downloaded local model now have full access to the file transcription feature.
- The audio language picker automatically filters to languages supported by the active local engine.
- Model loading happens on demand — if the selected model isn't in memory, it loads automatically before transcription begins.
March 9, 2026
Mouse button hotkeys, web integrations, and agent tool upgrades
- Any mouse button (M4/M5 side buttons, middle click, etc.) can now be assigned as the recording trigger — hold for walkie-talkie, tap for hands-free.
- Mouse button events are intercepted and suppressed so they don't fire their default action (e.g. Back/Forward) in other apps.
- Web integration services were added for Gmail, Google Calendar, and Microsoft services via OAuth.
- The Vox Agent gained enhanced calendar, reminders, email, and AppleScript tools for richer voice-driven actions.
- A device switch confirmation sheet was added for smoother multi-device handling.
March 3-4, 2026
Core stability work focused on freezes and agent recovery
- Main-thread freezes caused by mouse monitoring and audio-level storms were fixed.
- Engine health checks now include a startup grace period to avoid false positives.
- The Vox Agent loop can now detect and recover from malformed Gemini tool calls.
- Audio engine race conditions in the correction monitor were addressed.
February 28-March 1, 2026
Plans, API key flexibility, and usage messaging were refined
- Pro, Pro+, Local, and trial users got clearer plan management paths with a View Plans button.
- Usage messaging switched from exact token estimates to simpler relative usage labels.
- Pro, Pro+, and trial users can optionally bring their own Gemini API key.
- Pricing and token budgets were simplified around the lower Pro price point.
February 28, 2026
Vox Agent became much more capable and less noisy
- New agent tools shipped alongside better PDF and presentation creation flows.
- Math rendering and cloud backend support expanded the output quality of generated documents.
- Brief action confirmations moved to toast notifications instead of heavier interruptions.
- The agent stopped opening unnecessary browser tabs for simple knowledge questions.
- Screenshot sending in iMessage was fixed and content safety rules were added.
February 26-27, 2026
Document, spreadsheet, and style workflows expanded fast
- Live spreadsheet editing shipped for Numbers and Excel, including chart support.
- Presentation and PDF editing tools were added, followed by real file creation workflows.
- Per-app writing style was applied directly to Vox Agent composition.
- Style switcher chips were added to the floating toolbar and synced with the dashboard.
- Email reply workflows and AppleScript feedback loops were fixed.
February 24-25, 2026
Screenshots, image generation, and subscription packaging matured
- Screenshot capture moved to ScreenCaptureKit for more reliable visual context.
- Full-page capture and read-screen fallback tooling were added for the agent.
- Generated images can now render inline in the Vox Result window with a copy action.
- Image generation switched to Gemini's native image API.
- Pro+ and unified token-budget subscription routing were introduced.
February 23, 2026
Vox Agent rolled out in phases and became a headline feature
- The initial Vox Agent release shipped with multi-step tools and confirmation handling.
- Conversation memory, prompt improvements, and new tool phases followed the same day.
- Document-generation tools, command hotkeys, file reading, and language support expanded the feature set.
- Screenshot behavior, errors, hotkeys, and sound handling were tightened immediately after launch.
February 8-19, 2026
The macOS app foundation filled in updates, local models, and dictation polish
- Sparkle auto-updates, release pipeline work, and download packaging shipped.
- Bring-your-own Gemini key support and local transcription models were added.
- Automatic language detection, style matching improvements, and correction learning kept improving output quality.
- Custom hotkeys, hide-from-dock behavior, Hey Vox voice commands, and audio transcription all landed in this stretch.