Speech-to-text and AI text processing for macOS. Transcribe audio using on-device AI models or cloud APIs (Groq, OpenAI, xAI/Grok), then transform the result with reusable workflows. Your voice data stays on your Mac with local models - or use cloud APIs for faster processing.
TypeWhisper 1.7.0 is the current stable release for macOS. It adds iCloud sync
for History and Inbox, a visual keyboard shortcut editor, Undo Last Dictation
and Restore Raw Transcript, new indicator styles, expanded dictation recovery,
Finder transcription, and compatible local model imports.
Development on main targets 1.8; daily builds use the v1.8.0-daily.* line.
See the 1.7.0 release notes,
release readiness guide,
support matrix, and
release validation process for the shipped feature
set and ongoing 1.x maintenance gates. Manual ship checks are tracked per
candidate with the
release checklist issue template.
typewhisper-demo-full-with-voiceover.mp4
The localized macOS screenshot workflow is documented in docs/screenshot-automation.md.
- iCloud sync and a redesigned History - History and Inbox sync between your Macs through TypeWhisper's private iCloud container, and History groups entries by device with filters and search
- Undo and recovery - Undo Last Dictation and Restore Raw Transcript are available from the menu bar and as global hotkeys, and Dictation Recovery keeps the last three successful dictations for up to 24 hours so an incomplete provider response can be retried
- Shortcuts and cancellation - Shortcuts are set on a visual Mac keyboard that follows the active input source and shows conflicts, and cancellation can use double Escape, single Escape, immediate cancellation, or be disabled
- Indicator themes - Classic, Glass, and Light indicators with a live preview on the new Appearance settings page
- Import and export - Vocabulary import from Wispr Flow, Handy, and
compatible CSV files, export or deletion of all app data in Advanced
settings, and settings backups from the CLI with
typewhisper exportandtypewhisper import - More ways to transcribe - File transcription from Finder, web media through the optional Web Link plugin, custom local speech models from Hugging Face or a local folder, the Canary ASR plugin, and Parakeet Ultra
- Text handling - A minimum threshold for number formatting, dictionary corrections before workflow LLM processing, and segmented processing for long dictations
- Local and cloud engines - Choose from WhisperKit, Parakeet (TDT v3 and Ultra), Apple SpeechAnalyzer, Granite Speech, Qwen3 ASR, Voxtral, Cohere Transcribe, Groq Whisper, OpenAI Whisper, Soniox, Smallest Pulse, xAI/Grok STT, OpenAI Compatible, and additional bundled or community providers
- On-device or cloud - Keep audio on your Mac with a local engine, or explicitly configure a cloud API for faster processing and additional model choices
- Streaming preview - See partial transcription in real-time while speaking (WhisperKit)
- Short-clip handling - Better retention of brief utterances and fewer false no-speech discards
- File transcription - Batch-process multiple audio/video files with drag & drop, or start from Finder
- Custom local models - Import compatible speech models from a Hugging Face repository or a local folder in the settings of local plugins
- Subtitle export - Export transcriptions as SRT or WebVTT with timestamps
- System-wide - Push-to-talk, toggle, or hybrid mode via global hotkey, auto-pastes into any app
- Modifier-key hotkeys - Use a single modifier key (Command, Shift, Option, Control) as your hotkey
- Last-transcription actions - Copy or paste your latest transcription with configurable global hotkeys
- Undo and raw restore - Undo Last Dictation removes the last inserted dictation, and Restore Raw Transcript replaces it with the unprocessed text, without touching the clipboard
- Dictation Recovery - Keeps failed recordings and the last three successful dictations for up to 24 hours so they can be retried
- Cancellation - Cancel with double Escape, single Escape, or immediately, or turn cancellation off
- Visual shortcut editor - Set shortcuts on a Mac keyboard that follows the active input source and shows conflicts and unavailable keys
- Indicator styles - Choose Notch, Overlay, or Minimal in Classic, Glass, or Light, with optional live transcript preview where supported
- Sound feedback - Audio cues for recording start, transcription success, and errors
- Microphone selection - Choose a specific input device with live preview and improved recovery after route changes
- Workflows - Build reusable transformations for translation, rewriting, extraction, formatting, and app-specific automation. Workflows can run automatically by app, website, or app + website combinations, from a dedicated hotkey, as a global fallback, or manually from the Workflow Palette. Hotkey workflows can either start dictation or process the current selection/clipboard directly.
- LLM provider fallbacks - Order Apple Intelligence (macOS 26+), Groq, OpenAI / ChatGPT, xAI/Grok, Gemini, OpenAI Compatible, and local providers in one global provider/model list. Prompts and workflows inherit that order by default; a workflow with an explicit provider stays on that single provider
- Speech providers - System voices, xAI/Grok TTS, and experimental local Supertonic TTS can provide spoken feedback and readback
- Local prompt processing - The Local LLM (MLX) plugin runs on-device on Apple Silicon. It recommends the Gemma 4 E2B/E4B 4-bit models and offers Qwen3.5 2B and LFM2.5 2.6B as experimental lighter models
- Translation - Translate transcriptions on-device using Apple Translate
- Workflow triggers - Per-app, per-website, combined app + website, hotkey, global fallback, and manual palette-only triggers for language, task, engine, prompt, and auto-submit behavior. Website matching supports subdomains
- Dictionary - Terms improve cloud recognition accuracy. Corrections fix common transcription mistakes automatically. Auto-learns high-confidence local single-word manual corrections, while broader rewrites and deletions are skipped. Includes importable term packs and vocabulary import from Wispr Flow, Handy, and compatible CSV files
- Localized term packs - Built-in term pack names and descriptions are localized in English and German
- Snippets - Text shortcuts with trigger/replacement. Supports placeholders like
{{DATE}},{{TIME}}, and{{CLIPBOARD}} - History - Searchable transcription history with inline editing, correction detection, app context tracking, timeline grouping, device grouping, filters, bulk delete, multi-select export, auto-retention, and a standalone window accessible from the tray menu
- Meeting automation - Connect selected calendars for local reminders, start and stop countdowns, optional automatic Recorder sessions, and structured transcript output for supported meeting providers
- Correction learning - Learn deliberate manual corrections in supported target apps using conservative local matching and dictionary integration
- iCloud sync - Sync History and Inbox between your Macs automatically through TypeWhisper's private iCloud container
- Cloud Folder Sync - Sync Dictionary and Snippets data through a user-selected iCloud Drive, Dropbox, OneDrive, Syncthing, or custom folder
- Clear entitlement states - The Premium hub shows account access and the exact availability of each feature for the current license or signed-in Premium account
- Plugin system - Extend TypeWhisper with custom LLM providers, transcription engines, TTS providers, post-processors, memory providers, and action plugins. Bundled and registry integrations include local engines, major cloud providers, MCP Client, Obsidian, Linear, Script Runner, Webhook Notifications, and additional automation tools. See the plugin catalog
- Local model download controls - Bundled Qwen3, Granite, Voxtral, and Supertonic plugins support an optional HuggingFace token for higher rate limits and clearer download errors. Supertonic requires explicit OpenRAIL-M model-license acceptance before model assets download.
- HTTP API - Local REST API for integration with external tools and scripts
- CLI tool - Shell-friendly transcription and settings backup via the command line
- Discord claim service - Optional external service for Polar supporter and GitHub Sponsors Discord role claims
- Home dashboard - Usage statistics, activity chart, and onboarding tutorial
- Statistics and backups - Inspect local aggregate usage and export or restore supported settings and user data without uploading them to TypeWhisper
- Your data - Export all app data as a ZIP or delete it in Advanced settings
- Auto-update - Built-in updates via Sparkle with stable, release-candidate, and daily channels
- Universal binary - Runs natively on Apple Silicon and Intel Macs
- Widgets - Desktop widgets for usage stats, last transcription, activity chart, and transcription history
- Multilingual UI - English, German, Japanese, and Simplified Chinese
- Launch at Login - Start automatically with macOS
brew install --cask typewhisper/tap/typewhisperDownload the latest DMG from GitHub Releases.
Stable direct-download releases use the default Sparkle channel. Release candidates and daily builds are published as GitHub prereleases, update the shared Sparkle appcast on their own channels, and are excluded from Homebrew.
Installed builds can switch channels in Settings -> About via the Update Channel picker.
- Install TypeWhisper from Homebrew or the latest DMG.
- Open Settings and grant Microphone plus Accessibility access (System Settings > Privacy & Security > Accessibility, named Device Control and Data Access on macOS 27 and later).
- Pick an engine and, if needed, download a local model.
- Trigger the global hotkey and complete your first dictation.
These steps are for official TypeWhisper release builds on macOS. They remove the app itself, its local state, widget data, and stored secrets so you can reinstall from a clean slate.
If you installed via Homebrew, you can optionally start with:
brew uninstall --cask typewhisperThat removes the app bundle, but it does not reliably remove all files in ~/Library or TypeWhisper entries in Keychain.
If ~/Library is hidden in Finder, use Go -> Go to Folder... and paste the paths below.
- Quit TypeWhisper if it is running.
- Delete the app bundle:
rm -rf /Applications/TypeWhisper.app
- Delete app data and plugins:
rm -rf ~/Library/Application\ Support/TypeWhisper
- Delete preferences:
rm -f ~/Library/Preferences/com.typewhisper.mac.plist - Delete widget and app group data used by official releases:
rm -rf ~/Library/Group\ Containers/2D8ALY3LCL.com.typewhisper.mac
- Remove TypeWhisper secrets from Keychain:
- In Keychain Access, search for
com.typewhisper.mac.apikeyand delete matching items. - This includes API and plugin secrets stored under the
com.typewhisper.mac.apikey.*service prefix. - Also remove the license items stored under service
com.typewhisper.mac.apikey.license, especially thepolar-licenseandpolar-supporteraccounts.
- In Keychain Access, search for
- If you installed the CLI tool from Settings > Advanced, remove it too:
rm -f /usr/local/bin/typewhisper
- Optional: if you want to remove exported user files as well, delete:
rm -rf ~/Documents/TypeWhisper\ Recordings
- Restart your Mac, then install the latest build again.
If a fresh install still crashes immediately after these steps, please open an issue and include your macOS version, how you installed TypeWhisper, and whether the crash happens on first launch or after granting permissions.
- macOS 14.0 (Sonoma) or later
- Apple Silicon (M1 or later) recommended
- 8 GB RAM minimum, 16 GB+ recommended for larger models
- Some features (Apple Translate, improved Settings UI) require macOS 15+. Apple Intelligence and SpeechAnalyzer require macOS 26+.
TypeWhisper includes the Local LLM (MLX) plugin for on-device prompt processing on Apple Silicon. It recommends the dense Gemma 4 E2B 4-bit and E4B 4-bit models. Qwen3.5 2B and LFM2.5 2.6B need less memory and are available as experimental models; larger Gemma 4 variants are experimental as well. The plugin replaces the former Gemma 4 plugin (com.typewhisper.gemma4) but installs separately: existing Gemma 4 installations do not update to it, and downloaded models, the HuggingFace token, and the model selection are not carried over. Workflows and prompt actions that used Gemma 4 switch to Local LLM (MLX) once the Gemma 4 plugin is removed.
| RAM | Recommended Models |
|---|---|
| < 8 GB | Whisper Tiny, Whisper Base |
| 8-16 GB | Whisper Small, Whisper Large v3 Turbo, Parakeet TDT v3, Voxtral Mini 4B |
| > 16 GB | Whisper Large v3 |
-
Clone the repository:
git clone https://github.com/TypeWhisper/typewhisper-mac.git cd typewhisper-mac -
Open in Xcode 16+:
open TypeWhisper.xcodeproj
-
Select the TypeWhisper scheme and build (Cmd+B). Swift Package dependencies (WhisperKit, FluidAudio, Sparkle, TypeWhisperPluginSDK) resolve automatically.
-
Run the app. It appears as a menu bar icon - open Settings to download a model.
-
Run the automated checks before shipping changes:
xcodebuild test -project TypeWhisper.xcodeproj -scheme TypeWhisper -destination 'platform=macOS,arch=arm64' -parallel-testing-enabled NO CODE_SIGN_IDENTITY='-' CODE_SIGNING_REQUIRED=NO CODE_SIGNING_ALLOWED=NO swift test --package-path TypeWhisperPluginSDK
The HTTP API is an advanced local automation surface. It binds to 127.0.0.1 only, is disabled by default, and is intended for local tools and scripts.
Enable the API server in Settings > Advanced (default port: 8978).
The API server requires a bearer token by default. TypeWhisper creates the token once, keeps it in the Keychain, and writes it to an owner-only discovery file while the server runs. The command line tool, the Raycast extension, the MCP server, and the Pi extension read it automatically. Scripts can read it like this:
TYPEWHISPER_API_TOKEN="$(jq -r '.token' "$HOME/Library/Application Support/TypeWhisper/api-discovery.json")"
curl -H "Authorization: Bearer $TYPEWHISPER_API_TOKEN" http://localhost:8978/v1/modelsThe X-TypeWhisper-API-Token header works as well. Every endpoint except GET /v1/status needs the token; the examples below leave the header out for brevity. Settings > Advanced > Copy API Token copies it to the clipboard.
If the API server is turned on without a token when you update to 1.8.0, it keeps running without one until Require API Token is turned on, and Settings shows a warning. A server that was off during the update requires the token once it is turned on.
Independent of the token, the server rejects requests that browsers send on behalf of other websites: a Host header other than 127.0.0.1, localhost, or [::1], an Origin from a website not served from this Mac (including null), and Sec-Fetch-Site: cross-site unless the Origin is a page served from this Mac. Such requests get 403. Requests without these headers, such as those from curl or scripts, are not affected.
curl http://localhost:8978/v1/status{
"status": "ready",
"engine": "whisper",
"model": "openai_whisper-large-v3_turbo",
"supports_streaming": true,
"supports_translation": true
}curl -X POST http://localhost:8978/v1/transcribe \
-F "file=@recording.wav" \
-F "language=en"
curl -X POST http://localhost:8978/v1/transcribe \
-F "file=@recording.wav" \
-F "language_hint=de" \
-F "language_hint=en"{
"text": "Hello, world!",
"language": "en",
"duration": 2.5,
"processing_time": 0.8,
"engine": "whisper",
"model": "openai_whisper-large-v3_turbo"
}Optional parameters:
language- ISO 639-1 code (e.g.,en,de). Omit for full auto-detection.language_hint- Repeatable, ordered language hint for restricted auto-detection. Hint-aware engines receive the full list; other engines use the first hint as the requested language. Do not combine withlanguage.task-transcribe(default) ortranslate(translates to English, WhisperKit only).target_language- ISO 639-1 code for translation target language (e.g.,es,fr). Uses Apple Translate.apply_corrections- Boolean, defaulttrue. Set tofalseto return raw transcription text without Dictionary Corrections. For raw body uploads, sendx-apply-corrections: false.detect_speakers- Boolean, defaultfalse. Adds aspeaker(Speaker 1,Speaker 2, …) to each segment of averbose_jsonresponse. Needs Premium. For raw body uploads, sendx-detect-speakers: true.speaker_count- Number of speakers, when known. Omit it to detect the number. For raw body uploads, sendx-speaker-count.
Uploads to /v1/transcribe are limited to 256 MiB, including stdin uploads from the CLI. Requests above that size return 413 Payload Too Large. Local CLI file paths use a direct handoff to the running TypeWhisper app instead of uploading the file bytes.
curl http://localhost:8978/v1/models{
"models": [
{
"id": "openai_whisper-large-v3_turbo",
"engine": "whisper",
"name": "Large v3 Turbo",
"status": "ready",
"selected": true,
"downloaded": true,
"loaded": true
}
]
}Download (if needed), load, and select a model. The request returns after the model is ready or the load fails:
curl -X POST http://localhost:8978/v1/models/load \
-H "Content-Type: application/json" \
-d '{"engine":"whisper","model":"openai_whisper-large-v3_turbo"}'Unload the active model while keeping its downloaded files:
curl -X POST http://localhost:8978/v1/models/unload \
-H "Content-Type: application/json" \
-d '{"engine":"whisper"}'Delete downloaded model files:
curl -X DELETE "http://localhost:8978/v1/models?engine=whisper&model=openai_whisper-large-v3_turbo"Deleting the final downloaded model for a plugin disables that plugin. If API authentication is enabled, include the same bearer token used by the other protected endpoints.
# Search history
curl "http://localhost:8978/v1/history?q=meeting&limit=10&offset=0"
# Delete entry
curl -X DELETE "http://localhost:8978/v1/history?id=<uuid>"Entries with speaker detection carry speaker_state. Once a speaker transcript exists, they also carry speakers. Each speaker has an id and a name: the confirmed name, or Speaker N. When a voice profile suggested a name that nobody has confirmed yet, it is returned as suggested_name, and name stays Speaker N. Add include=speaker_segments to get the speaker segments. Entries that came from the Recorder carry the recording_id used by GET /v1/recorder/recordings and the .transcript-ready.json sidecar.
# List recognition terms
curl http://localhost:8978/v1/dictionary/terms
# Merge terms, or set replace=true to replace all terms
curl -X PUT http://localhost:8978/v1/dictionary/terms \
-H "Content-Type: application/json" \
-d '{"terms":["TypeWhisper","WhisperKit"],"replace":false}'
# Delete one term
curl -X DELETE http://localhost:8978/v1/dictionary/terms \
-H "Content-Type: application/json" \
-d '{"term":"TypeWhisper"}'
# List post-transcription corrections
curl http://localhost:8978/v1/dictionary/corrections
# Add or update one correction by original text
curl -X PUT http://localhost:8978/v1/dictionary/corrections \
-H "Content-Type: application/json" \
-d '{"original":"teh","replacement":"the","caseSensitive":false}'
# Delete one correction
curl -X DELETE http://localhost:8978/v1/dictionary/corrections \
-H "Content-Type: application/json" \
-d '{"original":"teh"}'The settings endpoints use the same JSON backup schema and merge/skip behavior as Settings > Advanced > Backup & Restore. Import applies every category present in the backup. These examples include the bearer token from the discovery file (see Authentication):
TYPEWHISPER_API_TOKEN="$(jq -r '.token' "$HOME/Library/Application Support/TypeWhisper/api-discovery.json")"
# Export the current settings backup
(
settings_backup_tmp="$(mktemp ./typewhisper-settings.json.tmp.XXXXXX)" || exit
trap 'rm -f "$settings_backup_tmp"' EXIT
curl --fail --silent --show-error http://localhost:8978/v1/settings/export \
-H "Authorization: Bearer $TYPEWHISPER_API_TOKEN" \
--output "$settings_backup_tmp" && \
mv "$settings_backup_tmp" typewhisper-settings.json
)
# Import all categories from a settings backup
curl --fail --silent --show-error -X POST http://localhost:8978/v1/settings/import \
-H "Authorization: Bearer $TYPEWHISPER_API_TOKEN" \
-H "Content-Type: application/json" \
--data-binary @typewhisper-settings.json
# Import and overwrite workflows, profiles, prompt actions with the same name and hotkeys
curl --fail --silent --show-error -X POST "http://localhost:8978/v1/settings/import?mode=replace" \
-H "Authorization: Bearer $TYPEWHISPER_API_TOKEN" \
-H "Content-Type: application/json" \
--data-binary @typewhisper-settings.jsonAn import skips history entries, workflows, profiles, and prompt actions that already exist unchanged, so importing a backup onto the Mac it came from does not duplicate them. The default mode=merge adds everything else and only fills empty hotkey slots. mode=replace also overwrites existing workflows, profiles, and prompt actions with the same name and the hotkeys contained in the backup. Neither mode deletes anything.
GET /v1/settings/audio returns the microphone priority list and the recording audio options from Settings > Dictation. PATCH /v1/settings/audio changes any subset of them through the same code path as the settings window: the change applies to the next recording, persists across restarts, and shows up in the open settings window. TypeWhisper for Windows serves the same contract.
# Show available inputs, the priority list, and the input the next recording would use
curl http://localhost:8978/v1/settings/audio
# Record from BlackHole and turn off ducking
curl -X PATCH http://localhost:8978/v1/settings/audio \
-H "Content-Type: application/json" \
-d '{"input_priority":[{"id":"BlackHole2ch_UID","name":"BlackHole 2ch"}],"audio_ducking_enabled":false}'
# Use the macOS default input
curl -X PATCH http://localhost:8978/v1/settings/audio \
-H "Content-Type: application/json" \
-d '{"input_priority":[]}'{
"input_devices": [{"id": "BlackHole2ch_UID", "name": "BlackHole 2ch", "is_system_default": false}],
"input_priority": [{"id": "AppleUSBAudioEngine:HyperX:QuadCast 2", "name": "HyperX QuadCast 2"}],
"active_input": {"id": "BuiltInMicrophoneDevice", "name": "MacBook Pro Microphone"},
"audio_ducking_enabled": true,
"audio_ducking_level": 0.2,
"pause_media_during_recording": false,
"sound_feedback_enabled": true
}| Field | Writable | Description |
|---|---|---|
input_devices |
no | Input devices that are connected now. id is the CoreAudio device UID. |
input_priority |
yes | The saved priority list, including devices that are not connected. Each entry needs an id; name is optional and is replaced by the current name while the device is connected. [] records from the macOS default input. |
active_input |
no | The input the next recording would use: the first connected entry of input_priority, otherwise the macOS default input. null when there is no input. |
audio_ducking_enabled |
yes | Reduce the system volume during recording. |
audio_ducking_level |
yes | Share of the current volume kept during recording, from 0 (mute) to 1. The settings slider offers 0 to 0.5. |
pause_media_during_recording |
yes | Pause media playback during recording. |
sound_feedback_enabled |
yes | Play start, success, and error sounds. |
A successful PATCH returns the full state, as GET does. It returns 400 for invalid JSON, a value of the wrong type or range, a duplicate or empty device ID, a read-only field, or an unknown field; in that case nothing changes. TypeWhisper for macOS supports every writable field of the shared contract. A field that only another platform offers is an unknown field here, and the error message names it together with the fields macOS accepts. While a dictation or the recorder is recording or still processing, PATCH returns 409 and changes nothing; GET keeps working.
To change the settings temporarily, save the GET response, send your changes, and later send the saved writable fields back:
saved="$(curl --fail --silent http://localhost:8978/v1/settings/audio)"
# ... change settings, record ...
jq 'del(.input_devices, .active_input)' <<<"$saved" | \
curl --fail --silent -X PATCH http://localhost:8978/v1/settings/audio \
-H "Content-Type: application/json" --data-binary @-# List all workflow-backed rules
curl http://localhost:8978/v1/rules
# Toggle a workflow-backed rule on/off
curl -X PUT "http://localhost:8978/v1/rules/toggle?id=<uuid>"# Start dictation (returns session id)
curl -X POST http://localhost:8978/v1/dictation/start
# Stop dictation (returns same session id)
curl -X POST http://localhost:8978/v1/dictation/stop
# Check whether dictation is currently recording
curl http://localhost:8978/v1/dictation/status
# Fetch status/result for a specific dictation session
curl "http://localhost:8978/v1/dictation/transcription?id=<uuid>"Dictation control records microphone audio for system-wide insertion. A completed dictation session returns text that TypeWhisper can paste back into the active app.
For the last 100 dictations since the app started, the transcription response also contains a latency object with the timings of each dictation phase, from the start request to verified insertion and clipboard restoration. It never contains transcript text, audio or app content. Other sessions return latency: null.
Recorder control uses the same recorder path as the TypeWhisper UI, including microphone capture, optional system audio capture, mixing, finalization, and final transcription. Use it for automations that need a saved recording file or meeting/system-audio transcription without auto-pasting into another app.
# Start recorder with microphone and system audio
curl -X POST "http://localhost:8978/v1/recorder/start?mic=true&system_audio=true"
# Stop the active API recorder session
curl -X POST http://localhost:8978/v1/recorder/stop
# Check whether the recorder is currently recording
curl http://localhost:8978/v1/recorder/status
# Fetch status/result for a specific recorder session
curl "http://localhost:8978/v1/recorder/session?id=<uuid>"POST /v1/recorder/start accepts optional query flags:
mic-true,false,1, or0. If omitted, TypeWhisper uses the current recorder microphone setting.system_audio-true,false,1, or0. If omitted, TypeWhisper uses the current recorder system-audio setting.
At least one source must be enabled. If both resolved sources are disabled, the API returns 400 Bad Request.
Start response:
{
"id": "8F8C1F45-6D03-44D2-A38C-0C4DE4F7E5F7",
"status": "recording"
}Stop response:
{
"id": "8F8C1F45-6D03-44D2-A38C-0C4DE4F7E5F7",
"status": "finalizing"
}Status response:
{
"recording": true
}Session response:
{
"id": "8F8C1F45-6D03-44D2-A38C-0C4DE4F7E5F7",
"status": "completed",
"text": "Meeting notes from the recording.",
"output_file": "/Users/alex/Documents/TypeWhisper Recordings/Recording 2026-05-20 14-30-00.m4a"
}Recorder sessions move through recording -> finalizing -> completed or failed. If recorder transcription is disabled or produces no transcript, text is omitted and output_file still points to the finalized recording when available. Failed sessions include an error field.
Conflict and lookup behavior:
- Starting while the recorder is already recording or finalizing returns
409 Conflict. - Stopping without an active API recorder session returns
409 Conflict. - Polling with a missing or invalid
idreturns400 Bad Request. - Polling a valid but unknown session id returns
404 Not Found.
GET /v1/recorder/recordings returns the latest successfully saved transcript for each recording, including recordings started manually, by the calendar integration, or through the API. Results survive app restarts and do not depend on live preview or an API session.
curl "http://localhost:8978/v1/recorder/recordings?since=1791100000.123456"{
"recordings": [
{
"source": "recorder",
"recording_id": "8F8C1F45-6D03-44D2-A38C-0C4DE4F7E5F7",
"completion_id": "59CBC40C-274A-47A9-804C-8774AB504401",
"completed_at": 1791100000.123456,
"text": "Meeting notes from the recording.",
"audio_file": "/Users/alex/Documents/TypeWhisper Recordings/Meeting.m4a",
"transcript_file": "/Users/alex/Documents/TypeWhisper Recordings/Meeting.txt"
}
]
}since accepts Unix seconds or an ISO 8601 timestamp and filters inclusively by successful completion time, not the recording's start time. Results are ordered oldest completion first. After processing a response, retain its last completed_at and deduplicate by completion_id; querying inclusively avoids losing completions with equal timestamps. Invalid timestamps return 400; unreadable completion receipts return 500 so a consumer does not silently advance past them.
Each successful retranscription keeps the recording_id, assigns a new completion_id, and replaces the previous result. This is a list of the latest saved results, not a revision history. Failed attempts leave the last successful result available. markdown_file is included only when a Markdown transcript was saved and still exists; currently that applies to calendar recordings with meeting metadata. Audio and transcript paths refer to existing files.
Completion receipts are saved alongside recordings as <audio filename>.transcript-ready.json. A separate <audio filename>.recording-id.json preserves the recording ID if the receipt is damaged, so retranscription can repair it without changing the ID. Both files are committed with the transcript and removed when the recording is deleted in TypeWhisper. Existing recordings become available here after their next successful transcription. Recordings without a saved transcript are omitted. The endpoint uses the same API token as the other private routes.
The Plugin SDK emits recorderTranscriptReady after the transcript and receipt have been saved. It does not emit the dictation transcriptionCompleted event or run Recorder Workflows. Event delivery is best effort while TypeWhisper and the plugin are running; use the API to catch up on missed completions.
In Webhook Notifications, enable Also send completed Recorder transcripts for each destination that should receive this event. The JSON body has the same fields as an entry above. Webhook retries keep the same completion_id.
In Script Runner, enable Also run for completed Recorder transcripts for each export command. Both options default to off, including for existing configurations, and operate independently of dictation rule/workflow filters. Recorder scripts receive the saved original through stdin and these environment variables:
| Variable | Value |
|---|---|
TYPEWHISPER_SOURCE |
recorder |
TYPEWHISPER_RECORDING_ID |
Stable recording UUID |
TYPEWHISPER_COMPLETION_ID |
UUID for this successful save |
TYPEWHISPER_COMPLETED_AT |
Completion time in Unix seconds |
TYPEWHISPER_AUDIO_FILE |
Audio file path |
TYPEWHISPER_TRANSCRIPT_FILE |
Plain-text transcript path |
TYPEWHISPER_MARKDOWN_FILE |
Markdown path, when available; otherwise unset |
Recorder scripts export the original transcript independently; stdout is ignored and may be empty. A nonzero exit status is logged as a failure. The existing five-second script timeout applies, so enqueue longer processing in a separate worker. Ordinary dictation scripts continue to transform text through stdout.
TypeWhisper includes a command-line tool for shell-friendly transcription. It is part of the advanced automation surface and connects to the running local API server.
Install via Settings > Advanced > CLI Tool > Install. This places the typewhisper binary in /usr/local/bin.
typewhisper status # Show server status
typewhisper models # List available models
typewhisper transcribe file.wav # Transcribe an audio file
typewhisper export settings.json # Export all supported settings
typewhisper import settings.json # Import all categories in a backup
typewhisper audio # Show microphone priority, ducking and sound settings
typewhisper audio set changes.json # Change audio settings (see Audio Settings above)| Option | Description |
|---|---|
--port <N> |
Server port (default: auto-detect) |
--json |
Output as JSON |
--language <code> |
Source language (e.g. en, de) |
--language-hint <code> |
Repeatable, ordered language hint for restricted auto-detection; engines without hint support use the first hint |
--task <task> |
transcribe (default) or translate |
--translate-to <code> |
Target language for translation |
--no-corrections |
Return raw transcription text without Dictionary Corrections |
--replace |
import only: overwrite workflows, profiles, and prompt actions with the same name and replace hotkeys |
# Transcribe with language and JSON output
typewhisper transcribe recording.wav --language de --json
# Restrict auto-detection to a shortlist
typewhisper transcribe recording.wav --language-hint de --language-hint en
# Pipe audio from stdin
cat audio.wav | typewhisper transcribe -
# Use in a script
typewhisper transcribe meeting.m4a --json | jq -r '.text'
# Keep a portable backup in a dotfiles repository
typewhisper export ~/.config/typewhisper/settings.json
# Restore it and receive a machine-readable import summary
typewhisper import ~/.config/typewhisper/settings.json --json
# Apply edits made to that file, including changed hotkeys
typewhisper import ~/.config/typewhisper/settings.json --replace
# Record from BlackHole until you switch back
echo '{"input_priority":[{"id":"BlackHole2ch_UID"}]}' | typewhisper audio set -The CLI requires the API server to be running (Settings > Advanced) and follows the documented command and flag surface for the current stable release.
Dictionary Corrections apply to CLI transcriptions by default. Use --no-corrections when a script needs the raw engine output.
Local file paths are handed to the running TypeWhisper app directly, so large files do not need to fit inside an HTTP upload body. Stdin usage (typewhisper transcribe -) still uses the regular /v1/transcribe upload endpoint and is limited to 256 MiB.
Workflows let you configure transcription, transformation, and automation behavior per application, website, combined app + website context, hotkey, global fallback, or manual palette-only workflow. For example:
- Mail - German language, Whisper Large v3
- Slack - English language, Parakeet TDT v3
- Terminal - English language, auto-submit enabled
- github.com - English cleanup workflow that matches in any browser
- docs.google.com - German dictation workflow that translates to English
Create workflows in Settings > Workflows. Choose a template, then use Automatic to enable app, website, hotkey, or any combination of those trigger components. Always stays the global fallback, and Manual keeps the workflow palette-only. Hotkey workflows choose whether the shortcut starts dictation or processes the current selection/clipboard through the same insertion path as the Workflow Palette. Spoken language can be left on full auto-detect, fixed to one exact language, or restricted to a shortlist of likely languages for better detection accuracy. Website patterns support subdomain matching - e.g. google.com also matches docs.google.com.
LLM Prompt workflows inherit the ordered global LLM fallback list shown at the top of Settings > Workflows. TypeWhisper moves to the next provider/model when an inherited attempt is unavailable, cannot restore, is rate-limited, hits a network or API error, or returns no text. Selecting a provider inside a workflow disables that fallback behavior for the workflow and makes one strict provider call instead.
When you start dictating, TypeWhisper matches the active app and browser URL against enabled workflows with the following priority:
- App + URL match - highest specificity (e.g. Chrome + github.com)
- URL-only match - cross-browser workflows (e.g. github.com in any browser)
- App-only match - generic app workflows (e.g. all of Chrome)
- Always fallback - global workflow when no more specific workflow matches
Hotkeys are direct workflow shortcuts, not context conditions in the app/URL matching order. Manual workflows are excluded from automatic dictation matching. They appear only in the Workflow Palette and use the existing Workflow Palette hotkey.
The active workflow name is shown as a badge in the indicator, together with a short explanation of why it matched.
Multiple engines can be loaded simultaneously during a session for instant switching between workflows. Note that loading multiple local models increases memory usage. Set Auto-unload model to Never if you want TypeWhisper to restore previously loaded local models at launch; any active auto-unload policy keeps startup lazy and reloads the selected/downloaded model on first real use. Cloud engines (Groq, OpenAI, xAI/Grok) have negligible memory overhead.
TypeWhisper supports plugins for adding custom LLM providers, transcription engines, TTS providers, post-processors, and action plugins. Plugins are macOS .bundle files placed in ~/Library/Application Support/TypeWhisper/Plugins/.
Bundled engines and integrations (WhisperKit, Parakeet, SpeechAnalyzer, Granite, Qwen3, Voxtral, Supertonic, Groq, OpenAI, xAI/Grok, OpenAI Compatible, Gemini, Linear, Webhook, and more) are implemented as plugins and serve as reference implementations.
See TypeWhisperPluginSDK/Plugins/README.md for the full plugin development guide, including the event bus, host services API, and manifest format.
TypeWhisper/
├── typewhisper-cli/ # Command-line tool (transcription and settings backup)
├── PluginRegistry/ # Source registry entries for community plugin feeds
├── Plugins/ # Redirect docs and legacy entrypoint for moved first-party plugin sources
├── TypeWhisperPluginSDK/ # Plugin SDK (Swift package)
│ ├── Plugins/ # First-party plugin sources and manifests
├── TypeWhisperWidgetExtension/ # WidgetKit widgets (stats, activity, history)
├── TypeWhisperWidgetShared/ # Shared widget data models
├── App/ # App entry point, dependency injection
├── Models/ # Data models (TranscriptionResult, Profile, PromptAction, etc.)
├── Services/
│ ├── Cloud/ # KeychainService, WavEncoder (shared cloud utilities)
│ ├── LLM/ # Apple Intelligence provider (cloud LLM providers are plugins)
│ ├── HTTPServer/ # Local REST API (HTTPServer, APIRouter, APIHandlers)
│ ├── ModelManagerService # Transcription dispatch (delegates to plugins)
│ ├── AudioRecordingService
│ ├── AudioFileService # Audio/video - 16kHz PCM conversion
│ ├── HotkeyService
│ ├── TextInsertionService
│ ├── WorkflowService # Workflow matching and persistence
│ ├── HistoryService # Transcription history persistence (SwiftData)
│ ├── DictionaryService # Custom term corrections
│ ├── SnippetService # Text snippets with placeholders
│ ├── PromptActionService # Prompt action persistence (SwiftData)
│ ├── PromptProcessingService # LLM orchestration for prompt execution
│ ├── PluginManager # Plugin discovery, loading, and lifecycle
│ ├── PluginRegistryService # Plugin marketplace (download, install, update)
│ ├── PostProcessingPipeline # Priority-based text processing chain
│ ├── EventBus # Typed publish/subscribe event system
│ ├── TranslationService # On-device translation via Apple Translate
│ ├── SubtitleExporter # SRT/VTT export
│ └── SoundService # Audio feedback for recording events
├── ViewModels/ # MVVM view models with Combine
├── Views/ # SwiftUI views
└── Resources/ # Info.plist, entitlements, localization, sounds
Patterns: MVVM with ServiceContainer singleton for dependency injection. ViewModels use a static _shared pattern. Localization via String(localized:) with Localizable.xcstrings.
GPLv3 - see LICENSE for details. Commercial licensing available - see LICENSE-COMMERCIAL.md.

























