Search...Search plugins and themes...
⌘K
Sign in
  • Get started
  • Download
  • Pricing
  • Enterprise
  • Account
  • Obsidian
  • Overview
  • Sync
  • Publish
  • Canvas
  • Mobile
  • Web Clipper
  • CLI
  • Learn
  • Help
  • Developers
  • Changelog
  • About
  • Roadmap
  • Blog
  • Resources
  • System status
  • License overview
  • Terms of service
  • Privacy policy
  • Security
  • Community
  • Plugins
  • Themes
  • Discord
  • Forum / 中文论坛
  • Merch store
  • Brand guidelines
Follow us
DiscordTwitterBlueskyThreadsMastodonYouTubeGitHub
© 2026 Obsidian

Speech Kit

Alexander BrittainAlexander Brittain4k downloads

Offline and private speech-to-text, text-to-speech, and translation for desktop notes. Dictate notes, transcribe meetings, and more. NVIDIA, Whisper, and natural voices running on your machine.

Add to Obsidian
Speech Kit screenshot
  • Overview
  • Scorecard
  • Updates37
Speech Kit — Speech and language toolkit for Obsidian

Dictate live. Transcribe audio and video files. Bring YouTube captions into your notes. Translate text. Listen to notes. One plugin inside the editor where your notes already live.

Local Dictation is now Speech Kit. It is the same plugin with the same local-first foundation, now with a name that fits what it has become. Existing installs, settings, and hotkeys carry over automatically.

Install Speech Kit from Obsidian Community Plugins

What it does

  • Dictate: Capture long sessions with live text and a choice of speech models.
  • Transcribe files: Drop in audio or video and turn its audio into a local transcript.
  • Import YouTube captions: Paste a link and bring available captions into your note without downloading the video.
  • Translate and listen: Translate notes locally and hear them read aloud with natural voices.
  • Refine with AI: Run a custom preset on a complete transcript or note using a local model or your chosen provider.

Speech Kit translating an Obsidian note from English to Spanish and replacing the original text

Why Speech Kit?

Speech and language tools are usually fragmented. One tool handles dictation. Another transcribes meetings. Another reads text aloud. Another translates. Each brings its own settings, models, and hotkeys, and often its own cloud account, subscription, and privacy policy.

Speech Kit replaces that stack with one consistent workflow inside Obsidian: one model manager, one settings surface, and one set of commands.

Dictate an idea. Capture a meeting. Translate a passage. Listen to a note. Refine the result. It all happens inside the editor where your notes already live.

Choose the models that fit your workflow

Speech Kit is not tied to one speech engine or hosted API. It manages a growing catalog of models. Install only what you need, mix and match, and change models as your language, hardware, or priorities change.

You want Choose
Words on screen while you speak Moonshine streaming models
Multilingual live transcription Nemotron 3.5 ASR
The most accurate transcripts Whisper Large V3 Turbo, Cohere Transcribe, and other batch models
Natural local voices Pocket TTS or Supertonic 3
Fast offline translation Firefox Translations

The setup wizard installs the native engine and your first speech model. From there, Speech Kit manages the downloads and you choose how you work.

Dictate, transcribe, translate, listen, and refine

Dictate. Streaming words appear and revise in place while you speak. Finished text lands as Markdown at your cursor. Switch to a batch model when accuracy after each pause matters more than immediacy.

Transcribe files. Run Transcribe local audio file to turn an audio or video recording into text in the current note. Speech Kit extracts the audio with its media decoder and transcribes it with an installed batch speech model. Choose a language, model, timestamps, speaker labels, and paragraph style for that job. Your media stays on your computer.

Import YouTube captions. Run Transcribe YouTube video, paste a video link, and add its available captions to the current note. Choose the caption language and optional linked timestamps. This fast path needs an internet connection, but no speech model or video download. Videos without usable captions leave the note unchanged.

Both commands can optionally run one of your AI presets on the complete transcript. AI is separate from transcription; if you choose a remote provider, transcript text is sent to that provider. See the media transcription guide for setup and options.

Translate. Translate a selection or a whole note between English and seven other languages. Preview the result before replacing your text, inserting it into the note, or copying it. One local model pack covers every supported direction.

Listen. Read any note aloud with natural local voices. Control the voice, speed, and playback without leaving Obsidian.

Refine. Optional LLM tools can clean up, summarize, restructure, or transform text with your own prompts.

One toolkit across platforms

Many speech apps are limited to one operating system, one model, or one part of the workflow. Speech Kit brings the same toolkit to macOS, Windows, and Linux, with hardware acceleration and system-audio capture where available.

Platform Architecture Acceleration System audio
macOS Apple silicon Metal for Whisper macOS 14.2 or later
Windows x86-64, ARM64 (Snapdragon 8cx or newer) Optional NVIDIA CUDA (x86-64 only) Supported
Linux x86-64, ARM64 glibc (ARMv8.2 or newer, e.g. Raspberry Pi 5) Optional NVIDIA CUDA (x86-64 only) PulseAudio or PipeWire

Choose your platform. Choose your models. Keep one workflow inside Obsidian.

Getting started

  1. Install Speech Kit from Community Plugins.
  2. Follow the setup wizard to install the native engine and a speech model.
  3. Select Try dictation now, or start from the ribbon, command palette, or a hotkey.

Dictation, local file transcription, translation, and read aloud require no account, API key, usage credits, or cloud service. Once their models are installed, they continue working offline. YouTube caption import requires a connection to YouTube.

Optional LLM text tools are separate. You can connect a local or remote provider when you choose to use them.

Language support

Each feature is served by a different model, so coverage is tracked per feature rather than as a single list.

Language Transcription Live dictation Read aloud Translation Interface
English, Spanish, German, French, Portuguese, Italian, Dutch, Japanese ✅ ✅ ✅ ✅ ✅
Croatian ✅ ✅ ✅ — ✅
Serbian ✅ — — — —

✅ supported · — not yet available

Transcription coverage also depends on the model you select: multilingual models cover the full set above, while some smaller or specialized models are English-only. Translation runs through English in either direction, so every supported pair has English on one side.

Local-first, private by default

Speech Kit works without accounts, subscriptions, or required cloud services.

  • Your media stays on your machine. Dictation, local file transcription, read aloud, and translation run locally and continue working offline once their models are installed. YouTube caption import retrieves text from YouTube without downloading the video.
  • No account, telemetry, or metered usage. No API key, credit card, subscription, or usage credits to monitor.
  • LLM tools are optional. Add flexible language processing to your workflow using a local model or a remote provider you choose. Text leaves your device only when you explicitly use a remote provider, and audio is never uploaded.
  • Choose what works for you. Install high-quality models suited to your language, hardware, and workflow.
  • Transparent and open. Downloads are explicit, third-party licenses are documented, and Speech Kit is open source.

Support development

If Speech Kit is useful to you, please support development:

Development and project links

Speech Kit pairs a TypeScript plugin with a Rust native sidecar. See CONTRIBUTING.md for its architecture, setup, and development workflow.

Community Plugin · Latest release

Issues · License

Third-party component and model licenses are documented in THIRD_PARTY_NOTICES.md and shown before model download.

HealthExcellent
ReviewSatisfactory
About
Obsidian handles the notes. Speech Kit handles speech and language. Speech Kit - formerly Local Dictation - brings everything you need to work with speech and language into one place. Dictate ideas as they happen. Turn meetings and calls into useful notes. Translate notes across eight languages. Listen to your writing with natural voices. Clean up, summarize, and reshape transcripts when you need to. Built for desktop Obsidian on macOS, Windows, and Linux, Speech Kit lets you choose the models that fit your language, hardware, and priorities. Core dictation, transcription, translation, and read aloud run on your device and keep working offline - no account, subscription, API key, or metered usage. Optional LLM tools can use a local model or a remote provider you choose.
AudioAIWriting
Details
Current version
2026.10.1
Last updated
18 hours ago
Created
6 months ago
Updates
37 releases
Downloads
4k
Compatible with
Obsidian 1.11.5+
Platforms
Desktop only
License
MIT
Report bugRequest featureReport plugin
Author
Alexander BrittainAlexander Brittainbrittain9
GitHubbrittain9
  1. Community
  2. Plugins
  3. Audio
  4. Speech Kit

Related plugins

ChatGPT MD

A seamless integration of ChatGPT, OpenRouter.ai and local LLMs via Ollama into your notes.

Text Generator

Generate text content using GPT-3 (OpenAI).

Local GPT

Local Ollama and OpenAI-like GPT's assistance for maximum privacy and offline access.

Note Companion

AI-powered note organization and chat. Requires subscription or self-hosting with your own API keys.

BMO Chatbot

Generate and brainstorm ideas while creating your notes using Large Language Models (LLMs) such as OpenAI's "gpt-3.5-turbo" and "gpt-4".

GPT-3 Notes

Generate notes on any subject using OpenAI's GPT-3.5 and GPT-4 language models.

Claudian

Embeds Claude Code/Codex and other local Agents as AI collaborators in your vault.

LanguageTool Integration

Advanced grammar and spell checking, powered by LanguageTool.

Smart Connections

Find related notes and excerpts while writing. Your AI link building copilot displays relevant content in graph + list view. A local embedding model powers semantic search. Zero setup. No API key.

Fast Note Sync

Real-time sync of your vaults across server, mobile, and web; shareable with anyone; supports REST and MCP integrations to build your personal AI knowledge base.