Search...Search plugins and themes...
⌘K
Sign in
  • Get started
  • Download
  • Pricing
  • Enterprise
  • Account
  • Obsidian
  • Overview
  • Sync
  • Publish
  • Canvas
  • Mobile
  • Web Clipper
  • CLI
  • Learn
  • Help
  • Developers
  • Changelog
  • About
  • Roadmap
  • Blog
  • Resources
  • System status
  • License overview
  • Terms of service
  • Privacy policy
  • Security
  • Community
  • Plugins
  • Themes
  • Discord
  • Forum / 中文论坛
  • Merch store
  • Brand guidelines
Follow us
DiscordTwitterBlueskyThreadsMastodonYouTubeGitHub
© 2026 Obsidian

Speech Kit

Alexander BrittainAlexander Brittain2k downloads

Offline and private speech-to-text, text-to-speech, and translation for desktop notes. Dictate notes, transcribe meetings, and more. NVIDIA, Whisper, and natural voices running on your machine.

Speech Kit screenshot
  • Overview
  • Scorecard
  • Updates27
Speech Kit — Speech and language toolkit for Obsidian

Dictate live. Transcribe meetings. Translate text. Listen to notes. One plugin inside the editor where your notes already live.

Local Dictation is now Speech Kit. It is the same plugin with the same local-first foundation, now with a name that fits what it has become. Existing installs, settings, and hotkeys carry over automatically.

Install Speech Kit from Obsidian Community Plugins

What it does

  • 🎤 Speech: Dictate with live streaming text, or capture higher-accuracy transcripts from meetings, calls, and other audio.
  • 🔊 Voice: Listen to your notes with natural voices.
  • 🌍 Language: Dictate in ten languages and translate notes locally across eight.
  • 🧠 Models: Choose from a managed catalog of speech, voice, and translation models, with optional LLM text tools.

Speech Kit translating an Obsidian note from English to Spanish and replacing the original text

Why Speech Kit?

Speech and language tools are usually fragmented. One tool handles dictation. Another transcribes meetings. Another reads text aloud. Another translates. Each brings its own settings, models, and hotkeys, and often its own cloud account, subscription, and privacy policy.

Speech Kit replaces that stack with one consistent workflow inside Obsidian: one model manager, one settings surface, and one set of commands.

Dictate an idea. Capture a meeting. Translate a passage. Listen to a note. Refine the result. It all happens inside the editor where your notes already live.

Choose the models that fit your workflow

Speech Kit is not tied to one speech engine or hosted API. It manages a growing catalog of models. Install only what you need, mix and match, and change models as your language, hardware, or priorities change.

You want Choose
Words on screen while you speak Moonshine streaming models
Multilingual live transcription Nemotron 3.5 ASR
The most accurate transcripts Whisper Large V3 Turbo, Cohere Transcribe, and other batch models
Natural local voices Pocket TTS or Supertonic 3
Fast offline translation Firefox Translations

The setup wizard installs the native engine and your first speech model. From there, Speech Kit manages the downloads and you choose how you work.

Dictate, transcribe, translate, listen, and refine

Dictate. Streaming words appear and revise in place while you speak. Finished text lands as Markdown at your cursor. Switch to a batch model when accuracy after each pause matters more than immediacy.

Transcribe. Combine your microphone with system audio to capture meetings, calls, interviews, and videos. Add timestamps and optional on-device speaker labels.

Translate. Translate a selection or a whole note between English and seven other languages. Preview the result before replacing your text, inserting it into the note, or copying it. One local model pack covers every supported direction.

Listen. Read any note aloud with natural local voices. Control the voice, speed, and playback without leaving Obsidian.

Refine. Optional LLM tools can clean up, summarize, restructure, or transform text with your own prompts.

One toolkit across platforms

Many speech apps are limited to one operating system, one model, or one part of the workflow. Speech Kit brings the same toolkit to macOS, Windows, and Linux, with hardware acceleration and system-audio capture where available.

Platform Architecture Acceleration System audio
macOS Apple silicon Metal for Whisper macOS 14.2 or later
Windows x86-64 Optional NVIDIA CUDA Supported
Linux x86-64 glibc Optional NVIDIA CUDA PulseAudio or PipeWire

Choose your platform. Choose your models. Keep one workflow inside Obsidian.

Getting started

  1. Install Speech Kit from Community Plugins.
  2. Follow the setup wizard to install the native engine and a speech model.
  3. Select Try dictation now, or start from the ribbon, command palette, or a hotkey.

Dictation, transcription, translation, and read aloud require no account, API key, usage credits, or cloud service. Once their models are installed, they continue working offline.

Optional LLM text tools are separate. You can connect a local or remote provider when you choose to use them.

Language support

Each feature is served by a different model, so coverage is tracked per feature rather than as a single list.

Language Transcription Live dictation Read aloud Translation Interface
English, Spanish, German, French, Portuguese, Italian, Dutch, Japanese ✅ ✅ ✅ ✅ ✅
Croatian ✅ ✅ ✅ — ✅
Serbian ✅ — — — —

✅ supported · — not yet available

Transcription coverage also depends on the model you select: multilingual models cover the full set above, while some smaller or specialized models are English-only. Translation runs through English in either direction, so every supported pair has English on one side.

Local-first, private by default

Speech Kit works without accounts, subscriptions, or required cloud services.

  • Your work stays on your machine. Dictation, transcription, read aloud, and translation run locally and continue working offline once their models are installed.
  • No account, telemetry, or metered usage. No API key, credit card, subscription, or usage credits to monitor.
  • LLM tools are optional. Add flexible language processing to your workflow using a local model or a remote provider you choose. Text leaves your device only when you explicitly use a remote provider, and audio is never uploaded.
  • Choose what works for you. Install high-quality models suited to your language, hardware, and workflow.
  • Transparent and open. Downloads are explicit, third-party licenses are documented, and Speech Kit is open source.

Support development

If Speech Kit is useful to you, please support development:

Development and project links

Speech Kit pairs a TypeScript plugin with a Rust native sidecar. See CONTRIBUTING.md for its architecture, setup, and development workflow.

Community Plugin · Latest release

Issues · License

Third-party component and model licenses are documented in THIRD_PARTY_NOTICES.md and shown before model download.

HealthExcellent
ReviewNot scanned
About
Obsidian handles the notes. Speech Kit handles speech and language. Speech Kit - formerly Local Dictation - brings everything you need to work with speech and language into one place. Dictate ideas as they happen. Turn meetings and calls into useful notes. Translate notes across eight languages. Listen to your writing with natural voices. Clean up, summarize, and reshape transcripts when you need to. Built for desktop Obsidian on macOS, Windows, and Linux, Speech Kit lets you choose the models that fit your language, hardware, and priorities. Core dictation, transcription, translation, and read aloud run on your device and keep working offline - no account, subscription, API key, or metered usage. Optional LLM tools can use a local model or a remote provider you choose.
AudioAIWriting
Details
Current version
2026.8.4
Last updated
15 hours ago
Created
4 months ago
Updates
27 releases
Downloads
2k
Compatible with
Obsidian 1.11.5+
Platforms
Desktop only
License
MIT
Report bugRequest featureReport plugin
Author
Alexander BrittainAlexander Brittainbrittain9
GitHubbrittain9
  1. Community
  2. Plugins
  3. Audio
  4. Speech Kit

Related plugins

Text Generator

Generate text content using GPT-3 (OpenAI).

Smart Composer

AI chat with note context, smart writing assistance, and one-click edits for your vault.

Local GPT

Local Ollama and OpenAI-like GPT's assistance for maximum privacy and offline access.

ChatGPT MD

A seamless integration of ChatGPT, OpenRouter.ai and local LLMs via Ollama into your notes.

BMO Chatbot

Generate and brainstorm ideas while creating your notes using Large Language Models (LLMs) such as OpenAI's "gpt-3.5-turbo" and "gpt-4".

GPT-3 Notes

Generate notes on any subject using OpenAI's GPT-3.5 and GPT-4 language models.

Claudian

Embeds Claude Code/Codex and other local Agents as AI collaborators in your vault.

Smart Connections

Find related notes and excerpts while writing. Your AI link building copilot displays relevant content in graph + list view. A local embedding model powers semantic search. Zero setup. No API key.

Fast Note Sync

Real-time sync of your vaults across server, mobile, and web; shareable with anyone; supports REST and MCP integrations to build your personal AI knowledge base.

Copilot

Run AI agents such as Claude Code, Codex, and OpenCode inside your vault. Turn your second brain into a smart assistant that gets knowledge work done.