Karpathy's LLM Wiki implementation plugin for Obsidian - turns notes and PDFs into a linked, LLM-powered knowledge base with entity pages, concept pages, graph-powered Q&A, and local-first privacy.

An Obsidian plugin that turns your notes into a connected, queryable knowledge base — the Karpathy LLM Wiki idea, built into the editor where you already write.
Obsidian Review Perfect Score • Zero-embedding graph retrieval • 11-language native • Native PDF + images + Office ingest • Works with every provider • Local-first • No backend • GDPR-Friendly
<br>
English | 简体中文 | 繁體中文 | 日本語 | 한국어 | Deutsch | Français | Español | Português | Italiano | Русский
Official Site | Obsidian Marketplace | Blog | Discussions
🤔 Why this plugin? | 🚀 Quick Start | ✨ Features | 🌐 Ecosystem | 🛠️ Headless CLI | 🔍 How Retrieval Works | 🤖 Models | ❓ FAQ
← If this plugin has helped you, feel free to buy me a coffee♥️ or drop a star🌟↗
You write notes. They sit in folders. Finding what relates to what means remembering threads you forgot months ago.
Other open-source reimplementations of Karpathy's LLM Wiki idea exist — but none ships as a one-click Obsidian plugin. Most are CLI tools, Claude Code skills, or separate desktop apps; this one runs inside Obsidian — Graph View, ribbons, command palette included.
| Karpathy LLM Wiki (this plugin) | nashsu / llm_wiki | SamurAIGPT / llm-wiki-agent | atomicstrata / llm-wiki-compiler | |
|---|---|---|---|---|
| Delivery | ✅ One-click Obsidian plugin | 🟡 Tauri desktop app | 🟡 Claude Code / Codex / OpenCode / Gemini CLI skill | 🟡 TypeScript CLI pipeline |
| Dependencies | ✅ None — plugin only | 🟡 Python runtime + sqlite | 🟡 Claude Code / Codex / OpenCode runtime | ❌ Embedding model + vector DB (per their docs) |
| i18n (UI + wiki output) | ✅ 11 languages | 🟡 EN / 中文 | ❌ EN only | ❌ EN only |
| LLM providers | ✅ 16+ (Anthropic, OpenAI, Bedrock, Gemini, DeepSeek, Qwen, Grok, Kimi, GLM, MiniMax, Step, Hunyuan, MiMo, Gemma, Codex OAuth, Ollama, LM Studio, OpenRouter, Anthropic-Compatible, …) | 🟡 OpenAI-compatible | 🟡 Subscription via Claude Code / Codex | 🟡 OpenAI-compatible |
| Retrieval & query pipeline | ✅ PPR + Monte Carlo over [[wiki-link]] graph |
🟡 2-hop decay (4-signal heuristic) | 🟡 Louvain community detection | 🟡 BM25 + semantic over chunks |
| Graph visualization | ✅ Obsidian's native Graph View (built in, zero extra size) | 🟡 Custom sigma.js + graphology in desktop app | 🟡 vis.js graph.html (separate file) |
❌ Read-only browser viewer |
| Ingest formats | ✅ Markdown + PDF + images + Office (DOCX/PPTX/XLSX) — one switch flips between native PDF (Anthropic / OpenAI / Bedrock / Gemini) and the built-in MinerU multi-format backend | 🟡 Markdown + PDF | 🟡 Markdown / code files only | 🟡 Markdown only |
Three we left off the table: sdyckjq/llm-wiki-skill was a Codex skill that 404s today (deleted by author); atomicstrata is included even though its retrieval is chunk-based because it's the most active TypeScript alternative; nashsu ships the largest user base of the three (10k+ stars, others 2k+) but is a Tauri desktop app, not an Obsidian plugin.
[[wiki-link]] graph — built in, zero extra bundle size.✅ Yes, if you:
wiki/ within seconds.[[wiki-links]] back into your knowledge graph.[[wiki-links]] — every link you write already enriches retrieval; no separate tagging/embedding/indexing step.❌ No, if you:
Install. Obsidian → Settings → Community plugins → Browse → search "Karpathy LLM Wiki" → Install → Enable. Or visit the Community Plugin page and click Add to Obsidian.
Configure a provider. Open Settings → Karpathy LLM Wiki → pick a provider (OpenAI, Anthropic, Ollama, ChatGPT Plan (Codex OAuth), etc.) → enter API key (not needed for local) → click Test Connection → Save.
Ingest one note. Two ways:
Cmd+P/Ctrl+P → "Ingest single source" → pick any Markdown (or PDF, v1.25.0+) file.Your first wiki pages appear in wiki/sources/, wiki/entities/, wiki/concepts/ within seconds.
Query your wiki. Two ways:
Cmd+P/Ctrl+P → "Query wiki".A right-docked side panel opens (Copilot-style) where you can chat with your wiki. Answers carry [[wiki-links]] back into your knowledge graph.

That's it. The plugin modifies nothing in your original notes — only creates new pages under wiki/. Both Ingest and Query wiki are pinned to the left ribbon for one-click access anytime. (Cmd on macOS, Ctrl on Windows/Linux.)
| Command | What it does |
|---|---|
| 📥 Ingest single source | Cmd+P/Ctrl+P → "Ingest single source" — pick a Markdown or PDF (v1.25.0+) file, get entity/concept/wiki pages. Also: 🖱️ ribbon sticker icon on the active note. |
| 📂 Ingest from folder | Cmd+P/Ctrl+P → "Ingest from folder" — batch-ingest every note in a folder, with smart batch skip |
| 📑 Ingest multiple files | Cmd+P/Ctrl+P → "Ingest multiple files" — pick a subset via a two-pane file tree (with live queue + per-file cancel) |
| 🔍 Query wiki | Cmd+P/Ctrl+P → "Query wiki" — chat with your wiki in a right-docked side panel; answers carry [[wiki-links]]. Also: 🖱️ ribbon message-circle icon. |
| 🛠️ Lint wiki | Cmd+P/Ctrl+P → "Lint wiki" — full health scan: duplicates, dead links, empty pages, orphans, missing aliases, contradictions |
| ⚡ Smart Fix All | inside Lint Modal — one-click causal-order repair with per-phase report |
| 📋 Regenerate index | Cmd+P/Ctrl+P → "Regenerate index" — rebuild wiki/index.md with current pages and aliases |
| ⏹ Cancel | Cmd+P/Ctrl+P → "Cancel current ingestion" or click the status bar — stops cleanly at the next batch boundary |
| 📊 Ingestion history | Cmd+P/Ctrl+P → "View Ingestion History" — searchable UI for past ingestions, lint reports, maintenance runs |

| Before | After |
|---|---|
notes/machine-learning.md (a flat file) |
wiki/concepts/supervised-learning.md with [[bidirectional links]], aliases, source attribution, and an entry in wiki/index.md |
📖 Walkthroughs in GitHub Discussions → Guides. Found it useful? Star the repo to follow releases.
reviewed: true pages are protected from overwrite.Five on-ramps, switchable per ingest:
.md, drop it in your vault outside the wiki folder, and ingest as a regular Markdown note.![[image.png]] and  embed. Each image is analyzed with its nearest Markdown paragraphs, sent in 20 MiB visual-evidence packages, and capped at 10 MiB; remote URLs are never downloaded. For testable per-image output, optionally enable Save embedded image visual evidence to source page to add a collapsible audit section to the generated source page.Caveat for Office formats: Obsidian does not natively render .docx / .xlsx / .pptx (file-formats), so the practical workflow for Office files is: MinerU converts to .md, the plugin ingests that .md into wiki pages, and the original Office file is kept around for reference. Use a community plugin like Pandoc Plugin / Docxer / Md Importer / Office Reader if you need to inline-preview Office files.
Plumbing shared across all paths:
.obsidian/plugins/karpathywiki/pdf-cache/ stores converted Markdown keyed by content hash + model + converter version; 100 MB total / 1000 entries / 10 MB single-entry caps with LRU-by-mtime eviction.<basename>.pdf.md next to the source PDF (off by default — cache-only is the default).[illegible] / [figure: ...] anti-hallucination markers; markdown-fence-wrapping from small local models is auto-cleaned before cache write.sources/<slug>.md page now carries a Mentions in Source section built from the same verbatim quotes the extraction captured per entity/concept (the prose the model already proved it could see), so the underlying document is the one wiki page with a real, grounded trail back to its source text.📖 Full setup walkthroughs for all paths (cloud providers, oMLX hardware tiers, MinerU installation, cache housekeeping) → docs/PDF-OCR-GUIDE.md
[[wiki-link]] gives graph-aware multi-hop context.skipMentionOnlyCandidates, default off, Settings → Advanced). For sources whose language has a measured profile (de measured; en/fr/es/pt/nl/ko estimated with pinned edge cases; zh/ja character-script thresholds unmeasured), candidates named only inside parentheses / enumerations / short list items are pruned before they cost a page plus dedup and generation calls. Cross-language notes are not gated; wiki languages without a profile report once per ingest and never silently skip.wiki/.The plugin composes with the rest of your Obsidian stack — each tool below plugs into the [[wiki-link]] graph without code changes.
[[wiki-link]] becomes a node, every back-link an edge. Built in, zero extra bundle size.Ingest from folder command to batch-extract entities and concepts.LIST FROM "wiki/entities" WHERE contains(tags, "person")) or JS API. The plugin writes standard frontmatter (tags:, type:, aliases:) on every page, so Dataview queries work out of the box.wiki/, so git diff cleanly separates your edits from LLM-generated content.marp: true). Wiki pages are pure Markdown, so they render as slides without extra conversion.[[wiki-links]] without leaving the vault.Most users should ignore this section. The plugin's user-facing CLI lives in the sibling repo green-dalii/obsidian-llm-wiki-cli — install with npm i -g karpathywiki-cli and run karpathywiki-cli ingest --sources <path> --wiki <path> --provider <id> --key <key>.
What ships in this repo at tools/dev-instrument/ is the dev-only headless measurement instrument for engine contributors — it runs the real WikiEngine.ingestSource against a vault on disk with no Obsidian runtime, prints per-task token + wall-clock accounting — same numbers that drive the perf evidence in CLAUDE.md and release notes. See tools/dev-instrument/README.md for the entry command, env vars, measurement modes, and exit-code spec.
Most "AI search" plugins fragment your notes into chunks and embed them in a vector DB. We don't. Karpathy's argument against RAG is that chunking breaks the LLM's ability to reason across your whole knowledge graph — and that argument holds up in practice. Instead, we walk the graph you already maintain by writing [[wiki-links]].
When you ask "Who founded Microsoft?", Query Wiki runs five stages before any answer generation:
[[wiki-link]] graph starting from the candidate seed set. This is what gives graph-aware multi-hop context: "Bill Gates" → "Microsoft" → "competitors", not just literal title overlap.The cascade truncates at whichever stage returned enough signal — no fixed 5-stage cost; no LLM calls when lex is sufficient; semantic fallback only when lex + keyword scan alone isn't enough.
We use Monte Carlo PPR (Fogaras 2005) — 3,000 random walks × 50 steps each — with the dead-end rule from Haveliwala 2002. Cost is O(K × L) (K = walks, L = steps per walk), independent of the number of pages, so a 2,000-page vault sees the same expansion latency as a 200-page one.
PPR @5 = 27.1% vs pure-kNN baseline 24.1% on the project's own benchmark corpus (the only published retrieval benchmark in this open-source LLM-Wiki space).
We deliberately rejected the embedding path in Issue #175. The graph signal is already there — every [[wiki-link]] is a hand-curated "these are related" edge, and most providers we support (Ollama, LM Studio, Anthropic, Bedrock, Kimi, GLM, MiniMax) don't ship a /v1/embeddings endpoint at all. Adding an embedding model would mean a per-page download, a per-provider adapter, and zero benefit on retrieval quality.
Supported providers (16+, all from models.dev cross-check 2026-07):
| Provider | Series | Notes |
|---|---|---|
| Anthropic | Claude 5 series | Native PDF; /v1/messages protocol |
| OpenAI | GPT-5.6 series (Sol / Terra / Luna) | Native PDF; Platform API key |
| Google Gemini | Gemini 3.6 series | Native PDF (file parts since 1.5); OpenAI-compatible endpoint |
| DeepSeek | DeepSeek V4 series | OpenAI-compatible; lowest cost tier |
| Alibaba Qwen | Qwen3.7/3.8 series | OpenAI-compatible (DashScope) |
| xAI Grok | Grok 4 series | OpenAI-compatible; long context |
| Moonshot Kimi | Kimi K3 series | OpenAI-compatible; 2.8T MoE frontier |
| Zhipu GLM | GLM-5 series | OpenAI-compatible; strong bilingual |
| MiniMax | MiniMax M3 series | OpenAI-compatible; 1M context |
| Step (阶跃星辰) | Step 3 series (Flash) | OpenAI-compatible; fast inference |
| Tencent Hunyuan | Hy3 series | OpenAI-compatible; open-weight MoE |
| Xiaomi MiMo | MiMo V2.5 series | MIT open-source; flat pricing |
| Google Gemma | Gemma 4 series | Open-weight; 262K context |
| AWS Bedrock | Anthropic + OpenAI variants | Native PDF; VPC / compliance path; API key + SSO + IAM (v1.27.0, #425) |
| ChatGPT Plan (Codex OAuth) | Codex Responses API | Browser/device-code sign-in; SecretStorage |
| Local: Ollama, LM Studio, OpenRouter, Anthropic-Compatible | Any OpenAI-/Anthropic-protocol model | Custom OpenAI-Compatible + Anthropic-Compatible (Token Plan / Coding Plan) |
This plugin feeds the LLM your full Wiki context per query — so long-context models win. The full tiered table (cloud + local) lives in docs/MODEL-GUIDE.md, cross-checked against models.dev so the picks stay current.
/v1/embeddings is fine (most of our 16+ providers don't ship one).For PDF / image / Office ingest, see Document / PDF / Image ingest in Features — Anthropic, OpenAI, Bedrock, and Gemini read PDFs as file parts natively; the built-in MinerU backend (v1.27.0+) and Force PDF Support cover everything else.
Settings → Provider → Bedrock (Anthropic / OpenAI) now picks one of three auth modes; the provider row then asks for the inputs that mode actually needs:
karpathywiki-bedrock-sso in SecretStorage, exchanges it for temporary role credentials, and signs every request with hand-rolled SigV4 (no AWS SDK added). Account ID and role name are auto-detected when the SSO identity exposes exactly one of each; otherwise enter them in the provider settings.karpathywiki-bedrock-iam in SecretStorage; the in-memory cache memoizes per access-key to keep SigV4 signing within expiry.All three modes share the same Obsidian SecretStorage discipline (no credentials in data.json, logs, or docs) and the same zero-AWS-SDK hand-rolled OIDC + SigV4 path. Bedrock region is independent of auth mode and is configured in the same provider row.
📖 Full pick table (cloud + local + PDF OCR + Codex OAuth + quantization + hardware tiers) → docs/MODEL-GUIDE.md
Pick any note, folder, or selection; the LLM extracts entities and concepts and generates an interlinked wiki with [[bidirectional links]]. Ask questions and get conversational answers grounded in your notes, not the internet. Your original vault notes are never modified.
Install from Obsidian Community Plugins → pick a provider → Test Connection → run Ingest single source on any note. First wiki pages appear within seconds. See Quick Start.
✅ Backward compatible since v1.0.0. Set reviewed: true on any page to protect it from overwrite. Upgrading from v1.24.x doesn't rewrite your vault; v1.25.0's PDF ingest is cache-only by default, and v1.27.0 adds native PDF + images + Office ingest without changing the on-disk wiki layout.
✅ Yes. Anthropic, OpenAI, Bedrock, and Gemini read PDFs natively; the built-in MinerU backend (v1.27.0) covers everything else (PDF + images + Office). Full walkthrough — cloud providers, Apple Silicon OCR, Force PDF Support, cache housekeeping — in docs/PDF-OCR-GUIDE.md.
🚫 No backend, no analytics — the plugin runs entirely inside Obsidian. Only text you explicitly send for ingest/query leaves your device, and only to the LLM provider you configure. For complete data locality, use Ollama or LM Studio.
🌍 11 languages for both UI and wiki output. UI and wiki language are independent. Adding a 12th language is contributor-driven (PR #159 pattern).
🚫 No chunking. 🚫 No embeddings. 🚫 No vector DB. ✅ Personalized PageRank over your existing [[wiki-link]] graph — graph-aware multi-hop context, zero embedding cost, full local-model support.
Long-context models (≥200K tokens) work best. The Models section covers the principles; the full tiered table is in docs/MODEL-GUIDE.md.
Yes — PPR @5 = 27.1% vs pure-kNN baseline 24.1% on the project's own corpus. The full pipeline and benchmark script are described in How retrieval works.
Use Coarse or Minimal extraction granularity for batch ingest. Smart Batch Skip auto-detects already-ingested files. Auto-Maintenance is OFF by default. Lint shows counts before running fixes — nothing is charged without your approval.
Click the status bar (shows "Ingesting… click to cancel") or Cmd+P/Ctrl+P → "Cancel current ingestion". Stops cleanly at the next batch boundary.
GitHub Issues for bug reports · GitHub Discussions for questions and feature requests · Developer Console (Ctrl+Shift+I / Cmd+Option+I) for plugin logs.
This plugin is listed on the Obsidian Community Plugin Market and undergoes automated review for security and permissions.
mineru.net. That second one is the exception worth naming: it uploads the document to a jurisdiction you did not choose and that publishes no retention statement. Where your document is processed lists every path and where the file goes.For complete data locality, use Ollama or LM Studio. With a local provider, your data never leaves your machine.
If LLM-Wiki has become a meaningful part of your knowledge workflow:
Thanks to the following for supporting the project:
@jameses-cyber, @issaqua, Dikson Choi
karpathywiki-cli npm package. Runs the same WikiEngine against a vault on disk, no renderer. Install with npm i -g karpathywiki-cli. The in-tree tools/dev-instrument/ is the dev-only measurement instrument that drives the per-task cost numbers in this plugin's release notes.obsidian-plugin-dev workflow: scaffolds an Obsidian plugin workspace, drives the Red→Green TDD loop, runs the Six-Gate quality closure (lint/tsc/test/build/css-lint), and prepares a release-ready branch on feat/* or fix/*. Built so DSH-using contributors get the same scaffolding + gate experience without copy-pasting from CLAUDE.md.Apache License, Version 2.0 — see LICENSE, NOTICE and THIRD-PARTY-NOTICES.md.
Built on:
@ai-sdk/openai, @ai-sdk/anthropic, @ai-sdk/openai-compatible) via Obsidian requestUrlMaintainers: @green-dalii (author) · @DocTpoint (co-maintainer since September 2026)