Transcribe embedded images and PDFs in notes to editable Markdown using an OpenAI-compatible vision model that runs locally and fully offline. Stream live transcripts into a sidebar, create one transcript note per image without overwriting originals, and show model reasoning in expandable thinking blocks.