Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

F

dsh-voice-announcer

flashyiyi/dsh-voice-announcer

Voice announcements when a conversation ends (session title, round, outcome) plus live sentence-by-sentence reading of replies, via built-in edge-tts with zero third-party dependencies.

03 days agoVision, Voice & Multimodal
B

dsh-soundscape

berserk0501/dsh-soundscape

Dual-MediaPlayer daemon ($think loop + $fx one-shots) with volume ducking, mood-aware sounds (smooth/struggle/deep_think triggered by consecutive success/failure streaks and prolonged agent runtime), 15 built-in PCM WAV synthesised from sine/square waveforms, per-tool sound mapping with 51 categorised tools, custom WAV/MP3 drop-in with auto-conversion, and a full settings UI with dark-mode fix.

0last monthVision, Voice & MultimodalMIT
K

deepseek-plugin

klingai-dev/deepseek-plugin

Turn every idea into an image or video with Kling AI. In DeepSeek Harness, use natural language for text-to-image, image-to-image, text-to-video, image-to-video, reference-image creation, task tracking, and result previews.

0last monthVision, Voice & Multimodal
K

dsh-auto-vision

k2d5rqjpkg-art/dsh-auto-vision

Auto-switch the DeepSeek route to the vision model on demand: flash main session switches (A), pro keeps deep reasoning and delegates image reading to a vision subagent (B), subagents always switch, with fatal-failure fallback. No manual model switching.

0last monthVision, Voice & Multimodal
C

dsh-evidence

cooberped/dsh-evidence

Turns attached files into versioned evidence: `search_documents` builds a private local index (SQLite FTS5 after a startup capability probe, dependency-free JS fallback otherwise) and returns compact evidence blocks carrying an exact coordinate — PDF page, PPTX slide, text/DOCX line range, or quoted XLSX `Sheet!Range` — which `read_document` expands only while the content version still matches. Contiguous CJK runs are indexed as overlapping bigrams and queried as phrases, so word order is preserved; uploaded raster images take the native vision attachment path instead.

0last monthVision, Voice & MultimodalMIT
Z

dsh-tu4-inline-images

zehenk/dsh-tu4-inline-images

对话内联图片 DSH 插件 — 在 DeepSeek Harness (DSH) Web GUI 的对话中,出现本地图片路径即直接渲染为图片。 A DSH plugin that renders local image paths as inline images in DeepSeek Harness (DSH) web conversations. Security-first: loopback-only route, strong per-process token, multi-root realpath whitelist.

0last monthVision, Voice & MultimodalMIT
B

dsh-llm-capabilities

bamboostrip/dsh-llm-capabilities

DSH plugin: auto-detect and configure model capabilities (reasoningEfforts + input modalities) for llm-pi-ai. Successor to dsh-reasoning-efforts.

0last monthVision, Voice & MultimodalMIT
L

dsh-vision-toggle

lijian-ui/dsh-vision-toggle

Per-model vision (image input) toggle for DeepSeek Harness (dsh): list every configured model and flip a switch to enable/disable image support without hand-editing settings.yaml. 为 DeepSeek Harness 提供按模型的「支持图片」开关:无需手改 settings.yaml。

0last monthVision, Voice & MultimodalMIT
A

dsh-speech

allmodels-io/dsh-speech

Streaming speech-to-text for DeepSeek Harness using AllModels.io.

0last monthVision, Voice & MultimodalMIT
C

dsh-agnes

chaoliu615/dsh-agnes

Agnes AI image & video generation tools for DSH (agnes_image_generate / agnes_video_generate)

0last monthVision, Voice & MultimodalMIT
D

dsh-remote-deliver

demacia1314/dsh-remote-deliver

🚀 告别繁琐 SCP!远程部署 DSH 一键下载修改后的文件与图片预览交付插件

0last monthVision, Voice & Multimodal
W

dsh-dictation

wsl043/dsh-dictation

Editable local and desktop dictation for DeepSeek Harness

0last monthVision, Voice & MultimodalMIT
L

dsh-file-attach

lucasxingg/dsh-file-attach

Drag-and-drop PDF, Office, images, and text/code files into DSH conversations. The host extracts (and OCRs) them into the prompt; attach_* tools cover notebook cells, PDF-page OCR, image describe, and save.

0last monthVision, Voice & MultimodalMIT
J

dsh-csv-and-image-preview

jetecho/dsh-csv-and-image-preview

Preview images / SVG (and prepare CSV) in the DeepSeek Harness chat, rendered as real browser <img> elements. Preview-first workflow: show the user the asset, wait for approval, then apply the real change.

0last monthVision, Voice & MultimodalMIT
X

dsh-voice-input-space

xsakura666/dsh-voice-input-space

Voice input for DeepSeek Harness: hold Space to speak, release to insert. Zero dependencies, Web Speech API. / 语音输入:长按空格说话,松开上屏,零依赖。

0last monthVision, Voice & MultimodalMIT
N

dsh-voice

navid-kianfar/dsh-voice

Dictate prompts into the DeepSeek Harness Web Client — a microphone in the composer, with swappable transcription: hosted Whisper API, self-hosted server, or a fully offline whisper.cpp binary.

0last monthVision, Voice & MultimodalMIT
Z

dsh-video-understand

zeshuochen/dsh-video-understand

Subtitle-first video transcription and deterministic extractive Markdown summaries, with a faster-whisper large-v3 fallback when subtitles are unavailable.

0last monthVision, Voice & Multimodal
S

dsh-auto-vision

soarguo/dsh-auto-vision

Bridges images into text for DeepSeek Harness: when the session's selected model cannot see images, a configured vision model describes them and the descriptions enter the durable session history as folded context rows — your message stays untouched.

0last monthVision, Voice & MultimodalMIT
C

dsh-mmx

crazyma99/dsh-mmx

MiniMax CLI bridge for DeepSeek Harness (dsh): mmx-backed web search fallback and transparent image understanding, with first-run onboarding and a settings card.

0last monthVision, Voice & MultimodalMIT
Q

dsh-voice-mode

qishuilalala/dsh-voice-mode

Full-duplex voice mode for DeepSeek Harness: zipformer2 streaming ASR → editable draft, Edge TTS sentence-by-sentence read-aloud with live captions, true barge-in — on-device ASR, no API key. · DSH 语音双工对话:流式识别入草稿、按句朗读+实时字幕、开口即打断,识别本地推理、无需 API Key

0last monthVision, Voice & MultimodalMIT
Y

dsh-mimo-plugin

yyfather/dsh-mimo-plugin

MiMo (Xiaomi) tools as a DSH profile plugin: web search, image/audio/video understanding, ASR transcription, TTS, voice design and voice cloning — native agent tools with a Settings page for the API key.

0last monthVision, Voice & MultimodalMIT
B

dsh-plugin-88api-image

blackdm666/dsh-plugin-88api-image

88API Image Studio for DSH: four Image2 and Nano Banana models for text-to-image, multi-reference editing, 2K/4K output, and sequential batches.

0last monthVision, Voice & MultimodalMIT
X

computer-use-vision

xuanyuanluoxue/computer-use-vision

Windows computer-use capability for DeepSeek Harness: screenshot → vision model → simulated mouse/keyboard input, with self-evolving knowledge base.

0last monthVision, Voice & Multimodal
F

dsh-glm-vision

fightingfirefox/dsh-glm-vision

GLM 视觉模型插件:注册 glm-vision 供应商路由(glm-4.6v 系列,image+text 输入声明)并提供 glm_vision 工具,让 DeepSeek 等文本主模型直接调用智谱视觉模型看图。

0last monthVision, Voice & MultimodalMIT