Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

G

dsh-tts

goodandready/dsh-tts

Speaks agent replies in the DeepSeek Harness web UI through a provider fallback chain (OpenAI, ElevenLabs, Google, Azure, Groq, Deepgram, OpenRouter, Edge, Piper, eSpeak), so a failing or rate-limited provider falls through to the next instead of going silent.

26 days agoVision, Voice & MultimodalMIT
O

DSH-dseyes

okkay712/dsh-dseyes

Native image attachments for text-only DeepSeek in the Web GUI: pasted or dropped images appear as thumbnails in the session, and before dispatch the host reads them with the free Zhipu GLM-4V-Flash vision API (glm-4v-flash fallback chain) and substitutes the description, so DeepSeek answers about the image while the original is kept in history.

2last monthVision, Voice & MultimodalMIT
W

dsh-wm

waynejin0918/dsh-wm

World-model research toolkit: inspect frames, name 3D / pixel / latent routes, score pred vs GT, and RSI skills / wm.yaml.

22 months agoVision, Voice & MultimodalMIT
I

dsh-tool-visual-primitives

inkshadewoods/dsh-tool-visual-primitives

Analyzes conversation images through mode-specific prompts (caption, UI, document, grounding, topology, etc.) and injects structured evidence with coordinate primitives (boxes, points, refs) as text, with session-level caching for reuse across replay and compaction.

23 days agoVision, Voice & MultimodalMIT
V

dsh-smart-input

v-quest123456/dsh-smart-input

智能输入插件 for DeepSeek Harness — 语音输入 + 提示词优化

22 months agoVision, Voice & MultimodalMIT
G

dsh-vision-bridge

goodandready/dsh-vision-bridge

Routes images to a vision model of your choice - auto-rewrite, explicit tools, or hybrid - so a text-only chat model does not fail a turn that contains a picture.

26 days agoVision, Voice & MultimodalMIT
P

muxiva-dsh-voice

piyotahu/muxiva-dsh-voice

Local-first, full-duplex voice for DeepSeek Harness, orchestrated by Muxiva

22 months agoVision, Voice & MultimodalApache-2.0
P

dsh-voice-call

pandapolo/dsh-voice-call

Agent-initiated voice calls: `offer_call` rings the human (接听/拒接/稍后再说); accepted calls synthesize and play locally via CrispASR + Qwen3-TTS (9 speakers, 2 Chinese dialects), rejected calls return the decision to the agent.

23 days agoVision, Voice & MultimodalMIT
M

dsh-fish-tts

mari23333/dsh-fish-tts

Reads assistant replies aloud via Fish Audio API only (bring your own key): per-message read-aloud, auto-read toggle, and a settings page for model, voice reference_id, encrypted API key, and proxy.

22 days agoVision, Voice & MultimodalMIT
X

dsh-vision

xiaoshihou514/dsh-vision

Native vision capability extension, using either Zhipu (free) or Qwen-VL (local).

2last monthVision, Voice & MultimodalMIT
K

dsh-vision-recognizer

kaixinbaba/dsh-vision-recognizer

Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.

2last monthVision, Voice & MultimodalMIT
H

dsh-open-eyes

hyp6666/dsh-open-eyes

Vision bridge for text-only DeepSeek routes that analyzes attached and local images through configurable OpenAI Responses, Chat Completions, or Anthropic Messages endpoints while leaving image-capable routes native.

227 days agoVision, Voice & MultimodalMIT
5

dsh-youreyes

54xkeee/dsh-youreyes

Vision toolkit for text-only DeepSeek: model-invokable `vision` tool, wrapper adapters for deepseek/opencode-go (v4 flash/pro), Antigravity IDE quota (default, flash/pro) / any OpenAI-compatible VLM / Gemini / local Ollama channels, evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

22 months agoVision, Voice & MultimodalMIT
3

dsh-vision (vision-tool)

314857493/dsh-vision

Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.

228 days agoVision, Voice & MultimodalMIT
3

dsh-vision (vision-route)

314857493/dsh-vision

Registers a `deepseek-vision` provider route: the Web GUI accepts pasted images and transcribes them to text via the free Zhipu GLM vision API before delegating to the DeepSeek adapter.

228 days agoVision, Voice & MultimodalMIT
X

dsh-vision-bridge

ximengxiaolan/dsh-vision-bridge

Composer-attached images are transcribed to text by an OpenAI-compatible vision model before reaching text-only DeepSeek models.

22 months agoVision, Voice & MultimodalMIT
J

dsh-live2d-voice

john-walks-slow/dsh-live2d-voice

Live2D session view with realtime voice: continuous voice input, streaming TTS, subtitle translation, multi-model catalog, standalone entry / Live2D 实时语音会话视图:连续语音输入、流式 TTS、字幕翻译、多模型目录、独立入口

14 days agoVision, Voice & MultimodalMIT
W

dsh-multi-tts

wyr-233/dsh-multi-tts

Per-reply read-aloud with a multi-provider settings page — MiniMax or any OpenAI-compatible /audio/speech endpoint, with voice, emotion, speed and model selection plus an auto-read toggle.

122 days agoVision, Voice & Multimodal
V

dsh-live-voice

victorwads/dsh-live-voice

Local-first voice conversations for DSH, with local speech recognition and synthesis and optional external providers.

118 days agoVision, Voice & MultimodalGPL-3.0
C

dsh-vision-autoswitch

cultofluna/dsh-vision-autoswitch

DeepSeek Harness plugin: auto-route image-bearing requests to deepseek-v4-flash-vision-exp, then fall back to the original model.

1last monthVision, Voice & MultimodalMIT
N

dsh-dictate

navneetset/dsh-dictate

Live speech-to-text dictation into the dsh web composer via OpenRouter STT

127 days agoVision, Voice & MultimodalMIT
M

dsh-minimax-asr

moluyao/dsh-minimax-asr

MiniMax 语音识别(asr-1.0)+ 语音合成(speech-2.8-hd)的 DeepSeek Harness 全局插件:转写工具、任务结束后用喇叭念一句的播报、实时免手对话、303 个音色可选

119 days agoVision, Voice & MultimodalMIT
N

dsh-imgdraw

ninjasln-labs/dsh-imgdraw

Text-to-image for DeepSeek Harness: a `draw_image` model tool, an input-bar 生图 button with a prompt popup (async generation, 4-grid results, download / keep / delete), an /imgdraw image route, and persisted history. Backends: DashScope wan2.7-image (free

128 days agoVision, Voice & MultimodalMIT
D

dsh-voice-talk

duoduoqian708/dsh-voice-talk

Voice conversation mode for the DSH Web GUI: tap the mic to start a full-screen call, talk hands-free, and hear each reply read aloud as it streams. Bilingual (Chinese/English) UI, light and dark themes, adjustable speech rate, and switchable voices.

120 days agoVision, Voice & MultimodalMIT