Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

F

dsh-vision-proxy

flyvhidbwo/dsh-vision-proxy

DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed via the official deepseek-v4-flash-vision-exp by default (a pure-text V4-Pro brain can see images), with any OpenAI-compatible VLM or local Ollama as alternatives.

15last monthVision, Voice & MultimodalMIT
D

dsh-video-lens

dundunhan/dsh-video-lens

Give text-only DeepSeek Harness agents video understanding: scene-aware frame sampling + VLM + optional ASR transcript fused into timeline evidence. / 给纯文本模型的视频理解插件(场景感知抽帧 + VLM + 可选语音转录)

13last monthVision, Voice & MultimodalMIT
P

dsh-vision-opencode

poiuyjie/dsh-vision-opencode

Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.

1324 days agoVision, Voice & MultimodalMIT
Z

dsh-agnes-studio

zmm863-commits/dsh-agnes-studio

AI image and video studio as a floating DSH panel, so creation runs alongside the conversation instead of replacing it: text-to-image, image-to-image and multi-image composition at 1K-4K across eight aspect ratios; text-to-video and image-to-video with first-frame control at 4-12 seconds; short-drama mode that imports a script (.txt/.md/.json), breaks it into storyboard shots for preview and batch generation; and a prompt-expert workspace. Zero runtime dependencies.

123 days agoVision, Voice & Multimodal
B

dsh-voice-ai-girlfriend-plugin

beiyege-01/dsh-voice-ai-girlfriend-plugin

Voice AI girlfriend for the Web UI: FunASR mic input, Qwen3-TTS spoken replies, companion animation window, and two-way QQ chat (text/voice/image push) via NapCat.

1220 days agoVision, Voice & Multimodal
P

dsh-plugin-xiaomi-mimo-tts

ppy-web/dsh-plugin-xiaomi-mimo-tts

Adds Xiaomi MiMo text-to-speech to DSH Web with assistant-message read-aloud, PCM streaming, preset and custom voice design, browser speech fallback, playback controls, and optional UI sounds.

113 days agoVision, Voice & MultimodalMIT
A

dsh-speak

alan2z/dsh-speak

Zero-dependency, event-driven voice announcement plugin: no extra model, no token cost. Speaks with the system's built-in natural voice, supporting both Windows and macOS; final-reply announcements, approval & question alerts, optional event announcements (turn end, command done, goal change, tool errors, todo updates), replayable final replies, and a bilingual visual settings page.

117 days agoVision, Voice & MultimodalMIT
C

Gemini-Eyes

consolesun/gemini-eyes

MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.

11last monthVision, Voice & Multimodal
P

dsh-draw

perrylink/dsh-draw

Multi-engine text-to-image generation (OpenAI Images and Zhipu CogView presets) with per-session quota tracking, engine failover, credential-safe config, and a result card with regenerate.

98 days agoVision, Voice & MultimodalApache-2.0
S

dsh-deepseek-vision

siegfly/dsh-deepseek-vision

A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.

922 days agoVision, Voice & MultimodalMIT
L

dsh-vision

linenxi-ctrl/dsh-vision

External vision plugin for DeepSeek Harness: whale-button config panel, image recognition with auto-reply, and agent screenshot/recognize tools.

92 months agoVision, Voice & MultimodalMIT
A

dsh-highres-vision

azwosile/dsh-highres-vision

For the DeepSeek Harness native vision model deepseek-v4-flash-vision-exp, raises image admission limits to 32 MiB / 8192 px / 600 images and adds a highres_read tool that tiles large images, then returns the whole image plus 800x800 tiles through the host read_image tool.

822 days agoVision, Voice & MultimodalMIT
G

dsh-voice

goodandready/dsh-voice

Voice input for the web UI: dictation chunked by pauses and voice messages, each with its own provider fallback chain (Deepgram, Groq, HuggingFace, local whisper.cpp, or an OpenAI-compatible endpoint).

8yesterdayVision, Voice & MultimodalMIT
3

dsh-voice

3274375092/dsh-voice

Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.

82 days agoVision, Voice & MultimodalMIT
F

dsh-free-vision

fuzzysoul/dsh-free-vision

Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.

8last monthVision, Voice & MultimodalMIT
C

dsh-chat-imagine

corrinehu/dsh-chat-imagine

Automatically generates and displays images in the DSH chat via API channels or local CLIs (mmx / codex / agy), and can also recognize images using the corresponding CLI.

8last monthVision, Voice & MultimodalMIT
S

dsh-tool-vision

scorp1o117/dsh-tool-vision

Send local images or HTTP(S) image URLs to a configured OpenAI-compatible vision endpoint with the inspect_image tool and return its text response to the conversation.

83 days agoVision, Voice & Multimodal
W

dsh-appshots

wongyuye/dsh-appshots

Codex-style window capture for DSH Desktop on macOS and Windows: press both Command keys (macOS) or both Ctrl keys (Windows), or the camera button, to grab the frontmost window, attach it to the current chat, and inject a VoiceOver-style accessibility tree as hidden context.

715 days agoVision, Voice & MultimodalMIT
5

dsh-vision

54xkeee/dsh-vision

Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

72 months agoVision, Voice & MultimodalMIT
Y

dsh-duet

yums/dsh-duet

Full-duplex Chinese voice interaction for DSH Web: dictate and edit tasks, submit requests, manage sessions, answer DSH questions, and hear concise task-completion announcements.

64 days agoVision, Voice & MultimodalApache-2.0
Y

deepseek-harness-plugins (vision-bridge)

yinxe/deepseek-harness-plugins

Let text-only models see images: when a picture arrives with a placeholder like \[image omitted because this model accepts text only], the model calls the vision_describe tool and the plugin forwards the image reference plus the question to a multimodal model, retrying with a fallback model on failure. Configured in the settings page and persisted via the official settings API.

62 days agoVision, Voice & MultimodalMIT
Z

dsh-voice-input-plugin

zhangbo-cn/dsh-voice-input-plugin

Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.

6last monthVision, Voice & MultimodalMIT
G

dsh-ocr-local

grelvan/dsh-ocr-local

Local OCR fallback for text-only routes: when the session model declares it cannot accept images, the attached image is cached locally and its path injected so the model can call ocr_image — PP-OCRv5 + ONNX Runtime on CPU, no API key, images never leave the machine. Silent when the model can see images.

55 days agoVision, Voice & Multimodal
N

vision-exp-tile

nicholas023/vision-exp-tile

Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.

5last monthVision, Voice & MultimodalMIT