Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

86 plugins found

L

modlens

liustack/modlens

Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).

3.1k2 hours agoVision, Voice & MultimodalMIT
Y

dsh-vision-router

ysr666/dsh-vision-router

Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.

7401 hour agoVision, Voice & MultimodalMIT
A

dsh-vision-toolkit

anionex/dsh-vision-toolkit

Vision tasks for text-only models: intent-aware image Q&A, long-screenshot OCR, UI reproduction, grounding, and pixel diff.

6945 hours agoVision, Voice & MultimodalMIT
J

picturereader

jing-hy/picturereader

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

206 hours agoVision, Voice & MultimodalMIT
L

dsh-vision

linenxi-ctrl/dsh-vision

External vision plugin for DeepSeek Harness: whale-button config panel, image recognition with auto-reply, and agent screenshot/recognize tools.

123 days agoVision, Voice & MultimodalMIT
M

dsh-media-skills

mjorgin/dsh-media-skills

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

11yesterdayVision, Voice & MultimodalMIT
F

dsh-vision-proxy

flyvhidbwo/dsh-vision-proxy

DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed to text via any OpenAI-compatible VLM before reaching the text-only DeepSeek — a keyed fast path (default qwen3.7-flash; DashScope/Zhipu/OpenRouter or any OpenAI-compatible endpoint) with your own key, or local Ollama auto-detected with zero config.

112 days agoVision, Voice & MultimodalMIT
P

dsh-vision-opencode

poiuyjie/dsh-vision-opencode

Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.

1010 hours agoVision, Voice & MultimodalMIT
J

dsh-visual-plugin

jyh20030112/dsh-visual-plugin

Gives text-only models vision: forwards user images to an OpenAI-compatible vision model and shows the descriptions in a Web UI right panel.

911 hours agoVision, Voice & MultimodalMIT
C

Gemini-Eyes

consolesun/gemini-eyes

MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.

84 days agoVision, Voice & Multimodal
1

dsh-plugin-tts

1624318455/dsh-plugin-tts

Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.

77 hours agoVision, Voice & MultimodalMIT
D

dsh-imagegen

dickpy/dsh-imagegen

AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.

73 hours agoVision, Voice & MultimodalApache-2.0
C

dsh-chat-imagine

corrinehu/dsh-chat-imagine

Automatically generates and displays images in the DSH chat via API channels or local CLIs (supports mmx / codex / agy).

7yesterdayVision, Voice & MultimodalMIT
5

dsh-vision

54xkeee/dsh-vision

Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

72 days agoVision, Voice & MultimodalMIT
Z

dsh-voice-input-plugin

zhangbo-cn/dsh-voice-input-plugin

Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.

6yesterdayVision, Voice & MultimodalMIT
S

dsh-deepseek-vision

siegfly/dsh-deepseek-vision

A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.

623 hours agoVision, Voice & MultimodalMIT
M

dsh-windows-ocr

maxwell-feng/dsh-windows-ocr

Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

6yesterdayVision, Voice & MultimodalMIT
3

dsh-voice

3274375092/dsh-voice

Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.

43 days agoVision, Voice & MultimodalMIT
G

deepseek-vision (dsh-plugin-deepseek-vision)

gou-gee/deepseek-vision

Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.

42 days agoVision, Voice & MultimodalMIT
A

dsh-guide-dog

atropinoltt/dsh-guide-dog

MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.

4yesterdayVision, Voice & MultimodalMIT
H

dsh-her-eyes

huashenglian/dsh-her-eyes

一个可以让ai自动调用VLM(多模态模型)进行视觉分析的dsh插件。A dsh plugin that allows AI to automatically invoke VLMs (multimodal models) for visual analysis.

45 days agoVision, Voice & MultimodalMIT
A

dsh-speak

alan2z/dsh-speak

Voice-announce the final reply on Windows (SAPI5 natural voices) and macOS (system voice); skips reasoning and tool calls, one-line npm install.

32 days agoVision, Voice & MultimodalMIT
S

dsh-plugin-multimodal

shinjiyu/dsh-plugin-multimodal

Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.

32 days agoVision, Voice & MultimodalMIT
N

free-vision-skill

niyongsheng/free-vision-skill

Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.

33 days agoVision, Voice & MultimodalMIT