Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

311 plugins found

Z

dsh-web-ui (dsh-tool-describe-image)

zhu1090093659/dsh-web-ui

A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.

8k8 days agoVision, Voice & MultimodalApache-2.0
D

ipollowork

devin-axis/ipollowork

iPolloWork HyperFrames Video Studio and 27 editable video templates as a native DeepSeek Harness conversation view.

4.2k2 months agoVision, Voice & Multimodal
L

modlens

liustack/modlens

Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).

4.1k7 days agoVision, Voice & MultimodalMIT
Y

dsh-vision-router

ysr666/dsh-vision-router

Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.

1.1k18 hours agoVision, Voice & MultimodalMIT
A

dsh-vision-toolkit

anionex/dsh-vision-toolkit

Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.

8853 days agoVision, Voice & MultimodalMIT
T

tongflow

tong-io/tongflow

TongFlow studio plugin for DeepSeek Harness (dsh): film-crew style project model, agent-authored TongFlow workflows, deterministic media generation, embedded canvas.

8702 months agoVision, Voice & MultimodalAGPL-3.0
S

dsh-image-gen

shanliuling/dsh-image-gen

Native conversational image generation for DeepSeek Harness: ask the agent to create an image, and it handles generation and keeps the result directly in the conversation.

5694 days agoVision, Voice & MultimodalApache-2.0
O

watch-skill

oxbshw/watch-skill

DeepWatch's capabilities, installable into an existing DeepSeek Harness profile

37820 days agoVision, Voice & MultimodalMIT
D

dsh-imagegen

dickpy/dsh-imagegen

AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.

994 days agoVision, Voice & MultimodalApache-2.0
F

dsh-comfyui

fandc520/dsh-comfyui

Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.

947 days agoVision, Voice & MultimodalMIT
P

dsh-omi-voice

polinnizhong/dsh-omi-voice

In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.

74last monthVision, Voice & MultimodalMIT
C

deepseek-harness-video-director

chiphoton/deepseek-harness-video-director

🎬AI-Powered Director: MiniMax-H3 Video Generation Plugin Directed by DeepSeek-Harness. 🤖More than Prompting. 🪄Canvas UI.

696 days agoVision, Voice & MultimodalMIT
S

dsh-design-qa

sunxin-ai/dsh-design-qa

Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.

4427 days agoVision, Voice & MultimodalMIT
J

picturereader

jing-hy/picturereader

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

374 days agoVision, Voice & MultimodalMIT
P

dsh-voice-scribe

pensivefei/dsh-voice-scribe

Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.

344 days agoVision, Voice & MultimodalMIT
O

dsh-vision

oil-oil/dsh-vision

Near-native image understanding for DeepSeek Harness

242 months agoVision, Voice & MultimodalMIT
D

dsh-web-ui (dsh-tool-describe-image)

damonkoy/dsh-web-ui

Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.

222 months agoVision, Voice & MultimodalApache-2.0
W

dsh-ears

wiziscool/dsh-ears

Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.

213 days agoVision, Voice & MultimodalMIT
1

dsh-plugin-tts

1624318455/dsh-plugin-tts

Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.

217 days agoVision, Voice & MultimodalMIT
M

dsh-media-skills

mjorgin/dsh-media-skills

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

1913 days agoVision, Voice & MultimodalMIT
D

dsh-aigc-canvas

dsh-external/dsh-aigc-canvas

AIGC canvas plugin (cordis).

1723 days agoVision, Voice & Multimodal
Q

dsh-voice-mode (dsh-voice-mode)

qishuilalala/dsh-voice-mode

Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.

16yesterdayVision, Voice & MultimodalMIT
P

dsh-talk

perrylink/dsh-talk

Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.

1610 days agoVision, Voice & MultimodalApache-2.0
J

dsh-visual-plugin

jyh20030112/dsh-visual-plugin

Gives text-only models vision: forwards user images to an OpenAI-compatible vision model and shows the descriptions in a Web UI right panel.

164 days agoVision, Voice & MultimodalMIT