Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

B

dsh-ocr-local

balcoz/dsh-ocr-local

Local OCR for DeepSeek Harness: paste/attach an image, get its text via PP-OCRv5 + ONNX Runtime, fully offline. TUI (cc-tui) and Web. / DeepSeek Harness 本地 OCR 插件:图片转文字,PP-OCRv5 + ONNX Runtime,完全离线,支持 TUI 与 Web。

52 months agoVision, Voice & Multimodal
S

dsh-vision-bridge

sfyyy/dsh-vision-bridge

On-demand vision for text-only DSH sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model

52 months agoVision, Voice & MultimodalMIT
0

dsh-voice-input

0nt-one/dsh-voice-input

Mic button in the composer tool row: Web Speech API speech-to-text (Chrome/Edge), language switching, and optional auto-send, zero dependencies.

52 months agoVision, Voice & MultimodalMIT
T

dsh-plugin-appshot

tauruswood/dsh-plugin-appshot

Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the composer for agent queries.

55 days agoVision, Voice & MultimodalMIT
P

dsh-screenshot

paicat1/dsh-screenshot

Zero-dependency screen capture for DSH: Lightweight — zero deps, zero binaries; Stage & shoot — one-click full screen, window layout, hover-snap capture of occluded windows; Agent self-service — path-only delivery; paths are universal, pair with modlens (optional) for one-call structured evidence.

525 days agoVision, Voice & MultimodalMIT
A

dsh-guide-dog

atropinoltt/dsh-guide-dog

MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.

52 months agoVision, Voice & MultimodalMIT
E

canvas-workbench

elangan1997-cmyk/canvas-workbench

本地生图工作台,零订阅费:开自己的 API 生图(任意 OpenAI 兼容接口,零订阅),画布排版+修图/擦除/去背景/OCR/转矢量,可编辑 PSD/AI 交付,Photoshop/Illustrator 图层级双向桥接

43 days agoVision, Voice & MultimodalMIT
Y

dsh-tts-bridge

yuuyuko-uu/dsh-tts-bridge

Reads DSH conversations aloud using the DeepSeek web page's built-in read-aloud, driven by a small browser extension.

43 days agoVision, Voice & Multimodal
C

dsh-cycle-image-gen

chengzzzi44/dsh-cycle-image-gen

Image generation tool for DeepSeek Harness over any OpenAI-compatible images endpoint, with an inline Web UI gallery.

422 days agoVision, Voice & MultimodalMIT
V

tripo3d-plugin-dsh

vast-ai-research/tripo3d-plugin-dsh

Tripo 3D skills for DeepSeek Harness — generate engine-ready 3D assets from a text prompt or a reference image via tripo-cli, with texturing, rigging, retopology and format conversion in one pass.

424 days agoVision, Voice & Multimodal
L

deepseek-vl-support

limccn/deepseek-vl-support

Give DeepSeek (text-only) models vision in Claude Code, Codex, and Agent Plugins clients: describe images via any OpenAI-compatible vision endpoint.

4last monthVision, Voice & MultimodalMIT
L

screenshot-feedback-hook-mcp (dsh-plugin)

lkh081231/screenshot-feedback-hook-mcp

Adds a take screenshot tool plus two optional automatic capture points, putting the screen into the conversation as an image block; requires an image-capable model.

4last monthVision, Voice & MultimodalMIT
A

dsh-pdf-reader

angeloszou/dsh-pdf-reader

A content-aware PDF reading plugin for vision models: profiles each page for figures (vector and raster), tables, formula risk and double-column layout, then applies content-aware hybrid extraction, rendering figure/table/formula pages as high-DPI region crops. Provides a low-resolution preview to understand the page layout, and renders a specified region at high resolution. Packaged as multiple tools for agents.

4last monthVision, Voice & MultimodalMIT
N

dsh-hos-scrcpy

ns-zzj/dsh-hos-scrcpy

Control a HarmonyOS phone from the DSH web UI with AI: live H.264 screen mirroring, mouse touch and system keys, hilog streaming, and agentic tools that let the model read the screen, locate UI controls, then tap, long-press, press keys, or type.

42 days agoVision, Voice & MultimodalMIT
X

dsh-vision-hub (tool-vision)

xing666173/dsh-vision-hub

Enhanced vision toolbox: 14 pixel-level vision tools (describe, ground, detect, crop, pixel-diff, OCR, long-screenshot OCR, vectorize, colors, cutout, screenshot, present, materialize, html-screenshot) driven by one OpenAI-compatible endpoint, with clean \[图片: path] bridge markers, content-safety classification and rate-limit auto-retry.

4last monthVision, Voice & Multimodal
D

dsh-image-pathify

dami9527/dsh-image-pathify

Lets text-only models handle pasted chat images, with a native vision experience, batch image viewing, and a built-in OpenAI-compatible analyze_image tool; vision-capable models are unaffected.

48 days agoVision, Voice & MultimodalMIT
S

dsh-voice

stardustlc666/dsh-voice

Voice tools: free edge-tts neural speech synthesis, OpenAI-compatible ASR transcription, voice list, batch voice preview and health self-check.

413 hours agoVision, Voice & MultimodalMIT
H

dsh-voice

haoku123/dsh-voice

Full-duplex voice mode for the Web UI: tap-to-toggle or hold-to-talk dictation (send key or `Ctrl`) with a live caption, host-side SenseVoice ASR via sherpa-onnx, sentence-by-sentence spoken replies, and speaking interrupts playback and the running turn (true barge-in). No API key.

4last monthVision, Voice & MultimodalMIT
F

dsh-chatvoice

fuzzysoul/dsh-chatvoice

Free voice closed loop for the Web UI: browser SpeechRecognition mic input with live interim results plus read-aloud speaker buttons and auto-read for assistant replies, zero configuration and no API key.

42 months agoVision, Voice & MultimodalMIT
X

dsh-draw-router

xiaozhe7772222/dsh-draw-router

Unified image generation router for DeepSeek Harness (DSH): auto-discovers image models from any OpenAI-compatible endpoint, provides draw_image and draw_list_sources tools, supports SenseNova, StepFun, Agnes, Qwen, Flux, SD, Imagen and more.

4last monthVision, Voice & MultimodalMIT
N

free-vision-skill

niyongsheng/free-vision-skill

Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.

4last monthVision, Voice & MultimodalMIT
G

deepseek-vision (dsh-plugin-deepseek-vision)

gou-gee/deepseek-vision

Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.

426 days agoVision, Voice & MultimodalMIT
F

dsh-plugin-deepeye

favio8/dsh-plugin-deepeye

DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.

42 months agoVision, Voice & Multimodal
H

dsh-her-eyes

huashenglian/dsh-her-eyes

一个可以让ai自动调用VLM(多模态模型)进行视觉分析的dsh插件。A dsh plugin that allows AI to automatically invoke VLMs (multimodal models) for visual analysis.

42 months agoVision, Voice & MultimodalMIT