- Home
- Categories
- Vision, Voice & Multimodal
Vision, Voice & Multimodal
Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.
312 plugins found
dsh-ocr-local
balcoz/dsh-ocr-local
Local OCR for DeepSeek Harness: paste/attach an image, get its text via PP-OCRv5 + ONNX Runtime, fully offline. TUI (cc-tui) and Web. / DeepSeek Harness 本地 OCR 插件:图片转文字,PP-OCRv5 + ONNX Runtime,完全离线,支持 TUI 与 Web。
dsh-vision-bridge
sfyyy/dsh-vision-bridge
On-demand vision for text-only DSH sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
dsh-voice-input
0nt-one/dsh-voice-input
Mic button in the composer tool row: Web Speech API speech-to-text (Chrome/Edge), language switching, and optional auto-send, zero dependencies.
dsh-plugin-appshot
tauruswood/dsh-plugin-appshot
Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the composer for agent queries.
dsh-screenshot
paicat1/dsh-screenshot
Zero-dependency screen capture for DSH: Lightweight — zero deps, zero binaries; Stage & shoot — one-click full screen, window layout, hover-snap capture of occluded windows; Agent self-service — path-only delivery; paths are universal, pair with modlens (optional) for one-call structured evidence.
dsh-guide-dog
atropinoltt/dsh-guide-dog
MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.
canvas-workbench
elangan1997-cmyk/canvas-workbench
本地生图工作台,零订阅费:开自己的 API 生图(任意 OpenAI 兼容接口,零订阅),画布排版+修图/擦除/去背景/OCR/转矢量,可编辑 PSD/AI 交付,Photoshop/Illustrator 图层级双向桥接
dsh-tts-bridge
yuuyuko-uu/dsh-tts-bridge
Reads DSH conversations aloud using the DeepSeek web page's built-in read-aloud, driven by a small browser extension.
dsh-cycle-image-gen
chengzzzi44/dsh-cycle-image-gen
Image generation tool for DeepSeek Harness over any OpenAI-compatible images endpoint, with an inline Web UI gallery.
tripo3d-plugin-dsh
vast-ai-research/tripo3d-plugin-dsh
Tripo 3D skills for DeepSeek Harness — generate engine-ready 3D assets from a text prompt or a reference image via tripo-cli, with texturing, rigging, retopology and format conversion in one pass.
deepseek-vl-support
limccn/deepseek-vl-support
Give DeepSeek (text-only) models vision in Claude Code, Codex, and Agent Plugins clients: describe images via any OpenAI-compatible vision endpoint.
screenshot-feedback-hook-mcp (dsh-plugin)
lkh081231/screenshot-feedback-hook-mcp
Adds a take screenshot tool plus two optional automatic capture points, putting the screen into the conversation as an image block; requires an image-capable model.
dsh-pdf-reader
angeloszou/dsh-pdf-reader
A content-aware PDF reading plugin for vision models: profiles each page for figures (vector and raster), tables, formula risk and double-column layout, then applies content-aware hybrid extraction, rendering figure/table/formula pages as high-DPI region crops. Provides a low-resolution preview to understand the page layout, and renders a specified region at high resolution. Packaged as multiple tools for agents.
dsh-hos-scrcpy
ns-zzj/dsh-hos-scrcpy
Control a HarmonyOS phone from the DSH web UI with AI: live H.264 screen mirroring, mouse touch and system keys, hilog streaming, and agentic tools that let the model read the screen, locate UI controls, then tap, long-press, press keys, or type.
dsh-vision-hub (tool-vision)
xing666173/dsh-vision-hub
Enhanced vision toolbox: 14 pixel-level vision tools (describe, ground, detect, crop, pixel-diff, OCR, long-screenshot OCR, vectorize, colors, cutout, screenshot, present, materialize, html-screenshot) driven by one OpenAI-compatible endpoint, with clean \[图片: path] bridge markers, content-safety classification and rate-limit auto-retry.
dsh-image-pathify
dami9527/dsh-image-pathify
Lets text-only models handle pasted chat images, with a native vision experience, batch image viewing, and a built-in OpenAI-compatible analyze_image tool; vision-capable models are unaffected.
dsh-voice
stardustlc666/dsh-voice
Voice tools: free edge-tts neural speech synthesis, OpenAI-compatible ASR transcription, voice list, batch voice preview and health self-check.
dsh-voice
haoku123/dsh-voice
Full-duplex voice mode for the Web UI: tap-to-toggle or hold-to-talk dictation (send key or `Ctrl`) with a live caption, host-side SenseVoice ASR via sherpa-onnx, sentence-by-sentence spoken replies, and speaking interrupts playback and the running turn (true barge-in). No API key.
dsh-chatvoice
fuzzysoul/dsh-chatvoice
Free voice closed loop for the Web UI: browser SpeechRecognition mic input with live interim results plus read-aloud speaker buttons and auto-read for assistant replies, zero configuration and no API key.
dsh-draw-router
xiaozhe7772222/dsh-draw-router
Unified image generation router for DeepSeek Harness (DSH): auto-discovers image models from any OpenAI-compatible endpoint, provides draw_image and draw_list_sources tools, supports SenseNova, StepFun, Agnes, Qwen, Flux, SD, Imagen and more.
free-vision-skill
niyongsheng/free-vision-skill
Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.
deepseek-vision (dsh-plugin-deepseek-vision)
gou-gee/deepseek-vision
Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.
dsh-plugin-deepeye
favio8/dsh-plugin-deepeye
DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.
dsh-her-eyes
huashenglian/dsh-her-eyes
一个可以让ai自动调用VLM(多模态模型)进行视觉分析的dsh插件。A dsh plugin that allows AI to automatically invoke VLMs (multimodal models) for visual analysis.