- Home
- Categories
- Vision, Voice & Multimodal
Vision, Voice & Multimodal
Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.
312 plugins found
dsh-plugin-vision-toolkit
yytbit/dsh-plugin-vision-toolkit
Vision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images
dsh-yali-image-generator
pptt121212/dsh-yali-image-generator
DeepSeek-Harness 图像生成插件。申请 Yali AI API Key:https://api.yaliai.com/
deepsee
chang416/deepsee
DeepSee: DeepSeek Harness vision, multi-model routing, and Gemini visual self-checks before delivery
dsh-sight
fu3rte/dsh-sight
Plug-in vision for text-only DeepSeek Harness (dsh) models: a `vision` tool with built-in cheap/free VLM presets, multi-image batch analysis, paste-to-hint image admission, and a web settings page with hot-reload.
dsh-friend
wanghehe123/dsh-friend
人格化伴侣插件 for DeepSeek Harness:角色卡、语音、Live2D、本地记忆与工作陪伴。
dsh-voice-webspeech
anweat/dsh-voice-webspeech
Browser Web Speech API voice input: zero server, zero keys, zero model downloads (Edge=Azure, Chrome=Google speech).
dsh-plugins
retiredphysicist/dsh-plugins
DeepSeek Harness plugin: web browsing tools (markdown / screenshot / pdf / crawl) powered by Cloudflare Browser Run — real headless Chrome with JS rendering, login sessions and WebMCP support.
dsh-kite
weibaohui/dsh-kite
Animated kite that responds to agent activity, with configurable kite frames, patterns and colors, plus user-supplied image textures.
dsh-zcode-cli-proxy
kyle123740/dsh-zcode-cli-proxy
ZCode CLI app-server relay for DeepSeek Harness, with GLM streaming, image input and session reuse using an existing Start Plan login; requires the ZCode client.
dsh-fal-imagegen
enchanted0911/dsh-fal-imagegen
fal.ai native image generation for DSH: a FAL_KEY settings card that follows DSH's language (zh/en, English default) plus Agent tools (fal_generate_image async / fal_edit_image / fal_get_image_task / fal_list_image_models) that call queue.fal.run directly
dsh-video-to-notes
yll-kb/dsh-video-to-notes
Opt-in DeepSeek Harness skill bundle that turns course, lecture, tutorial, documentary, meeting, and talk videos into structured study notes.
dsh-voice-danmaku
weizhida/dsh-voice-danmaku
这是dsh的插件,语音发送弹幕。在玩游戏时通过语音输入在b站发弹幕,不切出游戏可以正常操作
dsh-xiaozhi
toddpan/dsh-xiaozhi
Connects the Xiaozhi voice assistant to DSH Web over MCP: DSH is the tool provider, weaving 35 DSH Web endpoints into 16 voice-friendly tools for workspaces, sessions, chat, models, settings and files — outbound WebSocket to the Xiaozhi MCP access point by default (no public IP or port forwarding needed), several devices bound at once, plus a DSH settings page with live status.
sh-volume-knob
jianghu-lao-yao/sh-volume-knob
Speaker button beside the composer microphone — one click scrolls to the start of your newest question, marks it with a blinking caret and reads from there through the newest agent reply (dsh-tts, browser voice as fallback); press-and-drag-right picks any other reading start position on the page, press-and-drag-up opens a vertical mixer for in-page media volume and system output volume.
lookover (dsh-look)
mengxiaoxian/lookover
Scene-awareness probe for DSH on macOS: a privacy-first, pull-model `look` tool that reads the frontmost non-self window (app info, AX title/selection, gated local OCR), plus a summon-hotkey snapshot captured the instant you press.
dsh-agnes-gen
ylhow06/dsh-agnes-gen
Agnes AI image/video generation tools (agnes_image / agnes_video) for DSH (DeepSeek Harness). Built-in cross-process RPM rate limiting, 429 backoff and local ffmpeg GIF conversion.
dsh-plugin-notify
goodandready/dsh-plugin-notify
DSH plugin: audio chimes, cross-session toasts, desktop push, and IM webhooks for turn completion, errors, and approvals.
dsh-say
fangqian616/dsh-say
Speaks your agent's reports in a voice you choose, compressing long reports before speaking. Character voices need no training - a 3-10 second reference clip clones one, and community-trained models work too - and no 6.4 GB GPT-SoVITS install: the plugin installs its own runtime and voice. An existing GPT-SoVITS can be used instead, and it is the same voice model.
dsh-read-aloud
cccc12138/dsh-read-aloud
Adds a speaker button immediately right of the Like button on every finalized assistant reply; click it to hear the reply through the browser speech engine, and hover it for a speed and voice panel.
dsh-voice-alert
loyalchiiina/dsh-voice-alert
Speaks or plays a sound at the end of every turn and plays a failure cue when a tool call or a turn errors. Ships 20 built-in effects (10 alert chimes + 10 nature sounds) and defaults to effect mode, so it needs no API key and no audio files; an optional Volcengine voice-clone route generates the three announcement clips in one click. Plays through winmm/waveOut as-is, never changing the system volume or mute state. Windows only.
dsh-screen-eye
davidekingsss/dsh-screen-eye
Agent screen capture on macOS and Windows: one tool captures the screen and returns the image itself, so the model sees it without a second call. On Windows a whole burst runs in one engine call; on macOS a resident helper drops a region capture to 13ms and a change check to 23ms. A second tool, macOS only, reports whether Screen Recording is granted and opens the settings pane that fixes it.
vision-exp-tile
nicholaskin/vision-exp-tile
Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.
dsh-stt-plugin
zemanzhang809/dsh-stt-plugin
A speech-to-text (voice input) plugin for DeepSeek Harness.
dsh-asr-voice
bittersmilezzz/dsh-asr-voice
开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。