Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

Y

dsh-plugin-vision-toolkit

yytbit/dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images

12 months agoVision, Voice & MultimodalMIT
P

dsh-yali-image-generator

pptt121212/dsh-yali-image-generator

DeepSeek-Harness 图像生成插件。申请 Yali AI API Key:https://api.yaliai.com/

12 months agoVision, Voice & MultimodalMIT
C

deepsee

chang416/deepsee

DeepSee: DeepSeek Harness vision, multi-model routing, and Gemini visual self-checks before delivery

12 months agoVision, Voice & MultimodalMIT
F

dsh-sight

fu3rte/dsh-sight

Plug-in vision for text-only DeepSeek Harness (dsh) models: a `vision` tool with built-in cheap/free VLM presets, multi-image batch analysis, paste-to-hint image admission, and a web settings page with hot-reload.

12 months agoVision, Voice & MultimodalMIT
W

dsh-friend

wanghehe123/dsh-friend

人格化伴侣插件 for DeepSeek Harness:角色卡、语音、Live2D、本地记忆与工作陪伴。

12 months agoVision, Voice & MultimodalMIT
A

dsh-voice-webspeech

anweat/dsh-voice-webspeech

Browser Web Speech API voice input: zero server, zero keys, zero model downloads (Edge=Azure, Chrome=Google speech).

1last monthVision, Voice & MultimodalMIT
R

dsh-plugins

retiredphysicist/dsh-plugins

DeepSeek Harness plugin: web browsing tools (markdown / screenshot / pdf / crawl) powered by Cloudflare Browser Run — real headless Chrome with JS rendering, login sessions and WebMCP support.

02 days agoVision, Voice & Multimodal
W

dsh-kite

weibaohui/dsh-kite

Animated kite that responds to agent activity, with configurable kite frames, patterns and colors, plus user-supplied image textures.

04 days agoVision, Voice & Multimodal
K

dsh-zcode-cli-proxy

kyle123740/dsh-zcode-cli-proxy

ZCode CLI app-server relay for DeepSeek Harness, with GLM streaming, image input and session reuse using an existing Start Plan login; requires the ZCode client.

03 days agoVision, Voice & MultimodalAGPL-3.0
E

dsh-fal-imagegen

enchanted0911/dsh-fal-imagegen

fal.ai native image generation for DSH: a FAL_KEY settings card that follows DSH's language (zh/en, English default) plus Agent tools (fal_generate_image async / fal_edit_image / fal_get_image_task / fal_list_image_models) that call queue.fal.run directly

09 days agoVision, Voice & MultimodalMIT
Y

dsh-video-to-notes

yll-kb/dsh-video-to-notes

Opt-in DeepSeek Harness skill bundle that turns course, lecture, tutorial, documentary, meeting, and talk videos into structured study notes.

012 days agoVision, Voice & MultimodalMIT
W

dsh-voice-danmaku

weizhida/dsh-voice-danmaku

这是dsh的插件,语音发送弹幕。在玩游戏时通过语音输入在b站发弹幕,不切出游戏可以正常操作

04 days agoVision, Voice & MultimodalMIT
T

dsh-xiaozhi

toddpan/dsh-xiaozhi

Connects the Xiaozhi voice assistant to DSH Web over MCP: DSH is the tool provider, weaving 35 DSH Web endpoints into 16 voice-friendly tools for workspaces, sessions, chat, models, settings and files — outbound WebSocket to the Xiaozhi MCP access point by default (no public IP or port forwarding needed), several devices bound at once, plus a DSH settings page with live status.

06 days agoVision, Voice & Multimodal
J

sh-volume-knob

jianghu-lao-yao/sh-volume-knob

Speaker button beside the composer microphone — one click scrolls to the start of your newest question, marks it with a blinking caret and reads from there through the newest agent reply (dsh-tts, browser voice as fallback); press-and-drag-right picks any other reading start position on the page, press-and-drag-up opens a vertical mixer for in-page media volume and system output volume.

08 days agoVision, Voice & MultimodalMIT
M

lookover (dsh-look)

mengxiaoxian/lookover

Scene-awareness probe for DSH on macOS: a privacy-first, pull-model `look` tool that reads the frontmost non-self window (app info, AX title/selection, gated local OCR), plus a summon-hotkey snapshot captured the instant you press.

011 days agoVision, Voice & MultimodalMIT
Y

dsh-agnes-gen

ylhow06/dsh-agnes-gen

Agnes AI image/video generation tools (agnes_image / agnes_video) for DSH (DeepSeek Harness). Built-in cross-process RPM rate limiting, 429 backoff and local ffmpeg GIF conversion.

011 days agoVision, Voice & MultimodalMIT
G

dsh-plugin-notify

goodandready/dsh-plugin-notify

DSH plugin: audio chimes, cross-session toasts, desktop push, and IM webhooks for turn completion, errors, and approvals.

011 days agoVision, Voice & MultimodalMIT
F

dsh-say

fangqian616/dsh-say

Speaks your agent's reports in a voice you choose, compressing long reports before speaking. Character voices need no training - a 3-10 second reference clip clones one, and community-trained models work too - and no 6.4 GB GPT-SoVITS install: the plugin installs its own runtime and voice. An existing GPT-SoVITS can be used instead, and it is the same voice model.

012 days agoVision, Voice & Multimodal
C

dsh-read-aloud

cccc12138/dsh-read-aloud

Adds a speaker button immediately right of the Like button on every finalized assistant reply; click it to hear the reply through the browser speech engine, and hover it for a speed and voice panel.

015 days agoVision, Voice & MultimodalMIT
L

dsh-voice-alert

loyalchiiina/dsh-voice-alert

Speaks or plays a sound at the end of every turn and plays a failure cue when a tool call or a turn errors. Ships 20 built-in effects (10 alert chimes + 10 nature sounds) and defaults to effect mode, so it needs no API key and no audio files; an optional Volcengine voice-clone route generates the three announcement clips in one click. Plays through winmm/waveOut as-is, never changing the system volume or mute state. Windows only.

04 days agoVision, Voice & MultimodalMIT
D

dsh-screen-eye

davidekingsss/dsh-screen-eye

Agent screen capture on macOS and Windows: one tool captures the screen and returns the image itself, so the model sees it without a second call. On Windows a whole burst runs in one engine call; on macOS a resident helper drops a region capture to 13ms and a change check to 23ms. A second tool, macOS only, reports whether Screen Recording is granted and opens the settings pane that fixes it.

014 days agoVision, Voice & MultimodalMIT
N

vision-exp-tile

nicholaskin/vision-exp-tile

Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.

022 days agoVision, Voice & MultimodalMIT
Z

dsh-stt-plugin

zemanzhang809/dsh-stt-plugin

A speech-to-text (voice input) plugin for DeepSeek Harness.

015 days agoVision, Voice & MultimodalMIT
B

dsh-asr-voice

bittersmilezzz/dsh-asr-voice

开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。

018 days agoVision, Voice & MultimodalMIT