视觉、语音与多模态
视觉、语音与多模态类插件让 DeepSeek-Harness(dsh)里纯文本的 DeepSeek 模型也能"看图""听声""开口"。包含把粘贴的截图交给智谱 GLM、Gemini、豆包或本地 Ollama 做 OCR 与图像描述的 vision 工具和供应商路由、基于浏览器 Web Speech API 或 Whisper 兼容接口的麦克风语音输入、使用 Edge TTS 或自定义音色的朗读回复与全双工语音模式,以及图片、视频、音乐生成工具。
共 312 个插件
dsh-tier-router
zhangzhangco/dsh-tier-router
Automatic tier-based model routing for DeepSeek Harness: a virtual `smart` model classifies each request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured. 三级难度 + 视觉自动路由,虚拟 smart 模型零配置接入。
dsh-voice-input
rio-promax/dsh-voice-input
DSH voice input plugin: browser realtime + VAD dictation + local/cloud ASR + DeepSeek AI polish
dsh-audio-visualizer
caesarjue/dsh-audio-visualizer
System-audio driven UI visualizer for DeepSeek Harness: a draggable 48-band spectrum chip plus a frame-wide bass glow (memory-only FFT).
dsh-plugin-image-picker
zg2017/dsh-plugin-image-picker
Adds a file-picker 'attach image' button to the composer toolbar, since DSH's own web client only supports paste and drag-and-drop for image attachments (v1 scope, per its own upstream design notes) - a real gap for touch devices with no drag-and-drop and
dsh-desktop
new-256/dsh-desktop
DSH (DeepSeek Harness) ComfyUI 桥接插件:让 Agent 直接驱动本地/局域网 ComfyUI 生图生视频 — 8 个全局模型工具、预设工作流模板(txt2img/img2img/Wan/SVD/H3)、任意 API 工作流逃生舱、设备能力守卫、模型注册表断点续传下载、ComfyUI 一键安装评估。ComfyUI bridge plugin for DeepSeek Harness: image & video generation tools driving a local
dsh-ui-tool-result-images
suntianc/dsh-ui-tool-result-images
DeepSeek Harness Web plugin that keeps image-bearing Tool results visible after Compact transcript folding
dsh-voice-mimo
ch1bug/dsh-voice-mimo
Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).
dsh-image-router
zhiwuli0228/dsh-image-router
在准入前把提示词里的图片换成视觉模型的文字分析,因此任何模型(包括纯文本模型)都能读图且会话不切换模型;另提供 describe_image 工具处理图片路径。
dsh-image-guard
mafeis/dsh-image-guard
发送前把请求中的历史图片裁剪至保留张数,并从上游 400 错误解析单次图片上限、按更少张数降级重试,保证含图会话不被图片数量限制打断。
dsh-mathmatic-symbol
jaxzhou/dsh-mathmatic-symbol
Three tools for DeepSeek Harness: typeset LaTeX formulas into images, draw mathematical figures from a declarative spec, and convert a formula or SVG into an image ready to embed in a document.
dsh-voice-input-qwen-asr
jsoncode/dsh-voice-input-qwen-asr
Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run ser
dsh-tts-flash
2021heei/dsh-tts-flash
AI 回复流式 TTS 朗读插件,思考等待期有文字与语音同步的趣味短语反馈。内置 Edge TTS,支持任意 OpenAI 兼容云端引擎。
dsh-image-generation
whites18/dsh-image-generation
在设置中配置多个生图供应商,对话通过 image_generate 调用当前选中的模型;图片保存到工作区 generate/image,并在对话中内联显示。
dsh-reelsmaker
aayan-cloud/dsh-reelsmaker
DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.
dsh-voice-input-en
mohith-das/dsh-voice-input-en
Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow
dsh-plugins (dsh-ding-sound)
wwweljf/dsh-plugins
DSH 对话结束提示音:内置网络热梗原声(你干嘛哎呦、鸡你太美、神鹰哥等),设置面板可试听/切换/随机,支持自定义音频目录。
dsh-kitt-voice
kittcat-lab/dsh-kitt-voice
Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.
dsh-plugin-show-image
justhalfbit/dsh-plugin-show-image
DeepSeek Harness (DSH) 会话内图片渲染插件:全局 show_image 工具 + 点击放大 lightbox。 | Inline image rendering plugin for DSH: global show_image tool + click-to-enlarge lightbox.
dsh-plugins
aetheri-ai/dsh-plugins
DeepSeek Harness plugin: a model-facing show_image tool that presents a local image to the human viewer in the conversation.
dsh-voice-control
sucriss/dsh-voice-control
DSH 网页端语音控制:按住说话将语音转文字输入发送框(可自动发送),用浏览器语音合成朗读助手回复;右键麦克风打开设置面板,支持 Ctrl+M 全局快捷键。
dsh-vision-pro-bridge
shainedemo/dsh-vision-pro-bridge
视觉桥:贴图先经 deepseek-v4-flash-vision-exp 转写为文字,再交给纯文本的 deepseek-v4-pro 回答,零第三方依赖。
dsh-maclens
harzva/dsh-maclens
苹果设备端 Vision 工具:本地 OCR(含中文)、图像分类、人脸检测、文档版面分析与综合描述,100% 离线、无需 API key,支持长截图切片。
dsh-image-generation (tool-image-generation)
ankye/dsh-image-generation
面向模型的图像生成工具,提供可配置通道和标准化图像参数。
dsh-vision-bridge
alaxrpg/dsh-vision-bridge
通过已配置的 DSH Provider 或 OpenAI 兼容端点提供图片输入与识别。