跳过主要内容

视觉、语音与多模态

视觉、语音与多模态类插件让 DeepSeek-Harness(dsh)里纯文本的 DeepSeek 模型也能"看图""听声""开口"。包含把粘贴的截图交给智谱 GLM、Gemini、豆包或本地 Ollama 做 OCR 与图像描述的 vision 工具和供应商路由、基于浏览器 Web Speech API 或 Whisper 兼容接口的麦克风语音输入、使用 Edge TTS 或自定义音色的朗读回复与全双工语音模式,以及图片、视频、音乐生成工具。

共 312 个插件

Z

dsh-tier-router

zhangzhangco/dsh-tier-router

Automatic tier-based model routing for DeepSeek Harness: a virtual `smart` model classifies each request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured. 三级难度 + 视觉自动路由,虚拟 smart 模型零配置接入。

017天前视觉、语音与多模态
R

dsh-voice-input

rio-promax/dsh-voice-input

DSH voice input plugin: browser realtime + VAD dictation + local/cloud ASR + DeepSeek AI polish

021天前视觉、语音与多模态MIT
C

dsh-audio-visualizer

caesarjue/dsh-audio-visualizer

System-audio driven UI visualizer for DeepSeek Harness: a draggable 48-band spectrum chip plus a frame-wide bass glow (memory-only FFT).

021天前视觉、语音与多模态MIT
Z

dsh-plugin-image-picker

zg2017/dsh-plugin-image-picker

Adds a file-picker 'attach image' button to the composer toolbar, since DSH's own web client only supports paste and drag-and-drop for image attachments (v1 scope, per its own upstream design notes) - a real gap for touch devices with no drag-and-drop and

0上个月视觉、语音与多模态MIT
N

dsh-desktop

new-256/dsh-desktop

DSH (DeepSeek Harness) ComfyUI 桥接插件:让 Agent 直接驱动本地/局域网 ComfyUI 生图生视频 — 8 个全局模型工具、预设工作流模板(txt2img/img2img/Wan/SVD/H3)、任意 API 工作流逃生舱、设备能力守卫、模型注册表断点续传下载、ComfyUI 一键安装评估。ComfyUI bridge plugin for DeepSeek Harness: image & video generation tools driving a local

023天前视觉、语音与多模态MIT
S

dsh-ui-tool-result-images

suntianc/dsh-ui-tool-result-images

DeepSeek Harness Web plugin that keeps image-bearing Tool results visible after Compact transcript folding

023天前视觉、语音与多模态MIT
C

dsh-voice-mimo

ch1bug/dsh-voice-mimo

Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).

019天前视觉、语音与多模态MIT
Z

dsh-image-router

zhiwuli0228/dsh-image-router

在准入前把提示词里的图片换成视觉模型的文字分析,因此任何模型(包括纯文本模型)都能读图且会话不切换模型;另提供 describe_image 工具处理图片路径。

020天前视觉、语音与多模态MIT
M

dsh-image-guard

mafeis/dsh-image-guard

发送前把请求中的历史图片裁剪至保留张数,并从上游 400 错误解析单次图片上限、按更少张数降级重试,保证含图会话不被图片数量限制打断。

019天前视觉、语音与多模态MIT
J

dsh-mathmatic-symbol

jaxzhou/dsh-mathmatic-symbol

Three tools for DeepSeek Harness: typeset LaTeX formulas into images, draw mathematical figures from a declarative spec, and convert a formula or SVG into an image ready to embed in a document.

021天前视觉、语音与多模态MIT
J

dsh-voice-input-qwen-asr

jsoncode/dsh-voice-input-qwen-asr

Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run ser

028天前视觉、语音与多模态
2

dsh-tts-flash

2021heei/dsh-tts-flash

AI 回复流式 TTS 朗读插件,思考等待期有文字与语音同步的趣味短语反馈。内置 Edge TTS,支持任意 OpenAI 兼容云端引擎。

023天前视觉、语音与多模态MIT
W

dsh-image-generation

whites18/dsh-image-generation

在设置中配置多个生图供应商,对话通过 image_generate 调用当前选中的模型;图片保存到工作区 generate/image,并在对话中内联显示。

05天前视觉、语音与多模态MIT
A

dsh-reelsmaker

aayan-cloud/dsh-reelsmaker

DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.

028天前视觉、语音与多模态
M

dsh-voice-input-en

mohith-das/dsh-voice-input-en

Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow

0上个月视觉、语音与多模态MIT
W

dsh-plugins (dsh-ding-sound)

wwweljf/dsh-plugins

DSH 对话结束提示音:内置网络热梗原声(你干嘛哎呦、鸡你太美、神鹰哥等),设置面板可试听/切换/随机,支持自定义音频目录。

0上个月视觉、语音与多模态MIT
K

dsh-kitt-voice

kittcat-lab/dsh-kitt-voice

Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.

028天前视觉、语音与多模态MIT
J

dsh-plugin-show-image

justhalfbit/dsh-plugin-show-image

DeepSeek Harness (DSH) 会话内图片渲染插件:全局 show_image 工具 + 点击放大 lightbox。 | Inline image rendering plugin for DSH: global show_image tool + click-to-enlarge lightbox.

0上个月视觉、语音与多模态MIT
A

dsh-plugins

aetheri-ai/dsh-plugins

DeepSeek Harness plugin: a model-facing show_image tool that presents a local image to the human viewer in the conversation.

0上个月视觉、语音与多模态MIT
S

dsh-voice-control

sucriss/dsh-voice-control

DSH 网页端语音控制:按住说话将语音转文字输入发送框(可自动发送),用浏览器语音合成朗读助手回复;右键麦克风打开设置面板,支持 Ctrl+M 全局快捷键。

0上个月视觉、语音与多模态MIT
S

dsh-vision-pro-bridge

shainedemo/dsh-vision-pro-bridge

视觉桥:贴图先经 deepseek-v4-flash-vision-exp 转写为文字,再交给纯文本的 deepseek-v4-pro 回答,零第三方依赖。

0上个月视觉、语音与多模态MIT
H

dsh-maclens

harzva/dsh-maclens

苹果设备端 Vision 工具:本地 OCR(含中文)、图像分类、人脸检测、文档版面分析与综合描述,100% 离线、无需 API key,支持长截图切片。

0上个月视觉、语音与多模态MIT
A

dsh-image-generation (tool-image-generation)

ankye/dsh-image-generation

面向模型的图像生成工具,提供可配置通道和标准化图像参数。

0上个月视觉、语音与多模态MIT
A

dsh-vision-bridge

alaxrpg/dsh-vision-bridge

通过已配置的 DSH Provider 或 OpenAI 兼容端点提供图片输入与识别。

05天前视觉、语音与多模态MIT