视觉、语音与多模态
视觉、语音与多模态类插件让 DeepSeek-Harness(dsh)里纯文本的 DeepSeek 模型也能"看图""听声""开口"。包含把粘贴的截图交给智谱 GLM、Gemini、豆包或本地 Ollama 做 OCR 与图像描述的 vision 工具和供应商路由、基于浏览器 Web Speech API 或 Whisper 兼容接口的麦克风语音输入、使用 Edge TTS 或自定义音色的朗读回复与全双工语音模式,以及图片、视频、音乐生成工具。
共 312 个插件
dsh-voice-announcer
flashyiyi/dsh-voice-announcer
对话结束语音播报(会话名、轮数、结果),回复生成时实时逐句朗读;内置 edge-tts,零第三方依赖。
dsh-soundscape
berserk0501/dsh-soundscape
双 MediaPlayer 守护进程($think 思考循环 + $fx 一次性音效)支持音量闪避、情绪感知音效(连续成功/失败/长时间运行自动触发 smooth/struggle/deep_think)、15 种正弦/方波合成内置音、51 个工具分类映射、自定义 WAV/MP3 直替自动转换,以及完整设置页含深色模式修复。
deepseek-plugin
klingai-dev/deepseek-plugin
一句话,让灵感从想法变成大片。在 DeepSeek Harness 中直接用自然语言调用可灵 AI,支持文生图、图生图、文生视频、图生视频、参考图创作、任务跟踪和结果预览。
dsh-auto-vision
k2d5rqjpkg-art/dsh-auto-vision
按需自动把 DeepSeek 路由切到视觉模型:flash 主会话直接切(A),pro 保持深推理并交给视觉子代理读图(B),子代理恒切,含失败回退,无需手动切模型。
dsh-evidence
cooberped/dsh-evidence
Turns attached files into versioned evidence: `search_documents` builds a private local index (SQLite FTS5 after a startup capability probe, dependency-free JS fallback otherwise) and returns compact evidence blocks carrying an exact coordinate — PDF page, PPTX slide, text/DOCX line range, or quoted XLSX `Sheet!Range` — which `read_document` expands only while the content version still matches. Contiguous CJK runs are indexed as overlapping bigrams and queried as phrases, so word order is preserved; uploaded raster images take the native vision attachment path instead.
dsh-tu4-inline-images
zehenk/dsh-tu4-inline-images
对话内联图片 DSH 插件 — 在 DeepSeek Harness (DSH) Web GUI 的对话中,出现本地图片路径即直接渲染为图片。 A DSH plugin that renders local image paths as inline images in DeepSeek Harness (DSH) web conversations. Security-first: loopback-only route, strong per-process token, multi-root realpath whitelist.
dsh-llm-capabilities
bamboostrip/dsh-llm-capabilities
DSH plugin: auto-detect and configure model capabilities (reasoningEfforts + input modalities) for llm-pi-ai. Successor to dsh-reasoning-efforts.
dsh-vision-toggle
lijian-ui/dsh-vision-toggle
Per-model vision (image input) toggle for DeepSeek Harness (dsh): list every configured model and flip a switch to enable/disable image support without hand-editing settings.yaml. 为 DeepSeek Harness 提供按模型的「支持图片」开关:无需手改 settings.yaml。
dsh-speech
allmodels-io/dsh-speech
Streaming speech-to-text for DeepSeek Harness using AllModels.io.
dsh-agnes
chaoliu615/dsh-agnes
Agnes AI image & video generation tools for DSH (agnes_image_generate / agnes_video_generate)
dsh-remote-deliver
demacia1314/dsh-remote-deliver
🚀 告别繁琐 SCP!远程部署 DSH 一键下载修改后的文件与图片预览交付插件
dsh-dictation
wsl043/dsh-dictation
Editable local and desktop dictation for DeepSeek Harness
dsh-file-attach
lucasxingg/dsh-file-attach
Drag-and-drop PDF, Office, images, and text/code files into DSH conversations. The host extracts (and OCRs) them into the prompt; attach_* tools cover notebook cells, PDF-page OCR, image describe, and save.
dsh-csv-and-image-preview
jetecho/dsh-csv-and-image-preview
Preview images / SVG (and prepare CSV) in the DeepSeek Harness chat, rendered as real browser <img> elements. Preview-first workflow: show the user the asset, wait for approval, then apply the real change.
dsh-voice-input-space
xsakura666/dsh-voice-input-space
DeepSeek Harness 语音输入:长按空格说话、松开上屏,零依赖,基于 Web Speech API。
dsh-voice
navid-kianfar/dsh-voice
在 DeepSeek Harness Web 客户端里口述提示词:输入区麦克风,转写可换用托管 Whisper API、自建服务或离线 whisper.cpp。
dsh-video-understand
zeshuochen/dsh-video-understand
字幕优先的视频转录与确定性抽取式 Markdown 总结;字幕不可用时回退到 faster-whisper `large-v3`。
dsh-auto-vision
soarguo/dsh-auto-vision
为无法识图的模型把图片转成文字:由配置的视觉模型生成描述,描述以折叠上下文条目写入会话历史,原始消息保持不变。
dsh-mmx
crazyma99/dsh-mmx
MiniMax CLI(mmx)桥接:把 mmx 的网络搜索回退与图像理解接入 DSH,贴图自动识图,首次使用提供安装引导与设置卡片。
dsh-voice-mode
qishuilalala/dsh-voice-mode
DeepSeek Harness 全双工语音模式:流式识别写入草稿、按句朗读并配实时字幕、开口即打断;识别在本地推理,无需 API Key。
dsh-mimo-plugin
yyfather/dsh-mimo-plugin
把小米 MiMo 能力封装为 DSH 原生插件:联网搜索、图片/音频/视频理解、ASR 转写、TTS、音色设计与克隆,并在设置页配置 API Key。
dsh-plugin-88api-image
blackdm666/dsh-plugin-88api-image
面向 dsh 的 88API 图像工作室:四款 Image2 与 Nano Banana 模型,支持文生图、多参考图编辑与 2K/4K 输出
computer-use-vision
xuanyuanluoxue/computer-use-vision
Windows 桌面识图+模拟操作能力:截图交给视觉模型解析后模拟鼠标键盘操作,内置自进化技巧库
dsh-glm-vision
fightingfirefox/dsh-glm-vision
GLM 视觉模型插件:注册 glm-vision 供应商路由(glm-4.6v 系列)并提供 glm_vision 工具,让文本主模型直接调用智谱视觉模型看图。