视觉、语音与多模态
视觉、语音与多模态类插件让 DeepSeek-Harness(dsh)里纯文本的 DeepSeek 模型也能"看图""听声""开口"。包含把粘贴的截图交给智谱 GLM、Gemini、豆包或本地 Ollama 做 OCR 与图像描述的 vision 工具和供应商路由、基于浏览器 Web Speech API 或 Whisper 兼容接口的麦克风语音输入、使用 Edge TTS 或自定义音色的朗读回复与全双工语音模式,以及图片、视频、音乐生成工具。
共 312 个插件
dsh-fal-image-gen
goodandready/dsh-fal-image-gen
DeepSeek Harness 的 FAL 图像生成插件:generate_image 工具对接 FAL 队列 API,生成的图片直接渲染在对话中。
dsh-bundle-vision
skillre/dsh-bundle-vision
DeepSeek Harness 视觉 bundle + 插件:describe_image 工具读取本地图片并交由任意已配置的多模态路由解读,无需改动核心。
dsh-mindsee
123cdxcc/dsh-mindsee
DeepSeek Harness 插件:以 MindSee 为后端,为 DeepSeek 提供图片理解能力。
dsh-voice-input
opensquad-ai/dsh-voice-input
SenseVoice 语音输入:输入框旁麦克风按钮,录音转文字自动填入,本地服务自动启动
dsh-gemini-multimodal
realalexandreai/dsh-gemini-multimodal
经 Gemini API 或本地 Antigravity CLI 提供多模态工具:图像/音频/视频/文档理解与生成
dsh-tool-image-gen
zhangjunjesse/dsh-tool-image-gen
DSH 工具插件:经 ToAPIs 异步 GPT-Image-2 API 生成图像(提交、轮询、下载)
soyo
ottohere-mourn/soyo
DSH 原生视频理解,多模态供应商可配置(本地或 API)。
dsh-voice-live
tangzheng202202/dsh-voice-live
基于火山流式 ASR/TTS 的实时双工语音:回复朗读、打断、唤醒词、实时字幕、30 个中文音色与先响应后思考;在 DSH monorepo 内构建。
dsh-voice-input
newdanew/dsh-voice-input
Web UI 语音输入:输入框一键麦克风按钮,基于 Web Speech API 语音转文字填入草稿,可选识别后自动发送。
dsh-easyvision
s3yf1337/dsh-easyvision
让纯文本模型拥有视觉:describe_image 工具调用 dsh 模型列表中的视觉模型描述图片,走 harness 自身的 LLM 运行时。
dsh-subagent-vision
niuniuaba/dsh-subagent-vision
让纯文本 DeepSeek 代理在同一会话内读图:委派给视觉子代理,发送时自动把图片转为文件路径。
dsh-unsloth-hands
microherox/dsh-unsloth-hands
把重复的文本与视觉(OCR、图像分析、对比)劳动交给本地运行的 Unsloth Desktop(Unsloth Studio)服务器:unsloth_run 与 unsloth_vision 两个工具,纯 HTTP 客户端,不拉起也不持有任何进程。
dsh-koboldcpp-hands
microherox/dsh-koboldcpp-hands
给在线模型装上本地双手:koboldcpp_run 与 koboldcpp_vision 工具把重复的文本与视觉(OCR、图像分析、对比)劳动交给本地 KoboldCpp(llama.cpp)服务器,并按需拉起与管理服务器生命周期。
dsh-windows-ocr
maxwell-feng/dsh-windows-ocr
本地 Windows 引擎(Windows.Media.Ocr)识别附加图片:只把识别出的文字发送给模型,图片字节不进入上下文;视觉直通可选开启。
dsh-tesseract-ocr
maxwell-feng/dsh-tesseract-ocr
本地 Tesseract 识别附加图片:只把识别出的文字发送给模型,图片字节不进入上下文;视觉直通可选开启。
dsh-vision-plugin
ld-1101/dsh-vision-plugin
让纯文本模型也能看图:发图自动用视觉模型描述(默认提示词),描述不足时对话模型自动生成更具体的提示词重新解析;系统模型/自定义双模式 + GUI 可视化配置、Key 脱敏安全、针对 DSH 0.1.0-rc.6 的宿主小补丁(见仓库 README)。
dsh-quicksight
isanti2016/dsh-quicksight
纯文本模型双层识图:优先本地快速 OCR(RapidOCR,离线),不足时回落视觉模型(modlens)。
dsh-vision-guard
good-boy4069/dsh-vision-guard
纯文本路由的透明图片护栏:贴图不再 400 卡死会话,附带 vision_analyze 工具(OCR/PDF/docx/pptx/视频)。
dsh-plugin-image-input
elohia/dsh-plugin-image-input
图片转文字输入插件:粘贴/拖拽图片自动转为结构化文字描述发送,给纯文本 LLM 提供图片输入接管(OpenAI 兼容视觉 API)。
dsh-vision-solution
br1nosense/dsh-vision-solution
为 DSH 纯文本模型补充视觉能力:识图技能(图片理解/OCR/文档解析,竞速池→自定义通道→本地)+ 幂等宿主补丁,让图片消息能送达模型侧。
dsh-plugin-vision
tdf1995/dsh-plugin-vision
为纯文本大模型提供视觉能力:通过免费的 Gemini 与 GLM 视觉 API 实现图像描述、OCR 与视觉问答。
vision-tool
haowencang/vision-tool
DeepSeek Harness 会话式 UI/UX 视觉评审:vision_review / vision_ask 工具,对接 OpenAI 兼容多模态端点。
dsh-nanobananapro
synmindai/dsh-nanobananapro
Generate images and videos in DeepSeek Harness through the NanoBananaPro API
dsh-seedance2
synmindai/dsh-seedance2
Generate images and Seedance videos in DeepSeek Harness through the Seedance 2 AI API