跳过主要内容

视觉、语音与多模态

视觉、语音与多模态类插件让 DeepSeek-Harness(dsh)里纯文本的 DeepSeek 模型也能"看图""听声""开口"。包含把粘贴的截图交给智谱 GLM、Gemini、豆包或本地 Ollama 做 OCR 与图像描述的 vision 工具和供应商路由、基于浏览器 Web Speech API 或 Whisper 兼容接口的麦克风语音输入、使用 Edge TTS 或自定义音色的朗读回复与全双工语音模式,以及图片、视频、音乐生成工具。

共 312 个插件

Y

dsh-plugin-vision-toolkit

yytbit/dsh-plugin-vision-toolkit

DeepSeek Harness 视觉工具包:glance/ground/detect/crop 等 CLI 工具,让纯文本 Agent 也能理解图像。

12个月前视觉、语音与多模态MIT
P

dsh-yali-image-generator

pptt121212/dsh-yali-image-generator

DeepSeek-Harness 图像生成插件。申请 Yali AI API Key:https://api.yaliai.com/

12个月前视觉、语音与多模态MIT
C

deepsee

chang416/deepsee

DeepSee: DeepSeek Harness vision, multi-model routing, and Gemini visual self-checks before delivery

12个月前视觉、语音与多模态MIT
F

dsh-sight

fu3rte/dsh-sight

为纯文本 DeepSeek Harness 模型补上视觉:vision 工具内置免费 VLM 预设,支持多图批量分析、粘贴图片与热重载设置页。

12个月前视觉、语音与多模态MIT
W

dsh-friend

wanghehe123/dsh-friend

DeepSeek Harness 的人格化伴侣插件:角色卡、语音对话、Live2D 舞台、本地记忆与工作陪伴。

12个月前视觉、语音与多模态MIT
A

dsh-voice-webspeech

anweat/dsh-voice-webspeech

浏览器 Web Speech API 语音输入:零服务端、零密钥、零模型下载(Edge=Azure 语音、Chrome=Google 语音)。

1上个月视觉、语音与多模态MIT
R

dsh-plugins

retiredphysicist/dsh-plugins

DeepSeek Harness plugin: web browsing tools (markdown / screenshot / pdf / crawl) powered by Cloudflare Browser Run — real headless Chrome with JS rendering, login sessions and WebMCP support.

0前天视觉、语音与多模态
W

dsh-kite

weibaohui/dsh-kite

Animated kite that responds to agent activity, with configurable kite frames, patterns and colors, plus user-supplied image textures.

04天前视觉、语音与多模态
K

dsh-zcode-cli-proxy

kyle123740/dsh-zcode-cli-proxy

ZCode CLI app-server relay for DeepSeek Harness, with GLM streaming, image input and session reuse using an existing Start Plan login; requires the ZCode client.

03天前视觉、语音与多模态AGPL-3.0
E

dsh-fal-imagegen

enchanted0911/dsh-fal-imagegen

fal.ai native image generation for DSH: a FAL_KEY settings card that follows DSH's language (zh/en, English default) plus Agent tools (fal_generate_image async / fal_edit_image / fal_get_image_task / fal_list_image_models) that call queue.fal.run directly

09天前视觉、语音与多模态MIT
Y

dsh-video-to-notes

yll-kb/dsh-video-to-notes

Opt-in DeepSeek Harness skill bundle that turns course, lecture, tutorial, documentary, meeting, and talk videos into structured study notes.

012天前视觉、语音与多模态MIT
W

dsh-voice-danmaku

weizhida/dsh-voice-danmaku

这是dsh的插件,语音发送弹幕。在玩游戏时通过语音输入在b站发弹幕,不切出游戏可以正常操作

04天前视觉、语音与多模态MIT
T

dsh-xiaozhi

toddpan/dsh-xiaozhi

把小智语音助手接入 DSH Web:DSH 作为 MCP 工具提供方,把 35 个 DSH Web 接口封装成 16 个语音友好工具,覆盖工作区、会话、对话、模型、设置与文件;默认出站 WebSocket 连到小智 MCP 接入点(无需公网 IP 和端口转发),可同时绑定多台设备,并自带实时刷新状态的 DSH 设置页。

06天前视觉、语音与多模态
J

sh-volume-knob

jianghu-lao-yao/sh-volume-knob

输入框话筒旁的扬声器按钮——单击翻到「你最新提问的开头」并闪烁光标,从那里读到最新回复结尾(走 dsh-tts,回退浏览器语音);按住右滑可在页面上挑选任意朗读起点,按住上滑调出竖式混音台,分别控制页内媒体音量与系统输出音量。

08天前视觉、语音与多模态MIT
M

lookover (dsh-look)

mengxiaoxian/lookover

macOS 场景感知探针:拉模型零轮询的 look 工具,读取最前非自身窗口(应用信息、AX 标题/选中文本、按需本地 OCR),支持召唤热键按下瞬间抓取快照注入对话。

011天前视觉、语音与多模态MIT
Y

dsh-agnes-gen

ylhow06/dsh-agnes-gen

Agnes AI image/video generation tools (agnes_image / agnes_video) for DSH (DeepSeek Harness). Built-in cross-process RPM rate limiting, 429 backoff and local ffmpeg GIF conversion.

011天前视觉、语音与多模态MIT
G

dsh-plugin-notify

goodandready/dsh-plugin-notify

DSH plugin: audio chimes, cross-session toasts, desktop push, and IM webhooks for turn completion, errors, and approvals.

011天前视觉、语音与多模态MIT
F

dsh-say

fangqian616/dsh-say

让你的 DSH 用你喜欢的声音开口说话、汇报内容,过长的汇报先压缩再念。角色声线不需要自己训练:一段 3-10 秒参考音就能克隆,也能直接用别人训练好的模型;不必为此装一整套 6.4 GB 的 GPT-SoVITS —— 运行时和声线都由插件自己装好。已有 GPT-SoVITS 的话也能直接接上,两者跑的是同一套声线。

012天前视觉、语音与多模态
C

dsh-read-aloud

cccc12138/dsh-read-aloud

在每条已定稿的助手回复里、紧挨点赞按钮右侧加一个小喇叭:点击用浏览器语音引擎朗读这条回复,鼠标悬停可调倍速与音色。

015天前视觉、语音与多模态MIT
L

dsh-voice-alert

loyalchiiina/dsh-voice-alert

每个 turn 结束自动语音/音效提醒,出错时播失败提示;内置 20 个音效(提醒 10 + 大自然 10)且默认「音效」模式,零配置开箱即用、不需要任何 API Key 与音频文件;进阶可用火山「声音复刻」克隆自己的音色并一键生成三条播报语音;走 winmm/waveOut 原样播放,绝不改动系统音量或静音状态。仅 Windows。

04天前视觉、语音与多模态MIT
D

dsh-screen-eye

davidekingsss/dsh-screen-eye

macOS 与 Windows 上的 agent 截屏:一个工具截屏并把图片本身返回给模型,无需再调第二次。Windows 上整段连拍在同一个引擎进程内完成;macOS 上常驻助手把区域截图降到 13ms、变化检测降到 23ms。另有一个仅 macOS 的工具:报告屏幕录制授权状态,并直接打开用于修复的系统设置面板。

014天前视觉、语音与多模态MIT
N

vision-exp-tile

nicholaskin/vision-exp-tile

Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.

022天前视觉、语音与多模态MIT
Z

dsh-stt-plugin

zemanzhang809/dsh-stt-plugin

A speech-to-text (voice input) plugin for DeepSeek Harness.

015天前视觉、语音与多模态MIT
B

dsh-asr-voice

bittersmilezzz/dsh-asr-voice

开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。

018天前视觉、语音与多模态MIT