跳过主要内容

插件

浏览、筛选并安装 DeepSeek-Harness 插件。

共 14 个插件

W

dsh-ears

wiziscool/dsh-ears

面向 DeepSeek Harness (dsh) 的语音输入插件:输入框的麦克风按钮把语音转成草稿文本,支持多种语音识别后端,可选经 dsh 自有 LLM 路由润色,并带原生设置页。

21前天视觉、语音与多模态MIT
C

dsh-bilibili

czx2244/dsh-bilibili

B站视频分析工具:提取元数据、字幕文稿(必剪/本地 ASR 兜底)、评论与弹幕,抓取清晰关键帧并可选本地视觉描述。

92个月前工具与能力MIT
I

dsh-video-understand

ilps2/dsh-video-understand

低成本视频理解工具:video_understand 把 B站链接/BV号/本地视频转成 AVIS 信息层(ASR+场景结构+对象轨迹+YOLO 标签)并输出摘要+问答。问题驱动动态路由分层(L0 ASR / L1 对象轨迹 / L2 关键帧 VLM)、语义层复用(重复提问直接查层)、单次问题预算上限。Python 引擎:核心层需 faster-whisper / opencv / yt-dlp(约 200-300MB);可选语义层另需约 2GB 的 torch / transformers / ultralytics。内置 doctor --fix 一键建 venv 并装齐两者。

8上个月工具与能力
G

dsh-voice

goodandready/dsh-voice

Web UI 语音输入:按停顿分段的听写与语音消息,各自拥有独立的服务商回退链(Deepgram、Groq、HuggingFace、本地 whisper.cpp 或 OpenAI 兼容接口)。

8前天视觉、语音与多模态MIT
S

dsh-voice

stardustlc666/dsh-voice

语音五工具:edge-tts 免费微软神经语音合成、OpenAI 兼容 ASR 转写、音色清单、批量音色试听与健康自检。

4前天视觉、语音与多模态MIT
B

dsh-stt-input

baisama-cloud/dsh-stt-input

Web UI 语音输入:输入框旁麦克风按钮语音转文字填入输入框;支持浏览器 Web Speech API 本地识别(零配置、无需密钥)与 OpenAI 兼容 Whisper API(OpenAI/Groq),模型与语言可在设置中选择。

3上个月视觉、语音与多模态MIT
Z

dsh-watch-video

zeshuochen/dsh-watch-video

字幕优先的视频转录,支持 SRT 导出与可取消任务控制;字幕不可用时回退到 faster-whisper `large-v3`。

2上个月视觉、语音与多模态
1

dsh-wsl-im

173787247/dsh-wsl-im

把飞书、企微、钉钉、QQ、Slack、Discord、Telegram、Mattermost 桥进 dsh agent(出站长连接/长轮询/Webhook),提供 im_status、可选本机 Whisper ASR、QQ/钉钉/TG/MM 纯文本出站,以及环境变量白名单(DSH_IM_*_ALLOWED_USER_IDS)。allowedUserIds 为空即任何人都能驱动 agent——对外暴露前请先设白名单。

117小时前开发与插件工具MIT
V

dsh-live-voice

victorwads/dsh-live-voice

本地优先的 DSH 语音对话:本地语音识别与合成、可选外部供应商,支持连续对话与自动朗读

115天前视觉、语音与多模态GPL-3.0
Z

dsh-voice

zhuiyueya/dsh-voice

Voice for DeepSeek Harness(dsh) — speech-to-text input + read-aloud TTS for text-only DeepSeek, zero API key.

12个月前视觉、语音与多模态MIT
B

dsh-asr-voice

bittersmilezzz/dsh-asr-voice

开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。

015天前视觉、语音与多模态MIT
K

dsh-kitt-voice

kittcat-lab/dsh-kitt-voice

Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.

025天前视觉、语音与多模态MIT
N

dsh-voice

navid-kianfar/dsh-voice

Dictate prompts into the DeepSeek Harness Web Client — a microphone in the composer, with swappable transcription: hosted Whisper API, self-hosted server, or a fully offline whisper.cpp binary.

0上个月视觉、语音与多模态MIT
Z

dsh-video-understand

zeshuochen/dsh-video-understand

字幕优先的视频转录与确定性抽取式 Markdown 总结;字幕不可用时回退到 faster-whisper `large-v3`。

0上个月视觉、语音与多模态