본문으로 건너뛰기

비전, 음성 및 멀티모달

비전, 음성 및 멀티모달 플러그인은 DeepSeek-Harness(dsh)의 텍스트 전용 DeepSeek 모델에 눈, 귀, 그리고 목소리를 부여합니다. 붙여넣은 스크린샷을 Zhipu GLM, Gemini, Doubao 또는 로컬 Ollama로 OCR·설명하는 비전 도구와 제공자 라우트, 브라우저 Web Speech API 또는 Whisper 호환 API를 이용한 마이크 음성 입력, Edge TTS나 커스텀 음성으로 답변을 읽어 주는 TTS, 전이중 음성 모드, 이미지·영상·음악 생성기를 살펴보세요.

플러그인 312개 찾음

Z

dsh-agnes-media

zlforward/dsh-agnes-media

123일 전비전, 음성 및 멀티모달MIT
F

dsh-funasr-voice

fenglin-ai/dsh-funasr-voice

Offline voice input for the DSH Web UI: mic to local FunASR (SenseVoiceSmall), one-click install, no cloud.

1지난달비전, 음성 및 멀티모달MIT
I

dshtools-sensevoice-input

ilovedyou6666-hub/dshtools-sensevoice-input

基于 SenseVoiceSmall(iic/SenseVoiceSmall)多语言语音理解模型的 DSH Desktop 本地语音输入插件。

129일 전비전, 음성 및 멀티모달MIT
Z

dsh-file-convert

zzy-12345678/dsh-file-convert

Local-first file conversion: 26 conversions across images, PDF (with OCR and experimental PDF→DOCX), data, audio/video and office docs; 7 tools, all local, no API keys.

1지난달비전, 음성 및 멀티모달
D

dsh-voice-input-npm

difimim/dsh-voice-input-npm

语音输入插件 for Deepseek Harness

1지난달비전, 음성 및 멀티모달MIT
J

dsh-live-voice

jstn-1g/dsh-live-voice

Consent-bound one-turn voice preview for DSH Web with a credential-free local synthetic demo, optional Qwen Audio, exact Session isolation, and explicit transcript-to-draft handoff without automatic submission.

127일 전비전, 음성 및 멀티모달MIT
Y

dsh-video-gen

yang-wudi/dsh-video-gen

Text-to-video and image-to-video generation via DashScope Wanx, Volcengine Seedance, Google Veo and OpenAI Sora, with tool results saved to the session workspace and a video gallery (grid, lightbox, download, delete) that survives DSH restarts.

1지난달비전, 음성 및 멀티모달
W

dsh-image-viewer

wsl043/dsh-image-viewer

Upgrades DSH image viewing with pointer-centered zoom, pan, galleries, downloads, keyboard support, and inline region notes.

1그저께비전, 음성 및 멀티모달MIT
J

dsh-freecanvas

justinqiuck/dsh-freecanvas

Install DSH FreeCanvas as a bundled DeepSeek Harness app with split layouts and managed local Agent connectivity. · 将 DSH FreeCanvas 作为内置应用安装到 DSH,支持分屏布局与本地 Agent 自动连接。

1지난달비전, 음성 및 멀티모달MIT
M

dsh-bilibili

moxingovo/dsh-bilibili

DeepSeek Harness plugin: Bilibili keyword video search, video metadata, subtitle transcripts, direct play URLs, and multimodal frame viewing (bilibili_search / bilibili_video / bilibili_subtitles / bilibili_playurl / bilibili_frames). Anonymous by default

1지난달비전, 음성 및 멀티모달MIT
L

dsh-pianist

laplace-bit/dsh-pianist

Piano performance plugin: ask the agent to play a piece and it renders on a Canvas2D grand piano with real Salamander Grand samples, an immersive stage, and an interactive 88-key keyboard.

123일 전비전, 음성 및 멀티모달MIT
A

dsh-image-preview

algerkong/dsh-image-preview

Image preview for DSH (DeepSeek Harness) web sessions: read_image results render as a thumbnail, click for full size in the built-in lightbox.

1지난달비전, 음성 및 멀티모달
J

dsh-mmroute

jmxsxwyzjdwl/dsh-mmroute

Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.

126일 전비전, 음성 및 멀티모달MIT
C

aura-vision

ck-epsilon/aura-vision

Free vision OCR with adaptive tile recognition for long documents and Markdown/Word/PNG/Excel export.

1지난달비전, 음성 및 멀티모달MIT
T

taxue-dsh-artisan

taxueseek/taxue-dsh-artisan

Integrated visual creation toolchain for DSH: prompt reverse-engineering and audit optimization plus multi-provider image generation with async background rendering.

1지난달비전, 음성 및 멀티모달MIT
H

dsh-vision-analysis

harvey-will/dsh-vision-analysis

DeepSeek Harness vision plugin: 8 analysis modes (describe, OCR, chart data, UI review, object detection, compare, code-gen, debug), any OpenAI- or Anthropic-compatible vision API, with a built-in free vision model and automatic rate-limit failover.

13일 전비전, 음성 및 멀티모달MIT
C

remotion-video-plugin

chenjie1129/remotion-video-plugin

Remotion video creation and verified rendering plugin for DeepSeek Harness

1지난달비전, 음성 및 멀티모달MIT
A

dsh-screenshot

alain-prot0s5/dsh-screenshot

Screenshot-to-input for DeepSeek Harness: composer camera button + global hotkey (Alt+A) + listener bound to the app lifecycle, configurable in settings. 截图自动粘贴到 DSH 输入框:相机按钮 + 全局快捷键 + 生命周期绑定 + 设置页配置。

1지난달비전, 음성 및 멀티모달MIT
A

dsh-composer-image-tools

ai-yucheng/dsh-composer-image-tools

聊天输入框图片工具(自研):上传图片(≤10MB 防烧 token) + 自定义区域截图(Electron desktopCapturer),注入 DSH 草稿图片轨。零外部依赖。

1지난달비전, 음성 및 멀티모달MIT
L

dsh-speech-input

liznee/dsh-speech-input

A microphone button for DeepSeek Harness that writes browser speech recognition into the composer draft.

1지난달비전, 음성 및 멀티모달MIT
S

dsh-voice-input-cn

schumchanvi/dsh-voice-input-cn

컴포저용 중국 대응 음성 입력입니다. 로컬 Python 브리지가 필요합니다 (pip install dashscope websockets, bridge/voice-bridge.py 실행) — 플러그인만으로는 작동하지 않습니다. 브라우저 마이크가 16kHz PCM을 브리지로 스트리밍하고, Alibaba Cloud DashScope ASR(paraformer-realtime-v2)를 실행합니다. 중간 결과가 커서 위치의 초안을 채우며, 무음 시 자동 정지, 선택적 자동 전송을 지원합니다.

116일 전비전, 음성 및 멀티모달MIT
A

dsh-client-vision (tool-vision)

ankye/dsh-client-vision

Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.

123일 전비전, 음성 및 멀티모달MIT
S

dsh-narrate

stuarthu/dsh-narrate

DeepSeek Harness (dsh) plugin: turn one idea into a narrated video cut from your own asset folder, stopping four times to ask you first.

1지난달비전, 음성 및 멀티모달MIT
A

dsh-audio-copilot

ai-yucheng/dsh-audio-copilot

Audio Copilot for DeepSeek Harness: transcribe audio (ASR) and synthesize speech (TTS) — gives text-only agents ears and a voice. Windows-local SAPI TTS out of the box; OpenAI-compatible ASR/TTS endpoints configurable. Includes an in-composer voice-input

1지난달비전, 음성 및 멀티모달MIT