본문으로 건너뛰기

비전, 음성 및 멀티모달

비전, 음성 및 멀티모달 플러그인은 DeepSeek-Harness(dsh)의 텍스트 전용 DeepSeek 모델에 눈, 귀, 그리고 목소리를 부여합니다. 붙여넣은 스크린샷을 Zhipu GLM, Gemini, Doubao 또는 로컬 Ollama로 OCR·설명하는 비전 도구와 제공자 라우트, 브라우저 Web Speech API 또는 Whisper 호환 API를 이용한 마이크 음성 입력, Edge TTS나 커스텀 음성으로 답변을 읽어 주는 TTS, 전이중 음성 모드, 이미지·영상·음악 생성기를 살펴보세요.

플러그인 312개 찾음

F

dsh-voice-announcer

flashyiyi/dsh-voice-announcer

Voice announcements when a conversation ends (session title, round, outcome) plus live sentence-by-sentence reading of replies, via built-in edge-tts with zero third-party dependencies.

03일 전비전, 음성 및 멀티모달
B

dsh-soundscape

berserk0501/dsh-soundscape

Dual-MediaPlayer daemon ($think loop + $fx one-shots) with volume ducking, mood-aware sounds (smooth/struggle/deep_think triggered by consecutive success/failure streaks and prolonged agent runtime), 15 built-in PCM WAV synthesised from sine/square waveforms, per-tool sound mapping with 51 categorised tools, custom WAV/MP3 drop-in with auto-conversion, and a full settings UI with dark-mode fix.

0지난달비전, 음성 및 멀티모달MIT
K

deepseek-plugin

klingai-dev/deepseek-plugin

Turn every idea into an image or video with Kling AI. In DeepSeek Harness, use natural language for text-to-image, image-to-image, text-to-video, image-to-video, reference-image creation, task tracking, and result previews.

0지난달비전, 음성 및 멀티모달
K

dsh-auto-vision

k2d5rqjpkg-art/dsh-auto-vision

Auto-switch the DeepSeek route to the vision model on demand: flash main session switches (A), pro keeps deep reasoning and delegates image reading to a vision subagent (B), subagents always switch, with fatal-failure fallback. No manual model switching.

0지난달비전, 음성 및 멀티모달
C

dsh-evidence

cooberped/dsh-evidence

Turns attached files into versioned evidence: `search_documents` builds a private local index (SQLite FTS5 after a startup capability probe, dependency-free JS fallback otherwise) and returns compact evidence blocks carrying an exact coordinate — PDF page, PPTX slide, text/DOCX line range, or quoted XLSX `Sheet!Range` — which `read_document` expands only while the content version still matches. Contiguous CJK runs are indexed as overlapping bigrams and queried as phrases, so word order is preserved; uploaded raster images take the native vision attachment path instead.

0지난달비전, 음성 및 멀티모달MIT
Z

dsh-tu4-inline-images

zehenk/dsh-tu4-inline-images

对话内联图片 DSH 插件 — 在 DeepSeek Harness (DSH) Web GUI 的对话中,出现本地图片路径即直接渲染为图片。 A DSH plugin that renders local image paths as inline images in DeepSeek Harness (DSH) web conversations. Security-first: loopback-only route, strong per-process token, multi-root realpath whitelist.

0지난달비전, 음성 및 멀티모달MIT
B

dsh-llm-capabilities

bamboostrip/dsh-llm-capabilities

DSH plugin: auto-detect and configure model capabilities (reasoningEfforts + input modalities) for llm-pi-ai. Successor to dsh-reasoning-efforts.

0지난달비전, 음성 및 멀티모달MIT
L

dsh-vision-toggle

lijian-ui/dsh-vision-toggle

Per-model vision (image input) toggle for DeepSeek Harness (dsh): list every configured model and flip a switch to enable/disable image support without hand-editing settings.yaml. 为 DeepSeek Harness 提供按模型的「支持图片」开关:无需手改 settings.yaml。

0지난달비전, 음성 및 멀티모달MIT
A

dsh-speech

allmodels-io/dsh-speech

Streaming speech-to-text for DeepSeek Harness using AllModels.io.

0지난달비전, 음성 및 멀티모달MIT
C

dsh-agnes

chaoliu615/dsh-agnes

Agnes AI image & video generation tools for DSH (agnes_image_generate / agnes_video_generate)

0지난달비전, 음성 및 멀티모달MIT
D

dsh-remote-deliver

demacia1314/dsh-remote-deliver

🚀 告别繁琐 SCP!远程部署 DSH 一键下载修改后的文件与图片预览交付插件

0지난달비전, 음성 및 멀티모달
W

dsh-dictation

wsl043/dsh-dictation

Editable local and desktop dictation for DeepSeek Harness

0지난달비전, 음성 및 멀티모달MIT
L

dsh-file-attach

lucasxingg/dsh-file-attach

Drag-and-drop PDF, Office, images, and text/code files into DSH conversations. The host extracts (and OCRs) them into the prompt; attach_* tools cover notebook cells, PDF-page OCR, image describe, and save.

0지난달비전, 음성 및 멀티모달MIT
J

dsh-csv-and-image-preview

jetecho/dsh-csv-and-image-preview

Preview images / SVG (and prepare CSV) in the DeepSeek Harness chat, rendered as real browser <img> elements. Preview-first workflow: show the user the asset, wait for approval, then apply the real change.

0지난달비전, 음성 및 멀티모달MIT
X

dsh-voice-input-space

xsakura666/dsh-voice-input-space

Voice input for DeepSeek Harness: hold Space to speak, release to insert. Zero dependencies, Web Speech API. / 语音输入:长按空格说话,松开上屏,零依赖。

0지난달비전, 음성 및 멀티모달MIT
N

dsh-voice

navid-kianfar/dsh-voice

Dictate prompts into the DeepSeek Harness Web Client — a microphone in the composer, with swappable transcription: hosted Whisper API, self-hosted server, or a fully offline whisper.cpp binary.

0지난달비전, 음성 및 멀티모달MIT
Z

dsh-video-understand

zeshuochen/dsh-video-understand

Subtitle-first video transcription and deterministic extractive Markdown summaries, with a faster-whisper large-v3 fallback when subtitles are unavailable.

0지난달비전, 음성 및 멀티모달
S

dsh-auto-vision

soarguo/dsh-auto-vision

Bridges images into text for DeepSeek Harness: when the session's selected model cannot see images, a configured vision model describes them and the descriptions enter the durable session history as folded context rows — your message stays untouched.

0지난달비전, 음성 및 멀티모달MIT
C

dsh-mmx

crazyma99/dsh-mmx

MiniMax CLI bridge for DeepSeek Harness (dsh): mmx-backed web search fallback and transparent image understanding, with first-run onboarding and a settings card.

0지난달비전, 음성 및 멀티모달MIT
Q

dsh-voice-mode

qishuilalala/dsh-voice-mode

Full-duplex voice mode for DeepSeek Harness: zipformer2 streaming ASR → editable draft, Edge TTS sentence-by-sentence read-aloud with live captions, true barge-in — on-device ASR, no API key. · DSH 语音双工对话:流式识别入草稿、按句朗读+实时字幕、开口即打断,识别本地推理、无需 API Key

0지난달비전, 음성 및 멀티모달MIT
Y

dsh-mimo-plugin

yyfather/dsh-mimo-plugin

MiMo (Xiaomi) tools as a DSH profile plugin: web search, image/audio/video understanding, ASR transcription, TTS, voice design and voice cloning — native agent tools with a Settings page for the API key.

0지난달비전, 음성 및 멀티모달MIT
B

dsh-plugin-88api-image

blackdm666/dsh-plugin-88api-image

88API Image Studio for DSH: four Image2 and Nano Banana models for text-to-image, multi-reference editing, 2K/4K output, and sequential batches.

0지난달비전, 음성 및 멀티모달MIT
X

computer-use-vision

xuanyuanluoxue/computer-use-vision

Windows computer-use capability for DeepSeek Harness: screenshot → vision model → simulated mouse/keyboard input, with self-evolving knowledge base.

0지난달비전, 음성 및 멀티모달
F

dsh-glm-vision

fightingfirefox/dsh-glm-vision

GLM 视觉模型插件:注册 glm-vision 供应商路由(glm-4.6v 系列,image+text 输入声明)并提供 glm_vision 工具,让 DeepSeek 等文本主模型直接调用智谱视觉模型看图。

0지난달비전, 음성 및 멀티모달MIT