본문으로 건너뛰기

비전, 음성 및 멀티모달

비전, 음성 및 멀티모달 플러그인은 DeepSeek-Harness(dsh)의 텍스트 전용 DeepSeek 모델에 눈, 귀, 그리고 목소리를 부여합니다. 붙여넣은 스크린샷을 Zhipu GLM, Gemini, Doubao 또는 로컬 Ollama로 OCR·설명하는 비전 도구와 제공자 라우트, 브라우저 Web Speech API 또는 Whisper 호환 API를 이용한 마이크 음성 입력, Edge TTS나 커스텀 음성으로 답변을 읽어 주는 TTS, 전이중 음성 모드, 이미지·영상·음악 생성기를 살펴보세요.

플러그인 86개 찾음

T

dsh-plugins (dsh-vision)

tzhr-invest/dsh-plugins

Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.

1어제비전, 음성 및 멀티모달MIT
P

dsh-screenshot

paicat1/dsh-screenshot

Standalone screen capture for DeepSeek Harness (dsh). Browser hotkeys for instant capture plus an agent-facing capture and read tool that lets the agent see and analyze any screen region.

1어제비전, 음성 및 멀티모달MIT
N

dsh-auto-vision

normanfxxkingrockwell/dsh-auto-vision

Auto-discovery vision bridge for text-only DeepSeek Harness agents: automatically finds an image-capable model from your configured providers and returns picture descriptions as plain text via a vision tool.

112시간 전비전, 음성 및 멀티모달MIT
M

dsh-koboldcpp-hands

microherox/dsh-koboldcpp-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.

117시간 전비전, 음성 및 멀티모달MIT
L

dsh-plugin-grok2api-media-tool

lsjspl/dsh-plugin-grok2api-media-tool

Gives dsh the ability to generate images and videos through the grok2api API.

111시간 전비전, 음성 및 멀티모달
L

dsh-eyes

leeminjing/dsh-eyes

On-demand vision for text-only DeepSeek models: upload images, and the model calls a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).

1그저께비전, 음성 및 멀티모달MIT
K

dsh-mindseye

kanchengw/dsh-mindseye

Vision plugin for text-only DeepSeek Harness models: native image paste, layered evidence memory and cache, and intent-driven tool selection.

15시간 전비전, 음성 및 멀티모달MIT
G

dsh-tool-vision

gloryxpnv/dsh-tool-vision

Local-first structured vision for text-only agents: images go to a local OpenAI-compatible VLM and come back as JSON evidence (summary, verbatim OCR, layout regions, entities/relations, colors, explicit uncertainty), with anti-hallucination fallback and an optional paste/upload bridge; zero cloud cost, images never leave the machine.

13일 전비전, 음성 및 멀티모달MIT
E

dsh-plugin-mm-vision

elohia/dsh-plugin-mm-vision

Synesthesia Encoder for DSH: a vision model translates images into compact structured spatial text (canvas/elements/percentage coordinates), giving text-only LLMs pixel-level image understanding via the `mm_vision` tool.

15일 전비전, 음성 및 멀티모달MIT
C

dsh-deepseek-vision

cheng-cheng9669/dsh-deepseek-vision

Reuses DeepSeek web's built-in vision mode for text-only models: the deepseek_vision tool drives the local deepseek-vision-cli browser automation (manual login helper, deep-think enabled, auto-closes browser) and returns image descriptions as text.

1그저께비전, 음성 및 멀티모달MIT
5

dsh-youreyes

54xkeee/dsh-youreyes

Vision toolkit for text-only DeepSeek: model-invokable `vision` tool, wrapper adapters for deepseek/opencode-go (v4 flash/pro), Antigravity IDE quota (default, flash/pro) / any OpenAI-compatible VLM / Gemini / local Ollama channels, evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

113시간 전비전, 음성 및 멀티모달MIT
3

dsh-vision (vision-tool)

314857493/dsh-vision

Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.

1어제비전, 음성 및 멀티모달MIT
3

dsh-vision (vision-route)

314857493/dsh-vision

Registers a `deepseek-vision` provider route: the Web GUI accepts pasted images and transcribes them to text via the free Zhipu GLM vision API before delegating to the DeepSeek adapter.

1어제비전, 음성 및 멀티모달MIT
1

dsh-vision-fallback

1helloman1/dsh-vision-fallback

Routes chat images to a fixed OpenAI-compatible vision model, returns factual observations to the selected main model, and reuses session-scoped observations across replay, compaction, and restarts.

112시간 전비전, 음성 및 멀티모달MIT
Z

dsh-voice

zhuiyueya/dsh-voice

DeepSeek Harness(dsh)용 음성 기능 — 텍스트 전용 DeepSeek을 위한 음성-텍스트 입력 + TTS 읽어주기, API 키 불필요

15일 전비전, 음성 및 멀티모달MIT
S

multimodal-bridge

spirit4471/multimodal-bridge

DeepSeek Harness 플러그인 번들: 텍스트 전용 모델을 위한 qwen_vision(Qwen-VL 이미지 이해)과 qwen_generate(Qwen-Image 텍스트-이미지 변환) 도구

15일 전비전, 음성 및 멀티모달MIT
P

dsh-yali-image-generator

pptt121212/dsh-yali-image-generator

DeepSeek-Harness 이미지 생성 플러그인. Yali AI API Key 신청: https://api.yaliai.com/

15일 전비전, 음성 및 멀티모달MIT
C

deepsee

chang416/deepsee

DeepSee: DeepSeek Harness 비전, 멀티 모델 라우팅, 전달 전 Gemini 시각 자체 검증

14일 전비전, 음성 및 멀티모달MIT
A

dsh-voice-webspeech

anweat/dsh-voice-webspeech

브라우저 Web Speech API 음성 입력: 서버 불필요, 키 불필요, 모델 다운로드 불필요(Edge=Azure, Chrome=Google 음성)

15일 전비전, 음성 및 멀티모달MIT
J

dsh-voice

jesse-njx/dsh-voice

음성으로 입력하고 음성으로 답변 받기: 받아쓰기 음성이 사용자 메시지로 변환(transcribe)되고, 에이전트가 답변을 소리 내어 읽어줌(speak), ~/.dsh/voice 아래 로컬 우선 저장

15일 전비전, 음성 및 멀티모달MIT
T

dsh-voice-live

tangzheng202202/dsh-voice-live

Real-time duplex voice over Volcengine streaming ASR/TTS: agent reply narration, barge-in, wake word, live captions, 30 Chinese voices and a reply-first acknowledgment; builds in the DSH monorepo.

0그저께비전, 음성 및 멀티모달MIT
N

dsh-voice-input

newdanew/dsh-voice-input

Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.

04일 전비전, 음성 및 멀티모달MIT
M

dsh-fish-tts

mari23333/dsh-fish-tts

Reads assistant replies aloud via Fish Audio API only (bring your own key): per-message read-aloud, auto-read toggle, and a settings page for model, voice reference_id, encrypted API key, and proxy.

0어제비전, 음성 및 멀티모달MIT
H

dsh-voice

haoku123/dsh-voice

Full-duplex voice mode for the Web UI: a composer mic (RMS endpoint detection) transcribes speech with whisper running locally in the browser, assistant replies stream back as spoken audio sentence-by-sentence, and speaking interrupts playback and the running turn (true barge-in). No API key.

0그저께비전, 음성 및 멀티모달MIT