본문으로 건너뛰기

비전, 음성 및 멀티모달

비전, 음성 및 멀티모달 플러그인은 DeepSeek-Harness(dsh)의 텍스트 전용 DeepSeek 모델에 눈, 귀, 그리고 목소리를 부여합니다. 붙여넣은 스크린샷을 Zhipu GLM, Gemini, Doubao 또는 로컬 Ollama로 OCR·설명하는 비전 도구와 제공자 라우트, 브라우저 Web Speech API 또는 Whisper 호환 API를 이용한 마이크 음성 입력, Edge TTS나 커스텀 음성으로 답변을 읽어 주는 TTS, 전이중 음성 모드, 이미지·영상·음악 생성기를 살펴보세요.

플러그인 86개 찾음

L

modlens

liustack/modlens

텍스트 전용 모델을 위한 비전 브리지: 이미지를 붙여넣으면 구조화된 JSON 증거(OCR, 레이아웃, 의미 분석)를 반환합니다.

3.1k2시간 전비전, 음성 및 멀티모달MIT
Y

dsh-vision-router

ysr666/dsh-vision-router

텍스트 전용 에이전트를 위한 무료 비전 기능: 키 없이 사용 가능한 내장 비전 체인과 픽셀 도구(Q&A, 그라운딩, 자르기, 픽셀 비교, 색상, OCR, SVG 트레이스, 컷아웃, 스크린샷)를 제공하며, 이미지를 붙여넣기만 하면 됩니다.

7402시간 전비전, 음성 및 멀티모달MIT
A

dsh-vision-toolkit

anionex/dsh-vision-toolkit

텍스트 전용 모델을 위한 비전 작업 모음: 의도 인식 이미지 Q&A, 긴 스크린샷 OCR, UI 재현, 그라운딩, 픽셀 비교를 제공합니다.

6945시간 전비전, 음성 및 멀티모달MIT
J

picturereader

jing-hy/picturereader

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

207시간 전비전, 음성 및 멀티모달MIT
L

dsh-vision

linenxi-ctrl/dsh-vision

DeepSeek Harness용 외부 비전 플러그인: 고래 버튼 설정 패널, 자동 답변이 포함된 이미지 인식, 에이전트 스크린샷/인식 도구를 제공합니다.

123일 전비전, 음성 및 멀티모달MIT
M

dsh-media-skills

mjorgin/dsh-media-skills

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

11어제비전, 음성 및 멀티모달MIT
F

dsh-vision-proxy

flyvhidbwo/dsh-vision-proxy

DeepSeek 두뇌 + 자동 이미지 텍스트 변환: GUI에서 이미지를 첨부하면 텍스트 전용 DeepSeek에 전달되기 전에 OpenAI 호환 VLM이 각 이미지를 텍스트로 변환합니다 — 자체 키를 사용하는 키 기반 고속 경로(기본값 qwen3.7-flash; DashScope/Zhipu/OpenRouter 또는 OpenAI 호환 엔드포인트라면 모두 가능)나, 설정 없이 자동 감지되는 로컬 Ollama를 사용할 수 있습니다.

11그저께비전, 음성 및 멀티모달MIT
P

dsh-vision-opencode

poiuyjie/dsh-vision-opencode

Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.

1011시간 전비전, 음성 및 멀티모달MIT
J

dsh-visual-plugin

jyh20030112/dsh-visual-plugin

dsh-visual-plugin: 텍스트 전용 모델에 눈을 달아줍니다. 사용자 이미지를 OpenAI 호환 비전 모델로 전달하고 결과를 Web UI 우측 패널에서 확인할 수 있습니다.

912시간 전비전, 음성 및 멀티모달MIT
C

Gemini-Eyes

consolesun/gemini-eyes

MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.

84일 전비전, 음성 및 멀티모달
1

dsh-plugin-tts

1624318455/dsh-plugin-tts

Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.

77시간 전비전, 음성 및 멀티모달MIT
D

dsh-imagegen

dickpy/dsh-imagegen

AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.

74시간 전비전, 음성 및 멀티모달Apache-2.0
C

dsh-chat-imagine

corrinehu/dsh-chat-imagine

Automatically generates and displays images in the DSH chat via API channels or local CLIs (supports mmx / codex / agy).

7어제비전, 음성 및 멀티모달MIT
5

dsh-vision

54xkeee/dsh-vision

Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

7그저께비전, 음성 및 멀티모달MIT
Z

dsh-voice-input-plugin

zhangbo-cn/dsh-voice-input-plugin

Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.

6어제비전, 음성 및 멀티모달MIT
S

dsh-deepseek-vision

siegfly/dsh-deepseek-vision

A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.

6어제비전, 음성 및 멀티모달MIT
M

dsh-windows-ocr

maxwell-feng/dsh-windows-ocr

Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

6어제비전, 음성 및 멀티모달MIT
3

dsh-voice

3274375092/dsh-voice

Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.

43일 전비전, 음성 및 멀티모달MIT
G

deepseek-vision (dsh-plugin-deepseek-vision)

gou-gee/deepseek-vision

Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.

4그저께비전, 음성 및 멀티모달MIT
A

dsh-guide-dog

atropinoltt/dsh-guide-dog

MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.

4어제비전, 음성 및 멀티모달MIT
H

dsh-her-eyes

huashenglian/dsh-her-eyes

AI가 VLM(멀티모달 모델)을 자동으로 호출해 시각 분석을 수행할 수 있게 해주는 dsh 플러그인입니다.

45일 전비전, 음성 및 멀티모달MIT
A

dsh-speak

alan2z/dsh-speak

Voice-announce the final reply on Windows (SAPI5 natural voices) and macOS (system voice); skips reasoning and tool calls, one-line npm install.

3그저께비전, 음성 및 멀티모달MIT
S

dsh-plugin-multimodal

shinjiyu/dsh-plugin-multimodal

Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.

3그저께비전, 음성 및 멀티모달MIT
N

free-vision-skill

niyongsheng/free-vision-skill

Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.

33일 전비전, 음성 및 멀티모달MIT