본문으로 건너뛰기

비전, 음성 및 멀티모달

비전, 음성 및 멀티모달 플러그인은 DeepSeek-Harness(dsh)의 텍스트 전용 DeepSeek 모델에 눈, 귀, 그리고 목소리를 부여합니다. 붙여넣은 스크린샷을 Zhipu GLM, Gemini, Doubao 또는 로컬 Ollama로 OCR·설명하는 비전 도구와 제공자 라우트, 브라우저 Web Speech API 또는 Whisper 호환 API를 이용한 마이크 음성 입력, Edge TTS나 커스텀 음성으로 답변을 읽어 주는 TTS, 전이중 음성 모드, 이미지·영상·음악 생성기를 살펴보세요.

플러그인 312개 찾음

T

dsh-vision-worker

try-works/dsh-vision-worker

DeepSeek Harness plugin: a vision worker over Cloudflare Workers AI (@cf/moonshotai/kimi-k2.6) that routes image requests from text-only callers, returns a versioned righthand.vision.v1 envelope, and supports follow-up questions.

0지난달비전, 음성 및 멀티모달Apache-2.0
F

dsh-dictate

franksong2702/dsh-dictate

Browser Web Speech dictation for the Composer: recognition needs no dedicated ASR server, key, or model download; reuses Session text and a configured DSH model for contextual phrase hints and optional transcript polishing.

0지난달비전, 음성 및 멀티모달MIT
M

meow-vision

meimiaoji-creator/meow-vision

meow-vision 是 DeepSeek Harness 的一款视觉插件,解决纯文本模型无视觉。另一方面vue页面开发视觉验证不闭环的问题。

0지난달비전, 음성 및 멀티모달MIT
Z

dsh-vision-bridge

zzdream67/dsh-vision-bridge

Let text-only models see images in DeepSeek Harness: intercepts the llm/stream waterfall and transparently substitutes each image with a vision model's description.

0지난달비전, 음성 및 멀티모달MIT
G

glm-vision-plugin

gelomen/glm-vision-plugin

DSH Web 插件:调用智谱 GLM 视觉模型分析图片(analyze_image 工具 + 插件设置卡片)。

0지난달비전, 음성 및 멀티모달
S

dsh-vision-link

sprainjinyu/dsh-vision-link

Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.

0지난달비전, 음성 및 멀티모달
R

dsh-omni-vision

renji004/dsh-omni-vision

Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.

02개월 전비전, 음성 및 멀티모달MIT
I

dsh-tool-accurate-vision

imkingjh999/dsh-tool-accurate-vision

Model-facing accurate_vision tool: precise image spatial reasoning via a vision model

02개월 전비전, 음성 및 멀티모달MIT
V

dsh-ocr-bridge

vuvanmai936-dot/dsh-ocr-bridge

Paste images into DeepSeek Harness chat and have them read by a free local backend (macOS Vision / Tesseract) before the text-only DeepSeek model answers

0지난달비전, 음성 및 멀티모달
P

dsh-plugin-image-tools

pasumao/dsh-plugin-image-tools

DSH 图片插件:ask_user_choice 图片/图文混合选项(Web GUI 渲染图片选择卡,可放大查看)+ show_images 在回复中内嵌图片(图片与文字混排)。图片来源支持本地路径 / http(s) URL / base64 data URI。纯插件实现,不改核心包。

0지난달비전, 음성 및 멀티모달MIT
H

win11-oneocr

hawkhai/win11-oneocr

Local Windows 11 OneOCR tool for DSH: `oneocr_recognize` returns OCR text and structured line/word polygons, confidence, rotation, and handwriting style.

0지난달비전, 음성 및 멀티모달
H

wechat-ocr

hawkhai/wechat-ocr

Local WeChat OCR tool for DSH: `wechat_ocr_recognize` returns recognized text and the engine structured result for a local image path.

0지난달비전, 음성 및 멀티모달
Z

dsh-vision

zoahdev/dsh-vision

vision_analyze tool: analyze a local image or URL with an OpenAI-compatible vision model.

02개월 전비전, 음성 및 멀티모달MIT
V

dsh-doubao-voice

vorpal-poem/dsh-doubao-voice

DeepSeek Harness 语音输入插件:火山引擎流式 ASR(豆包 Seed ASR),流式回填输入框。Voice input for DSH via Volcengine streaming ASR.

02개월 전비전, 음성 및 멀티모달MIT
T

dsh-vision-api-localorweb

tipsong/dsh-vision-api-localorweb

02개월 전비전, 음성 및 멀티모달MIT
J

dsh-autovision

junkrat9527/dsh-autovision

Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.

02개월 전비전, 음성 및 멀티모달MIT
L

dsh-mingmu

lab-sku/dsh-mingmu

明眸 VisionBridge - 자체 개발 시각 브리지: 블라인드 모델이 이미지를 받을 때 자동으로 시각 모델을 호출하여 인식하고, 인식된 텍스트를 메인 모델에 다시 공급하며, 사용자 인지 없이, 구성 가능, 업그레이드 시 손실 없음

02개월 전비전, 음성 및 멀티모달MIT
M

dsh-voice

motongv/dsh-voice

给 DeepSeek Harness(DSH / DeepSeek Hermes)加语音能力的社区插件:输入框语音输入(可配快捷键)+ 回答朗读(微软 Edge 神经网络音色,可换音色、可试听),无需 API Key。

02개월 전비전, 음성 및 멀티모달MIT
T

dsh-vision-ocr

timeflies-qyh/dsh-vision-ocr

DeepSeek Harness OCR plugin — offline image text recognition powered by PaddleOCR-json (primary) and RapidOCR-json (fallback). 让 DeepSeek Harness 直接识别图片中的文字,无需视觉模型、完全本地离线运行。

02개월 전비전, 음성 및 멀티모달MIT
S

dsh-media-guard

spyfree/dsh-media-guard

Deterministic aggregate media budgets and safe request projections for DeepSeek Harness (DSH)

02개월 전비전, 음성 및 멀티모달MIT
S

dsh-siliconflow-vision

shixiangyu2/dsh-siliconflow-vision

DSH 插件:通过硅基流动(SiliconFlow)视觉大模型识别/分析图片,支持本地文件路径、http(s) 图片 URL 与 base64 data URL。含持久化的粘贴识别面板(web 端)。

02개월 전비전, 음성 및 멀티모달MIT
A

dsh-plugin-multimodal-bridge

avaritiachaos/dsh-plugin-multimodal-bridge

Dynamic multimodal-to-text projection and cross-model vision bridge for DeepSeek Harness (dsh), allowing seamless hot-switching between vision models (Gemini/Claude) and text-only models (DeepSeek).

02개월 전비전, 음성 및 멀티모달MIT
1

intelligenteyes

1210350468/intelligenteyes

IntelligentEyes universal agent vision gateway for DeepSeek Harness

02개월 전비전, 음성 및 멀티모달MIT
B

vision-translation

bingl-li/vision-translation

Native dsh (DeepSeek Harness) Cordis plugin adapter for vision-translation: grounds images into <vision-context> via the Python CLI (PROTOCOL v1). Spawns cli.py, never re-implements core logic.

02개월 전비전, 음성 및 멀티모달MIT