Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

T

dsh-vision-worker

try-works/dsh-vision-worker

DeepSeek Harness plugin: a vision worker over Cloudflare Workers AI (@cf/moonshotai/kimi-k2.6) that routes image requests from text-only callers, returns a versioned righthand.vision.v1 envelope, and supports follow-up questions.

0last monthVision, Voice & MultimodalApache-2.0
F

dsh-dictate

franksong2702/dsh-dictate

Browser Web Speech dictation for the Composer: recognition needs no dedicated ASR server, key, or model download; reuses Session text and a configured DSH model for contextual phrase hints and optional transcript polishing.

0last monthVision, Voice & MultimodalMIT
M

meow-vision

meimiaoji-creator/meow-vision

meow-vision 是 DeepSeek Harness 的一款视觉插件,解决纯文本模型无视觉。另一方面vue页面开发视觉验证不闭环的问题。

0last monthVision, Voice & MultimodalMIT
Z

dsh-vision-bridge

zzdream67/dsh-vision-bridge

Let text-only models see images in DeepSeek Harness: intercepts the llm/stream waterfall and transparently substitutes each image with a vision model's description.

0last monthVision, Voice & MultimodalMIT
G

glm-vision-plugin

gelomen/glm-vision-plugin

DSH Web 插件:调用智谱 GLM 视觉模型分析图片(analyze_image 工具 + 插件设置卡片)。

0last monthVision, Voice & Multimodal
S

dsh-vision-link

sprainjinyu/dsh-vision-link

Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.

0last monthVision, Voice & Multimodal
R

dsh-omni-vision

renji004/dsh-omni-vision

Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.

02 months agoVision, Voice & MultimodalMIT
I

dsh-tool-accurate-vision

imkingjh999/dsh-tool-accurate-vision

Model-facing accurate_vision tool: precise image spatial reasoning via a vision model

02 months agoVision, Voice & MultimodalMIT
V

dsh-ocr-bridge

vuvanmai936-dot/dsh-ocr-bridge

Paste images into DeepSeek Harness chat and have them read by a free local backend (macOS Vision / Tesseract) before the text-only DeepSeek model answers

0last monthVision, Voice & Multimodal
P

dsh-plugin-image-tools

pasumao/dsh-plugin-image-tools

DSH 图片插件:ask_user_choice 图片/图文混合选项(Web GUI 渲染图片选择卡,可放大查看)+ show_images 在回复中内嵌图片(图片与文字混排)。图片来源支持本地路径 / http(s) URL / base64 data URI。纯插件实现,不改核心包。

0last monthVision, Voice & MultimodalMIT
H

win11-oneocr

hawkhai/win11-oneocr

Local Windows 11 OneOCR tool for DSH: `oneocr_recognize` returns OCR text and structured line/word polygons, confidence, rotation, and handwriting style.

0last monthVision, Voice & Multimodal
H

wechat-ocr

hawkhai/wechat-ocr

Local WeChat OCR tool for DSH: `wechat_ocr_recognize` returns recognized text and the engine structured result for a local image path.

0last monthVision, Voice & Multimodal
Z

dsh-vision

zoahdev/dsh-vision

vision_analyze tool: analyze a local image or URL with an OpenAI-compatible vision model.

02 months agoVision, Voice & MultimodalMIT
V

dsh-doubao-voice

vorpal-poem/dsh-doubao-voice

DeepSeek Harness 语音输入插件:火山引擎流式 ASR(豆包 Seed ASR),流式回填输入框。Voice input for DSH via Volcengine streaming ASR.

02 months agoVision, Voice & MultimodalMIT
T

dsh-vision-api-localorweb

tipsong/dsh-vision-api-localorweb

02 months agoVision, Voice & MultimodalMIT
J

dsh-autovision

junkrat9527/dsh-autovision

Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.

02 months agoVision, Voice & MultimodalMIT
L

dsh-mingmu

lab-sku/dsh-mingmu

明眸 VisionBridge - 自研视觉桥:瞎子模型收图时自动调用视觉模型识别,把识别文字喂回主模型,无感、可配置、升级不丢

02 months agoVision, Voice & MultimodalMIT
M

dsh-voice

motongv/dsh-voice

给 DeepSeek Harness(DSH / DeepSeek Hermes)加语音能力的社区插件:输入框语音输入(可配快捷键)+ 回答朗读(微软 Edge 神经网络音色,可换音色、可试听),无需 API Key。

02 months agoVision, Voice & MultimodalMIT
T

dsh-vision-ocr

timeflies-qyh/dsh-vision-ocr

DeepSeek Harness OCR plugin — offline image text recognition powered by PaddleOCR-json (primary) and RapidOCR-json (fallback). 让 DeepSeek Harness 直接识别图片中的文字,无需视觉模型、完全本地离线运行。

02 months agoVision, Voice & MultimodalMIT
S

dsh-media-guard

spyfree/dsh-media-guard

Deterministic aggregate media budgets and safe request projections for DeepSeek Harness (DSH)

02 months agoVision, Voice & MultimodalMIT
S

dsh-siliconflow-vision

shixiangyu2/dsh-siliconflow-vision

DSH 插件:通过硅基流动(SiliconFlow)视觉大模型识别/分析图片,支持本地文件路径、http(s) 图片 URL 与 base64 data URL。含持久化的粘贴识别面板(web 端)。

02 months agoVision, Voice & MultimodalMIT
A

dsh-plugin-multimodal-bridge

avaritiachaos/dsh-plugin-multimodal-bridge

Dynamic multimodal-to-text projection and cross-model vision bridge for DeepSeek Harness (dsh), allowing seamless hot-switching between vision models (Gemini/Claude) and text-only models (DeepSeek).

02 months agoVision, Voice & MultimodalMIT
1

intelligenteyes

1210350468/intelligenteyes

IntelligentEyes universal agent vision gateway for DeepSeek Harness

02 months agoVision, Voice & MultimodalMIT
B

vision-translation

bingl-li/vision-translation

Native dsh (DeepSeek Harness) Cordis plugin adapter for vision-translation: grounds images into <vision-context> via the Python CLI (PROTOCOL v1). Spawns cli.py, never re-implements core logic.

02 months agoVision, Voice & MultimodalMIT