Vai al contenuto principale

Visione, voce e multimodale

I plugin Visione, voce e multimodale danno occhi, orecchie e voce ai modelli DeepSeek solo testo di DeepSeek-Harness (dsh). Trova strumenti di visione e route di provider che eseguono OCR e descrivono gli screenshot incollati tramite Zhipu GLM, Gemini, Doubao o Ollama in locale, input vocale dal microfono tramite la Web Speech API del browser o API compatibili con Whisper, lettura ad alta voce con Edge TTS o voci personalizzate, modalità vocali full-duplex, e generatori di immagini, video e musica.

312 plugin trovati

T

dsh-vision-worker

try-works/dsh-vision-worker

DeepSeek Harness plugin: a vision worker over Cloudflare Workers AI (@cf/moonshotai/kimi-k2.6) that routes image requests from text-only callers, returns a versioned righthand.vision.v1 envelope, and supports follow-up questions.

0mese scorsoVisione, voce e multimodaleApache-2.0
F

dsh-dictate

franksong2702/dsh-dictate

Browser Web Speech dictation for the Composer: recognition needs no dedicated ASR server, key, or model download; reuses Session text and a configured DSH model for contextual phrase hints and optional transcript polishing.

0mese scorsoVisione, voce e multimodaleMIT
M

meow-vision

meimiaoji-creator/meow-vision

meow-vision 是 DeepSeek Harness 的一款视觉插件,解决纯文本模型无视觉。另一方面vue页面开发视觉验证不闭环的问题。

0mese scorsoVisione, voce e multimodaleMIT
Z

dsh-vision-bridge

zzdream67/dsh-vision-bridge

Let text-only models see images in DeepSeek Harness: intercepts the llm/stream waterfall and transparently substitutes each image with a vision model's description.

0mese scorsoVisione, voce e multimodaleMIT
G

glm-vision-plugin

gelomen/glm-vision-plugin

DSH Web 插件:调用智谱 GLM 视觉模型分析图片(analyze_image 工具 + 插件设置卡片)。

0mese scorsoVisione, voce e multimodale
S

dsh-vision-link

sprainjinyu/dsh-vision-link

Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.

0mese scorsoVisione, voce e multimodale
R

dsh-omni-vision

renji004/dsh-omni-vision

Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.

02 mesi faVisione, voce e multimodaleMIT
I

dsh-tool-accurate-vision

imkingjh999/dsh-tool-accurate-vision

Model-facing accurate_vision tool: precise image spatial reasoning via a vision model

02 mesi faVisione, voce e multimodaleMIT
V

dsh-ocr-bridge

vuvanmai936-dot/dsh-ocr-bridge

Paste images into DeepSeek Harness chat and have them read by a free local backend (macOS Vision / Tesseract) before the text-only DeepSeek model answers

0mese scorsoVisione, voce e multimodale
P

dsh-plugin-image-tools

pasumao/dsh-plugin-image-tools

DSH 图片插件:ask_user_choice 图片/图文混合选项(Web GUI 渲染图片选择卡,可放大查看)+ show_images 在回复中内嵌图片(图片与文字混排)。图片来源支持本地路径 / http(s) URL / base64 data URI。纯插件实现,不改核心包。

0mese scorsoVisione, voce e multimodaleMIT
H

win11-oneocr

hawkhai/win11-oneocr

Local Windows 11 OneOCR tool for DSH: `oneocr_recognize` returns OCR text and structured line/word polygons, confidence, rotation, and handwriting style.

0mese scorsoVisione, voce e multimodale
H

wechat-ocr

hawkhai/wechat-ocr

Local WeChat OCR tool for DSH: `wechat_ocr_recognize` returns recognized text and the engine structured result for a local image path.

0mese scorsoVisione, voce e multimodale
Z

dsh-vision

zoahdev/dsh-vision

vision_analyze tool: analyze a local image or URL with an OpenAI-compatible vision model.

02 mesi faVisione, voce e multimodaleMIT
V

dsh-doubao-voice

vorpal-poem/dsh-doubao-voice

DeepSeek Harness 语音输入插件:火山引擎流式 ASR(豆包 Seed ASR),流式回填输入框。Voice input for DSH via Volcengine streaming ASR.

02 mesi faVisione, voce e multimodaleMIT
T

dsh-vision-api-localorweb

tipsong/dsh-vision-api-localorweb

02 mesi faVisione, voce e multimodaleMIT
J

dsh-autovision

junkrat9527/dsh-autovision

Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.

02 mesi faVisione, voce e multimodaleMIT
L

dsh-mingmu

lab-sku/dsh-mingmu

明眸 VisionBridge - Ponte visivo autoprogettato: quando un modello cieco riceve immagini, richiama automaticamente un modello visivo per riconoscerle e restituisce il testo riconosciuto al modello principale; trasparente, configurabile e senza perdite durante gli aggiornamenti.

02 mesi faVisione, voce e multimodaleMIT
M

dsh-voice

motongv/dsh-voice

给 DeepSeek Harness(DSH / DeepSeek Hermes)加语音能力的社区插件:输入框语音输入(可配快捷键)+ 回答朗读(微软 Edge 神经网络音色,可换音色、可试听),无需 API Key。

02 mesi faVisione, voce e multimodaleMIT
T

dsh-vision-ocr

timeflies-qyh/dsh-vision-ocr

DeepSeek Harness OCR plugin — offline image text recognition powered by PaddleOCR-json (primary) and RapidOCR-json (fallback). 让 DeepSeek Harness 直接识别图片中的文字,无需视觉模型、完全本地离线运行。

02 mesi faVisione, voce e multimodaleMIT
S

dsh-media-guard

spyfree/dsh-media-guard

Deterministic aggregate media budgets and safe request projections for DeepSeek Harness (DSH)

02 mesi faVisione, voce e multimodaleMIT
S

dsh-siliconflow-vision

shixiangyu2/dsh-siliconflow-vision

DSH 插件:通过硅基流动(SiliconFlow)视觉大模型识别/分析图片,支持本地文件路径、http(s) 图片 URL 与 base64 data URL。含持久化的粘贴识别面板(web 端)。

02 mesi faVisione, voce e multimodaleMIT
A

dsh-plugin-multimodal-bridge

avaritiachaos/dsh-plugin-multimodal-bridge

Dynamic multimodal-to-text projection and cross-model vision bridge for DeepSeek Harness (dsh), allowing seamless hot-switching between vision models (Gemini/Claude) and text-only models (DeepSeek).

02 mesi faVisione, voce e multimodaleMIT
1

intelligenteyes

1210350468/intelligenteyes

IntelligentEyes universal agent vision gateway for DeepSeek Harness

02 mesi faVisione, voce e multimodaleMIT
B

vision-translation

bingl-li/vision-translation

Native dsh (DeepSeek Harness) Cordis plugin adapter for vision-translation: grounds images into <vision-context> via the Python CLI (PROTOCOL v1). Spawns cli.py, never re-implements core logic.

02 mesi faVisione, voce e multimodaleMIT