Zum Hauptinhalt springen

Vision, Sprache & Multimodal

Vision, Sprache & Multimodal-Plugins verleihen den reinen Text-DeepSeek-Modellen in DeepSeek-Harness (dsh) Augen, Ohren und eine Stimme. Finden Sie Vision-Tools und Provider-Routen, die eingefügte Screenshots über Zhipu GLM, Gemini, Doubao oder lokales Ollama per OCR erfassen und beschreiben, Mikrofon-Spracheingabe über die Web Speech API des Browsers oder Whisper-kompatible APIs, Vorlesen per Edge TTS oder eigenen Stimmen, Vollduplex-Sprachmodi sowie Generatoren für Bilder, Video und Musik.

312 Plugins gefunden

T

dsh-vision-worker

try-works/dsh-vision-worker

DeepSeek Harness plugin: a vision worker over Cloudflare Workers AI (@cf/moonshotai/kimi-k2.6) that routes image requests from text-only callers, returns a versioned righthand.vision.v1 envelope, and supports follow-up questions.

0letzten MonatVision, Sprache & MultimodalApache-2.0
F

dsh-dictate

franksong2702/dsh-dictate

Browser Web Speech dictation for the Composer: recognition needs no dedicated ASR server, key, or model download; reuses Session text and a configured DSH model for contextual phrase hints and optional transcript polishing.

0letzten MonatVision, Sprache & MultimodalMIT
M

meow-vision

meimiaoji-creator/meow-vision

meow-vision 是 DeepSeek Harness 的一款视觉插件,解决纯文本模型无视觉。另一方面vue页面开发视觉验证不闭环的问题。

0letzten MonatVision, Sprache & MultimodalMIT
Z

dsh-vision-bridge

zzdream67/dsh-vision-bridge

Let text-only models see images in DeepSeek Harness: intercepts the llm/stream waterfall and transparently substitutes each image with a vision model's description.

0letzten MonatVision, Sprache & MultimodalMIT
G

glm-vision-plugin

gelomen/glm-vision-plugin

DSH Web 插件:调用智谱 GLM 视觉模型分析图片(analyze_image 工具 + 插件设置卡片)。

0letzten MonatVision, Sprache & Multimodal
S

dsh-vision-link

sprainjinyu/dsh-vision-link

Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.

0letzten MonatVision, Sprache & Multimodal
R

dsh-omni-vision

renji004/dsh-omni-vision

Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.

0vor 2 MonatenVision, Sprache & MultimodalMIT
I

dsh-tool-accurate-vision

imkingjh999/dsh-tool-accurate-vision

Model-facing accurate_vision tool: precise image spatial reasoning via a vision model

0vor 2 MonatenVision, Sprache & MultimodalMIT
V

dsh-ocr-bridge

vuvanmai936-dot/dsh-ocr-bridge

Paste images into DeepSeek Harness chat and have them read by a free local backend (macOS Vision / Tesseract) before the text-only DeepSeek model answers

0letzten MonatVision, Sprache & Multimodal
P

dsh-plugin-image-tools

pasumao/dsh-plugin-image-tools

DSH 图片插件:ask_user_choice 图片/图文混合选项(Web GUI 渲染图片选择卡,可放大查看)+ show_images 在回复中内嵌图片(图片与文字混排)。图片来源支持本地路径 / http(s) URL / base64 data URI。纯插件实现,不改核心包。

0letzten MonatVision, Sprache & MultimodalMIT
H

win11-oneocr

hawkhai/win11-oneocr

Local Windows 11 OneOCR tool for DSH: `oneocr_recognize` returns OCR text and structured line/word polygons, confidence, rotation, and handwriting style.

0letzten MonatVision, Sprache & Multimodal
H

wechat-ocr

hawkhai/wechat-ocr

Local WeChat OCR tool for DSH: `wechat_ocr_recognize` returns recognized text and the engine structured result for a local image path.

0letzten MonatVision, Sprache & Multimodal
Z

dsh-vision

zoahdev/dsh-vision

vision_analyze tool: analyze a local image or URL with an OpenAI-compatible vision model.

0vor 2 MonatenVision, Sprache & MultimodalMIT
V

dsh-doubao-voice

vorpal-poem/dsh-doubao-voice

DeepSeek Harness 语音输入插件:火山引擎流式 ASR(豆包 Seed ASR),流式回填输入框。Voice input for DSH via Volcengine streaming ASR.

0vor 2 MonatenVision, Sprache & MultimodalMIT
T

dsh-vision-api-localorweb

tipsong/dsh-vision-api-localorweb

0vor 2 MonatenVision, Sprache & MultimodalMIT
J

dsh-autovision

junkrat9527/dsh-autovision

Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.

0vor 2 MonatenVision, Sprache & MultimodalMIT
L

dsh-mingmu

lab-sku/dsh-mingmu

明眸 VisionBridge - Selbstentwickelte visuelle Brücke: Wenn ein blindes Modell Bilder empfängt, wird automatisch ein visuelles Modell zur Erkennung aufgerufen und der erkannte Text an das Hauptmodell zurückgegeben; nahtlos, konfigurierbar, verlustfrei bei Upgrades.

0vor 2 MonatenVision, Sprache & MultimodalMIT
M

dsh-voice

motongv/dsh-voice

给 DeepSeek Harness(DSH / DeepSeek Hermes)加语音能力的社区插件:输入框语音输入(可配快捷键)+ 回答朗读(微软 Edge 神经网络音色,可换音色、可试听),无需 API Key。

0vor 2 MonatenVision, Sprache & MultimodalMIT
T

dsh-vision-ocr

timeflies-qyh/dsh-vision-ocr

DeepSeek Harness OCR plugin — offline image text recognition powered by PaddleOCR-json (primary) and RapidOCR-json (fallback). 让 DeepSeek Harness 直接识别图片中的文字,无需视觉模型、完全本地离线运行。

0vor 2 MonatenVision, Sprache & MultimodalMIT
S

dsh-media-guard

spyfree/dsh-media-guard

Deterministic aggregate media budgets and safe request projections for DeepSeek Harness (DSH)

0vor 2 MonatenVision, Sprache & MultimodalMIT
S

dsh-siliconflow-vision

shixiangyu2/dsh-siliconflow-vision

DSH 插件:通过硅基流动(SiliconFlow)视觉大模型识别/分析图片,支持本地文件路径、http(s) 图片 URL 与 base64 data URL。含持久化的粘贴识别面板(web 端)。

0vor 2 MonatenVision, Sprache & MultimodalMIT
A

dsh-plugin-multimodal-bridge

avaritiachaos/dsh-plugin-multimodal-bridge

Dynamic multimodal-to-text projection and cross-model vision bridge for DeepSeek Harness (dsh), allowing seamless hot-switching between vision models (Gemini/Claude) and text-only models (DeepSeek).

0vor 2 MonatenVision, Sprache & MultimodalMIT
1

intelligenteyes

1210350468/intelligenteyes

IntelligentEyes universal agent vision gateway for DeepSeek Harness

0vor 2 MonatenVision, Sprache & MultimodalMIT
B

vision-translation

bingl-li/vision-translation

Native dsh (DeepSeek Harness) Cordis plugin adapter for vision-translation: grounds images into <vision-context> via the Python CLI (PROTOCOL v1). Spawns cli.py, never re-implements core logic.

0vor 2 MonatenVision, Sprache & MultimodalMIT