- Accueil
- Catégories
- Vision, voix et multimodal
Vision, voix et multimodal
Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.
312 plugins trouvés
dsh-vision-worker
try-works/dsh-vision-worker
DeepSeek Harness plugin: a vision worker over Cloudflare Workers AI (@cf/moonshotai/kimi-k2.6) that routes image requests from text-only callers, returns a versioned righthand.vision.v1 envelope, and supports follow-up questions.
dsh-dictate
franksong2702/dsh-dictate
Browser Web Speech dictation for the Composer: recognition needs no dedicated ASR server, key, or model download; reuses Session text and a configured DSH model for contextual phrase hints and optional transcript polishing.
meow-vision
meimiaoji-creator/meow-vision
meow-vision 是 DeepSeek Harness 的一款视觉插件,解决纯文本模型无视觉。另一方面vue页面开发视觉验证不闭环的问题。
dsh-vision-bridge
zzdream67/dsh-vision-bridge
Let text-only models see images in DeepSeek Harness: intercepts the llm/stream waterfall and transparently substitutes each image with a vision model's description.
glm-vision-plugin
gelomen/glm-vision-plugin
DSH Web 插件:调用智谱 GLM 视觉模型分析图片(analyze_image 工具 + 插件设置卡片)。
dsh-vision-link
sprainjinyu/dsh-vision-link
Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.
dsh-omni-vision
renji004/dsh-omni-vision
Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.
dsh-tool-accurate-vision
imkingjh999/dsh-tool-accurate-vision
Model-facing accurate_vision tool: precise image spatial reasoning via a vision model
dsh-ocr-bridge
vuvanmai936-dot/dsh-ocr-bridge
Paste images into DeepSeek Harness chat and have them read by a free local backend (macOS Vision / Tesseract) before the text-only DeepSeek model answers
dsh-plugin-image-tools
pasumao/dsh-plugin-image-tools
DSH 图片插件:ask_user_choice 图片/图文混合选项(Web GUI 渲染图片选择卡,可放大查看)+ show_images 在回复中内嵌图片(图片与文字混排)。图片来源支持本地路径 / http(s) URL / base64 data URI。纯插件实现,不改核心包。
win11-oneocr
hawkhai/win11-oneocr
Local Windows 11 OneOCR tool for DSH: `oneocr_recognize` returns OCR text and structured line/word polygons, confidence, rotation, and handwriting style.
wechat-ocr
hawkhai/wechat-ocr
Local WeChat OCR tool for DSH: `wechat_ocr_recognize` returns recognized text and the engine structured result for a local image path.
dsh-vision
zoahdev/dsh-vision
vision_analyze tool: analyze a local image or URL with an OpenAI-compatible vision model.
dsh-doubao-voice
vorpal-poem/dsh-doubao-voice
DeepSeek Harness 语音输入插件:火山引擎流式 ASR(豆包 Seed ASR),流式回填输入框。Voice input for DSH via Volcengine streaming ASR.
dsh-vision-api-localorweb
tipsong/dsh-vision-api-localorweb
dsh-autovision
junkrat9527/dsh-autovision
Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.
dsh-mingmu
lab-sku/dsh-mingmu
明眸 VisionBridge - Pont visuel développé en interne : lorsqu'un modèle aveugle reçoit une image, il appelle automatiquement un modèle visuel pour la reconnaître et renvoie le texte reconnu au modèle principal ; transparent, configurable, sans perte lors des mises à jour.
dsh-voice
motongv/dsh-voice
给 DeepSeek Harness(DSH / DeepSeek Hermes)加语音能力的社区插件:输入框语音输入(可配快捷键)+ 回答朗读(微软 Edge 神经网络音色,可换音色、可试听),无需 API Key。
dsh-vision-ocr
timeflies-qyh/dsh-vision-ocr
DeepSeek Harness OCR plugin — offline image text recognition powered by PaddleOCR-json (primary) and RapidOCR-json (fallback). 让 DeepSeek Harness 直接识别图片中的文字,无需视觉模型、完全本地离线运行。
dsh-media-guard
spyfree/dsh-media-guard
Deterministic aggregate media budgets and safe request projections for DeepSeek Harness (DSH)
dsh-siliconflow-vision
shixiangyu2/dsh-siliconflow-vision
DSH 插件:通过硅基流动(SiliconFlow)视觉大模型识别/分析图片,支持本地文件路径、http(s) 图片 URL 与 base64 data URL。含持久化的粘贴识别面板(web 端)。
dsh-plugin-multimodal-bridge
avaritiachaos/dsh-plugin-multimodal-bridge
Dynamic multimodal-to-text projection and cross-model vision bridge for DeepSeek Harness (dsh), allowing seamless hot-switching between vision models (Gemini/Claude) and text-only models (DeepSeek).
intelligenteyes
1210350468/intelligenteyes
IntelligentEyes universal agent vision gateway for DeepSeek Harness
vision-translation
bingl-li/vision-translation
Native dsh (DeepSeek Harness) Cordis plugin adapter for vision-translation: grounds images into <vision-context> via the Python CLI (PROTOCOL v1). Spawns cli.py, never re-implements core logic.