Passer au contenu principal

Vision, voix et multimodal

Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.

86 plugins trouvés

L

modlens

liustack/modlens

Pont de vision pour les modèles texte seul : collez une image, obtenez des preuves JSON structurées (OCR, mise en page, sémantique).

3.1kil y a 4 heuresVision, voix et multimodalMIT
Y

dsh-vision-router

ysr666/dsh-vision-router

Vision gratuite pour les agents texte seul : chaîne de vision intégrée sans clé, plus des outils pixel (questions-réponses, ancrage, recadrage, diff pixel, couleurs, OCR, tracé SVG, détourage, captures d'écran) ; collez une image pour l'utiliser.

740il y a 4 heuresVision, voix et multimodalMIT
A

dsh-vision-toolkit

anionex/dsh-vision-toolkit

Tâches de vision pour les modèles texte seul : questions-réponses sur image sensibles à l'intention, OCR de captures d'écran longues, reproduction d'UI, ancrage (grounding) et diff pixel.

694il y a 7 heuresVision, voix et multimodalMIT
J

picturereader

jing-hy/picturereader

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

20il y a 9 heuresVision, voix et multimodalMIT
L

dsh-vision

linenxi-ctrl/dsh-vision

Plugin de vision externe pour DeepSeek Harness : panneau de configuration via le bouton baleine, reconnaissance d'image avec réponse automatique, et outils de capture d'écran/reconnaissance pour l'agent.

12il y a 3 joursVision, voix et multimodalMIT
M

dsh-media-skills

mjorgin/dsh-media-skills

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

11hierVision, voix et multimodalMIT
F

dsh-vision-proxy

flyvhidbwo/dsh-vision-proxy

Cerveau DeepSeek + transcription automatique d'images : joignez des images dans l'interface, chacune est transcrite en texte via n'importe quel VLM compatible OpenAI avant d'atteindre DeepSeek, qui ne traite que du texte — un chemin rapide avec clé (par défaut qwen3.7-flash ; DashScope/Zhipu/OpenRouter ou tout endpoint compatible OpenAI) avec votre propre clé, ou Ollama en local auto-détecté sans configuration.

11avant-hierVision, voix et multimodalMIT
P

dsh-vision-opencode

poiuyjie/dsh-vision-opencode

Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.

10il y a 13 heuresVision, voix et multimodalMIT
J

dsh-visual-plugin

jyh20030112/dsh-visual-plugin

Dsh-visual-plugin. Donnez des yeux à votre modèle texte seul : transmettez les images de l'utilisateur à n'importe quel modèle de vision compatible OpenAI et consultez les résultats dans un panneau droit de la Web UI.

9il y a 13 heuresVision, voix et multimodalMIT
C

Gemini-Eyes

consolesun/gemini-eyes

MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.

8il y a 4 joursVision, voix et multimodal
1

dsh-plugin-tts

1624318455/dsh-plugin-tts

Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.

7il y a 9 heuresVision, voix et multimodalMIT
D

dsh-imagegen

dickpy/dsh-imagegen

AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.

7il y a 6 heuresVision, voix et multimodalApache-2.0
C

dsh-chat-imagine

corrinehu/dsh-chat-imagine

Automatically generates and displays images in the DSH chat via API channels or local CLIs (supports mmx / codex / agy).

7hierVision, voix et multimodalMIT
5

dsh-vision

54xkeee/dsh-vision

Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

7il y a 3 joursVision, voix et multimodalMIT
Z

dsh-voice-input-plugin

zhangbo-cn/dsh-voice-input-plugin

Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.

6hierVision, voix et multimodalMIT
S

dsh-deepseek-vision

siegfly/dsh-deepseek-vision

A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.

6hierVision, voix et multimodalMIT
M

dsh-windows-ocr

maxwell-feng/dsh-windows-ocr

Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

6hierVision, voix et multimodalMIT
3

dsh-voice

3274375092/dsh-voice

Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.

4il y a 3 joursVision, voix et multimodalMIT
G

deepseek-vision (dsh-plugin-deepseek-vision)

gou-gee/deepseek-vision

Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.

4avant-hierVision, voix et multimodalMIT
A

dsh-guide-dog

atropinoltt/dsh-guide-dog

MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.

4hierVision, voix et multimodalMIT
H

dsh-her-eyes

huashenglian/dsh-her-eyes

Un plugin dsh qui permet à l'IA d'invoquer automatiquement des VLM (modèles multimodaux) pour l'analyse visuelle.

4il y a 5 joursVision, voix et multimodalMIT
A

dsh-speak

alan2z/dsh-speak

Voice-announce the final reply on Windows (SAPI5 natural voices) and macOS (system voice); skips reasoning and tool calls, one-line npm install.

3avant-hierVision, voix et multimodalMIT
S

dsh-plugin-multimodal

shinjiyu/dsh-plugin-multimodal

Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.

3avant-hierVision, voix et multimodalMIT
N

free-vision-skill

niyongsheng/free-vision-skill

Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.

3il y a 3 joursVision, voix et multimodalMIT