Passer au contenu principal

Vision, voix et multimodal

Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.

86 plugins trouvés

D

dsh-web-ui (dsh-tool-describe-image)

damonkoy/dsh-web-ui

Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.

3il y a 3 joursVision, voix et multimodalApache-2.0
G

dsh-vision-bridge

gxx182/dsh-vision-bridge

Plugin DeepSeek Harness qui relie les images de session à des API de vision interchangeables, tout en gardant DeepSeek comme modèle principal.

3il y a 12 heuresVision, voix et multimodalMIT
E

dsh-llm-vision-bridge

einskyle/dsh-llm-vision-bridge

Passerelle de vision native au fournisseur LLM : les images collées dans le chat sont décrites par un modèle de vision (Qwen3-VL via pi-ai/llama.cpp), et la description textuelle est transmise à DeepSeek texte seul pour la réponse — l'admission, le routage et la compaction des images passent tous par les mécanismes natifs du harness, avec un cache de descriptions LRU et une reprise sur erreur 503.

3il y a 4 joursVision, voix et multimodalMIT
X

dsh-vision-bridge

ximengxiaolan/dsh-vision-bridge

Les images jointes au composeur sont transcrites en texte par un modèle de vision compatible OpenAI avant d'atteindre les modèles DeepSeek texte seul.

3il y a 4 joursVision, voix et multimodalMIT
Q

dsh-mic-input

qt-chen/dsh-mic-input

Saisie vocale au micro pour le composeur : transcription en direct via la Web Speech API du navigateur, déduplication/poursuite automatique, ponctuation intelligente, réglages de langue et d'envoi automatique.

3avant-hierVision, voix et multimodalMIT
S

dsh-voice

stardustlc666/dsh-voice

Voice pair: free edge-tts neural speech synthesis + OpenAI-compatible ASR transcription.

2il y a 7 heuresVision, voix et multimodalMIT
N

voco-input-sh

nothree-code/voco-input-sh

Voice input for the Web UI: a mic button that drives local VocoType offline speech recognition and auto-inserts recognized text into the composer (auto-deploy, dedupe, continuous dictation).

2il y a 21 heuresVision, voix et multimodalMIT
0

dsh-voice-input

0nt-one/dsh-voice-input

Mic button in the composer tool row: Web Speech API speech-to-text (Chrome/Edge), language switching, and optional auto-send, zero dependencies.

2il y a 3 joursVision, voix et multimodalMIT
X

dsh-draw-router

xiaozhe7772222/dsh-draw-router

Universal image generation for DeepSeek Harness: auto-discovers image models from any OpenAI-compatible endpoint (SenseNova, StepFun, Agnes, Qwen, Flux, SD, Imagen and more), with agent tools and REST API.

2il y a 5 heuresVision, voix et multimodalMIT
X

dsh-vision

xiaoshihou514/dsh-vision

Native vision capability extension, using either Zhipu (free) or Qwen-VL (local).

2avant-hierVision, voix et multimodalMIT
W

visual-review

wang-bool/visual-review

Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.

2il y a 17 heuresVision, voix et multimodalMIT
N

dsh-llm-deepseek-vision

nagasakisoyo-ui/dsh-llm-deepseek-vision

Vision-augmented DeepSeek adapter: a vision-capable model describes image input, then a text-only DeepSeek model reasons over the description.

2il y a 4 joursVision, voix et multimodalMIT
M

dsh-vision-tools

moon09300731/dsh-vision-tools

Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.

2hierVision, voix et multimodalMIT
M

dsh-tesseract-ocr

maxwell-feng/dsh-tesseract-ocr

Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

2hierVision, voix et multimodalMIT
L

dsh-image-gen

leemancheung/dsh-image-gen

GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.

2il y a 3 joursVision, voix et multimodalMIT
H

dsh-open-eyes

hyp6666/dsh-open-eyes

Vision bridge for text-only DeepSeek routes that analyzes attached and local images through configurable OpenAI Responses, Chat Completions, or Anthropic Messages endpoints while leaving image-capable routes native.

2avant-hierVision, voix et multimodalMIT
H

dsh-vision-mix

haiziyao/dsh-vision-mix

Combine text, vision, and image-generation APIs into one Mix model with automatic routing: text-only requests go to the chat model, user images and agent screenshots go to the vision model, follow-ups keep using the same session image, and agents can generate or edit images with session-scoped call history.

2il y a 3 joursVision, voix et multimodalMIT
F

dsh-free-vision

fuzzysoul/dsh-free-vision

Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.

2il y a 8 heuresVision, voix et multimodalMIT
P

dsh-voice-call

pandapolo/dsh-voice-call

Agent-initiated voice calls: `offer_call` rings the human (接听/拒接/稍后再说); accepted calls synthesize and play locally via CrispASR + Qwen3-TTS (9 speakers, 2 Chinese dialects), rejected calls return the decision to the agent.

1hierVision, voix et multimodalMIT
F

dsh-chatvoice

fuzzysoul/dsh-chatvoice

Free voice closed loop for the Web UI: browser SpeechRecognition mic input with live interim results plus read-aloud speaker buttons and auto-read for assistant replies, zero configuration and no API key.

1il y a 3 joursVision, voix et multimodalMIT
B

dsh-stt-input

baisama-cloud/dsh-stt-input

Speech-to-text voice input for the web UI: a mic button in the composer transcribes speech into the draft via the browser Web Speech API (zero-config) or an OpenAI-compatible Whisper API (OpenAI / Groq), with a selectable model and language in Settings.

1hierVision, voix et multimodalMIT
X

dsh-image-vision

xsoc1/dsh-image-vision

Chat image-attachment bridge with a `view_image` tool for any OpenAI-compatible VLM (local Ollama or cloud): pasted/dropped images become `view_image` path markers before reaching text-only DeepSeek models.

1il y a 3 joursVision, voix et multimodal
W

mimo-vision

wulusai2333/mimo-vision

`describe_image` tool: a vision bridge that sends images to mimo-v2.5 through the opencode Zen API (credential `OPENCODE_GO_API_KEY`, free route first with paid fallback) and returns text descriptions for text-only models, with native passthrough and ImageMagick transcoding of SVG/TIFF/HEIC formats.

1avant-hierVision, voix et multimodalMIT
W

dsh-tool-vision

wanshichenguang/dsh-tool-vision

Adds an image_describe tool backed by the DashScope OpenAI-compatible vision API; on text-only sessions, pasted images are stored as local paths for the model and rendered inline in the chat transcript.

1avant-hierVision, voix et multimodalMIT