Passer au contenu principal

Vision, voix et multimodal

Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.

312 plugins trouvés

Z

dsh-tier-router

zhangzhangco/dsh-tier-router

Automatic tier-based model routing for DeepSeek Harness: a virtual `smart` model classifies each request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured. 三级难度 + 视觉自动路由,虚拟 smart 模型零配置接入。

0il y a 17 joursVision, voix et multimodal
R

dsh-voice-input

rio-promax/dsh-voice-input

DSH voice input plugin: browser realtime + VAD dictation + local/cloud ASR + DeepSeek AI polish

0il y a 21 joursVision, voix et multimodalMIT
C

dsh-audio-visualizer

caesarjue/dsh-audio-visualizer

System-audio driven UI visualizer for DeepSeek Harness: a draggable 48-band spectrum chip plus a frame-wide bass glow (memory-only FFT).

0il y a 21 joursVision, voix et multimodalMIT
Z

dsh-plugin-image-picker

zg2017/dsh-plugin-image-picker

Adds a file-picker 'attach image' button to the composer toolbar, since DSH's own web client only supports paste and drag-and-drop for image attachments (v1 scope, per its own upstream design notes) - a real gap for touch devices with no drag-and-drop and

0le mois dernierVision, voix et multimodalMIT
N

dsh-desktop

new-256/dsh-desktop

DSH (DeepSeek Harness) ComfyUI 桥接插件:让 Agent 直接驱动本地/局域网 ComfyUI 生图生视频 — 8 个全局模型工具、预设工作流模板(txt2img/img2img/Wan/SVD/H3)、任意 API 工作流逃生舱、设备能力守卫、模型注册表断点续传下载、ComfyUI 一键安装评估。ComfyUI bridge plugin for DeepSeek Harness: image & video generation tools driving a local

0il y a 23 joursVision, voix et multimodalMIT
S

dsh-ui-tool-result-images

suntianc/dsh-ui-tool-result-images

DeepSeek Harness Web plugin that keeps image-bearing Tool results visible after Compact transcript folding

0il y a 23 joursVision, voix et multimodalMIT
C

dsh-voice-mimo

ch1bug/dsh-voice-mimo

Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).

0il y a 19 joursVision, voix et multimodalMIT
Z

dsh-image-router

zhiwuli0228/dsh-image-router

Digests the images in a prompt with a vision model before admission, so any model — a text-only one included — can read them without the session ever switching models, and adds a describe_image tool for image paths.

0il y a 21 joursVision, voix et multimodalMIT
M

dsh-image-guard

mafeis/dsh-image-guard

Trims historical images in outgoing chat requests down to a recent-image count, learns the provider per-prompt image cap from HTTP 400 responses, and retries with fewer images so image-heavy sessions keep working.

0il y a 19 joursVision, voix et multimodalMIT
J

dsh-mathmatic-symbol

jaxzhou/dsh-mathmatic-symbol

Three tools for DeepSeek Harness: typeset LaTeX formulas into images, draw mathematical figures from a declarative spec, and convert a formula or SVG into an image ready to embed in a document.

0il y a 21 joursVision, voix et multimodalMIT
J

dsh-voice-input-qwen-asr

jsoncode/dsh-voice-input-qwen-asr

Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run ser

0il y a 28 joursVision, voix et multimodal
2

dsh-tts-flash

2021heei/dsh-tts-flash

Reads AI replies aloud as they stream, with voiced waiting phrases while the model thinks. Edge TTS built in, any OpenAI-compatible cloud engine supported.

0il y a 23 joursVision, voix et multimodalMIT
W

dsh-image-generation

whites18/dsh-image-generation

Configure multiple image providers in Settings and call image_generate with the one selected model; images save under generate/image and show inline in the conversation.

0il y a 5 joursVision, voix et multimodalMIT
A

dsh-reelsmaker

aayan-cloud/dsh-reelsmaker

DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.

0il y a 28 joursVision, voix et multimodal
M

dsh-voice-input-en

mohith-das/dsh-voice-input-en

Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow

0le mois dernierVision, voix et multimodalMIT
W

dsh-plugins (dsh-ding-sound)

wwweljf/dsh-plugins

Turn-end notification sound for DSH: built-in internet meme voices (ni gan ma ai yo, ji ni tai mei, shen ying ge, etc.), Settings panel with preview/switch/random, custom audio folder support.

0le mois dernierVision, voix et multimodalMIT
K

dsh-kitt-voice

kittcat-lab/dsh-kitt-voice

Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.

0il y a 28 joursVision, voix et multimodalMIT
J

dsh-plugin-show-image

justhalfbit/dsh-plugin-show-image

DeepSeek Harness (DSH) 会话内图片渲染插件:全局 show_image 工具 + 点击放大 lightbox。 | Inline image rendering plugin for DSH: global show_image tool + click-to-enlarge lightbox.

0le mois dernierVision, voix et multimodalMIT
A

dsh-plugins

aetheri-ai/dsh-plugins

DeepSeek Harness plugin: a model-facing show_image tool that presents a local image to the human viewer in the conversation.

0le mois dernierVision, voix et multimodalMIT
S

dsh-voice-control

sucriss/dsh-voice-control

Voice control for the DSH web composer: push-to-talk speech-to-text into the input box (with optional auto-send), spoken playback of assistant replies via the Web Speech API, right-click settings popover, and a global Ctrl+M hotkey.

0le mois dernierVision, voix et multimodalMIT
S

dsh-vision-pro-bridge

shainedemo/dsh-vision-pro-bridge

Vision bridge for text-only DeepSeek models: transcribes attached images with deepseek-v4-flash-vision-exp before they reach deepseek-v4-pro, with no third-party dependencies.

0le mois dernierVision, voix et multimodalMIT
H

dsh-maclens

harzva/dsh-maclens

Apple on-device Vision tools for text-only dsh models: local OCR (zh-Hans + 30 langs), image classification, face detection, document layout, and a combined describe — 100% offline, no API key, tall-screenshot slicing.

0le mois dernierVision, voix et multimodalMIT
A

dsh-image-generation (tool-image-generation)

ankye/dsh-image-generation

Model-facing image-generation tool with a configurable channel and normalized image parameters.

0le mois dernierVision, voix et multimodalMIT
A

dsh-vision-bridge

alaxrpg/dsh-vision-bridge

Adds image input and recognition through configured DSH providers or an OpenAI-compatible endpoint.

0il y a 5 joursVision, voix et multimodalMIT