- Accueil
- Catégories
- Vision, voix et multimodal
Vision, voix et multimodal
Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.
312 plugins trouvés
dsh-tier-router
zhangzhangco/dsh-tier-router
Automatic tier-based model routing for DeepSeek Harness: a virtual `smart` model classifies each request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured. 三级难度 + 视觉自动路由,虚拟 smart 模型零配置接入。
dsh-voice-input
rio-promax/dsh-voice-input
DSH voice input plugin: browser realtime + VAD dictation + local/cloud ASR + DeepSeek AI polish
dsh-audio-visualizer
caesarjue/dsh-audio-visualizer
System-audio driven UI visualizer for DeepSeek Harness: a draggable 48-band spectrum chip plus a frame-wide bass glow (memory-only FFT).
dsh-plugin-image-picker
zg2017/dsh-plugin-image-picker
Adds a file-picker 'attach image' button to the composer toolbar, since DSH's own web client only supports paste and drag-and-drop for image attachments (v1 scope, per its own upstream design notes) - a real gap for touch devices with no drag-and-drop and
dsh-desktop
new-256/dsh-desktop
DSH (DeepSeek Harness) ComfyUI 桥接插件:让 Agent 直接驱动本地/局域网 ComfyUI 生图生视频 — 8 个全局模型工具、预设工作流模板(txt2img/img2img/Wan/SVD/H3)、任意 API 工作流逃生舱、设备能力守卫、模型注册表断点续传下载、ComfyUI 一键安装评估。ComfyUI bridge plugin for DeepSeek Harness: image & video generation tools driving a local
dsh-ui-tool-result-images
suntianc/dsh-ui-tool-result-images
DeepSeek Harness Web plugin that keeps image-bearing Tool results visible after Compact transcript folding
dsh-voice-mimo
ch1bug/dsh-voice-mimo
Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).
dsh-image-router
zhiwuli0228/dsh-image-router
Digests the images in a prompt with a vision model before admission, so any model — a text-only one included — can read them without the session ever switching models, and adds a describe_image tool for image paths.
dsh-image-guard
mafeis/dsh-image-guard
Trims historical images in outgoing chat requests down to a recent-image count, learns the provider per-prompt image cap from HTTP 400 responses, and retries with fewer images so image-heavy sessions keep working.
dsh-mathmatic-symbol
jaxzhou/dsh-mathmatic-symbol
Three tools for DeepSeek Harness: typeset LaTeX formulas into images, draw mathematical figures from a declarative spec, and convert a formula or SVG into an image ready to embed in a document.
dsh-voice-input-qwen-asr
jsoncode/dsh-voice-input-qwen-asr
Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run ser
dsh-tts-flash
2021heei/dsh-tts-flash
Reads AI replies aloud as they stream, with voiced waiting phrases while the model thinks. Edge TTS built in, any OpenAI-compatible cloud engine supported.
dsh-image-generation
whites18/dsh-image-generation
Configure multiple image providers in Settings and call image_generate with the one selected model; images save under generate/image and show inline in the conversation.
dsh-reelsmaker
aayan-cloud/dsh-reelsmaker
DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.
dsh-voice-input-en
mohith-das/dsh-voice-input-en
Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow
dsh-plugins (dsh-ding-sound)
wwweljf/dsh-plugins
Turn-end notification sound for DSH: built-in internet meme voices (ni gan ma ai yo, ji ni tai mei, shen ying ge, etc.), Settings panel with preview/switch/random, custom audio folder support.
dsh-kitt-voice
kittcat-lab/dsh-kitt-voice
Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.
dsh-plugin-show-image
justhalfbit/dsh-plugin-show-image
DeepSeek Harness (DSH) 会话内图片渲染插件:全局 show_image 工具 + 点击放大 lightbox。 | Inline image rendering plugin for DSH: global show_image tool + click-to-enlarge lightbox.
dsh-plugins
aetheri-ai/dsh-plugins
DeepSeek Harness plugin: a model-facing show_image tool that presents a local image to the human viewer in the conversation.
dsh-voice-control
sucriss/dsh-voice-control
Voice control for the DSH web composer: push-to-talk speech-to-text into the input box (with optional auto-send), spoken playback of assistant replies via the Web Speech API, right-click settings popover, and a global Ctrl+M hotkey.
dsh-vision-pro-bridge
shainedemo/dsh-vision-pro-bridge
Vision bridge for text-only DeepSeek models: transcribes attached images with deepseek-v4-flash-vision-exp before they reach deepseek-v4-pro, with no third-party dependencies.
dsh-maclens
harzva/dsh-maclens
Apple on-device Vision tools for text-only dsh models: local OCR (zh-Hans + 30 langs), image classification, face detection, document layout, and a combined describe — 100% offline, no API key, tall-screenshot slicing.
dsh-image-generation (tool-image-generation)
ankye/dsh-image-generation
Model-facing image-generation tool with a configurable channel and normalized image parameters.
dsh-vision-bridge
alaxrpg/dsh-vision-bridge
Adds image input and recognition through configured DSH providers or an OpenAI-compatible endpoint.