- Inicio
- Categorías
- Visión, voz y multimodal
Visión, voz y multimodal
Los plugins de Visión, voz y multimodal dan ojos, oídos y voz a los modelos DeepSeek de solo texto en DeepSeek-Harness (dsh). Encuentra herramientas de visión y rutas de proveedor que hacen OCR y describen capturas pegadas mediante Zhipu GLM, Gemini, Doubao u Ollama local, entrada de voz por micrófono a través de la Web Speech API del navegador o APIs compatibles con Whisper, lectura en voz alta con Edge TTS o voces personalizadas, modos de voz full-duplex, y generadores de imágenes, vídeo y música.
312 plugins encontrados
dsh-tier-router
zhangzhangco/dsh-tier-router
Automatic tier-based model routing for DeepSeek Harness: a virtual `smart` model classifies each request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured. 三级难度 + 视觉自动路由,虚拟 smart 模型零配置接入。
dsh-voice-input
rio-promax/dsh-voice-input
DSH voice input plugin: browser realtime + VAD dictation + local/cloud ASR + DeepSeek AI polish
dsh-audio-visualizer
caesarjue/dsh-audio-visualizer
System-audio driven UI visualizer for DeepSeek Harness: a draggable 48-band spectrum chip plus a frame-wide bass glow (memory-only FFT).
dsh-plugin-image-picker
zg2017/dsh-plugin-image-picker
Adds a file-picker 'attach image' button to the composer toolbar, since DSH's own web client only supports paste and drag-and-drop for image attachments (v1 scope, per its own upstream design notes) - a real gap for touch devices with no drag-and-drop and
dsh-desktop
new-256/dsh-desktop
DSH (DeepSeek Harness) ComfyUI 桥接插件:让 Agent 直接驱动本地/局域网 ComfyUI 生图生视频 — 8 个全局模型工具、预设工作流模板(txt2img/img2img/Wan/SVD/H3)、任意 API 工作流逃生舱、设备能力守卫、模型注册表断点续传下载、ComfyUI 一键安装评估。ComfyUI bridge plugin for DeepSeek Harness: image & video generation tools driving a local
dsh-ui-tool-result-images
suntianc/dsh-ui-tool-result-images
DeepSeek Harness Web plugin that keeps image-bearing Tool results visible after Compact transcript folding
dsh-voice-mimo
ch1bug/dsh-voice-mimo
Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).
dsh-image-router
zhiwuli0228/dsh-image-router
Digests the images in a prompt with a vision model before admission, so any model — a text-only one included — can read them without the session ever switching models, and adds a describe_image tool for image paths.
dsh-image-guard
mafeis/dsh-image-guard
Trims historical images in outgoing chat requests down to a recent-image count, learns the provider per-prompt image cap from HTTP 400 responses, and retries with fewer images so image-heavy sessions keep working.
dsh-mathmatic-symbol
jaxzhou/dsh-mathmatic-symbol
Three tools for DeepSeek Harness: typeset LaTeX formulas into images, draw mathematical figures from a declarative spec, and convert a formula or SVG into an image ready to embed in a document.
dsh-voice-input-qwen-asr
jsoncode/dsh-voice-input-qwen-asr
Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run ser
dsh-tts-flash
2021heei/dsh-tts-flash
Reads AI replies aloud as they stream, with voiced waiting phrases while the model thinks. Edge TTS built in, any OpenAI-compatible cloud engine supported.
dsh-image-generation
whites18/dsh-image-generation
Configure multiple image providers in Settings and call image_generate with the one selected model; images save under generate/image and show inline in the conversation.
dsh-reelsmaker
aayan-cloud/dsh-reelsmaker
DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.
dsh-voice-input-en
mohith-das/dsh-voice-input-en
Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow
dsh-plugins (dsh-ding-sound)
wwweljf/dsh-plugins
Turn-end notification sound for DSH: built-in internet meme voices (ni gan ma ai yo, ji ni tai mei, shen ying ge, etc.), Settings panel with preview/switch/random, custom audio folder support.
dsh-kitt-voice
kittcat-lab/dsh-kitt-voice
Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.
dsh-plugin-show-image
justhalfbit/dsh-plugin-show-image
DeepSeek Harness (DSH) 会话内图片渲染插件:全局 show_image 工具 + 点击放大 lightbox。 | Inline image rendering plugin for DSH: global show_image tool + click-to-enlarge lightbox.
dsh-plugins
aetheri-ai/dsh-plugins
DeepSeek Harness plugin: a model-facing show_image tool that presents a local image to the human viewer in the conversation.
dsh-voice-control
sucriss/dsh-voice-control
Voice control for the DSH web composer: push-to-talk speech-to-text into the input box (with optional auto-send), spoken playback of assistant replies via the Web Speech API, right-click settings popover, and a global Ctrl+M hotkey.
dsh-vision-pro-bridge
shainedemo/dsh-vision-pro-bridge
Vision bridge for text-only DeepSeek models: transcribes attached images with deepseek-v4-flash-vision-exp before they reach deepseek-v4-pro, with no third-party dependencies.
dsh-maclens
harzva/dsh-maclens
Apple on-device Vision tools for text-only dsh models: local OCR (zh-Hans + 30 langs), image classification, face detection, document layout, and a combined describe — 100% offline, no API key, tall-screenshot slicing.
dsh-image-generation (tool-image-generation)
ankye/dsh-image-generation
Model-facing image-generation tool with a configurable channel and normalized image parameters.
dsh-vision-bridge
alaxrpg/dsh-vision-bridge
Adds image input and recognition through configured DSH providers or an OpenAI-compatible endpoint.