Saltar al contenido principal

Visión, voz y multimodal

Los plugins de Visión, voz y multimodal dan ojos, oídos y voz a los modelos DeepSeek de solo texto en DeepSeek-Harness (dsh). Encuentra herramientas de visión y rutas de proveedor que hacen OCR y describen capturas pegadas mediante Zhipu GLM, Gemini, Doubao u Ollama local, entrada de voz por micrófono a través de la Web Speech API del navegador o APIs compatibles con Whisper, lectura en voz alta con Edge TTS o voces personalizadas, modos de voz full-duplex, y generadores de imágenes, vídeo y música.

312 plugins encontrados

Z

dsh-tier-router

zhangzhangco/dsh-tier-router

Automatic tier-based model routing for DeepSeek Harness: a virtual `smart` model classifies each request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured. 三级难度 + 视觉自动路由,虚拟 smart 模型零配置接入。

0hace 17 díasVisión, voz y multimodal
R

dsh-voice-input

rio-promax/dsh-voice-input

DSH voice input plugin: browser realtime + VAD dictation + local/cloud ASR + DeepSeek AI polish

0hace 21 díasVisión, voz y multimodalMIT
C

dsh-audio-visualizer

caesarjue/dsh-audio-visualizer

System-audio driven UI visualizer for DeepSeek Harness: a draggable 48-band spectrum chip plus a frame-wide bass glow (memory-only FFT).

0hace 21 díasVisión, voz y multimodalMIT
Z

dsh-plugin-image-picker

zg2017/dsh-plugin-image-picker

Adds a file-picker 'attach image' button to the composer toolbar, since DSH's own web client only supports paste and drag-and-drop for image attachments (v1 scope, per its own upstream design notes) - a real gap for touch devices with no drag-and-drop and

0el mes pasadoVisión, voz y multimodalMIT
N

dsh-desktop

new-256/dsh-desktop

DSH (DeepSeek Harness) ComfyUI 桥接插件:让 Agent 直接驱动本地/局域网 ComfyUI 生图生视频 — 8 个全局模型工具、预设工作流模板(txt2img/img2img/Wan/SVD/H3)、任意 API 工作流逃生舱、设备能力守卫、模型注册表断点续传下载、ComfyUI 一键安装评估。ComfyUI bridge plugin for DeepSeek Harness: image & video generation tools driving a local

0hace 23 díasVisión, voz y multimodalMIT
S

dsh-ui-tool-result-images

suntianc/dsh-ui-tool-result-images

DeepSeek Harness Web plugin that keeps image-bearing Tool results visible after Compact transcript folding

0hace 23 díasVisión, voz y multimodalMIT
C

dsh-voice-mimo

ch1bug/dsh-voice-mimo

Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).

0hace 19 díasVisión, voz y multimodalMIT
Z

dsh-image-router

zhiwuli0228/dsh-image-router

Digests the images in a prompt with a vision model before admission, so any model — a text-only one included — can read them without the session ever switching models, and adds a describe_image tool for image paths.

0hace 21 díasVisión, voz y multimodalMIT
M

dsh-image-guard

mafeis/dsh-image-guard

Trims historical images in outgoing chat requests down to a recent-image count, learns the provider per-prompt image cap from HTTP 400 responses, and retries with fewer images so image-heavy sessions keep working.

0hace 19 díasVisión, voz y multimodalMIT
J

dsh-mathmatic-symbol

jaxzhou/dsh-mathmatic-symbol

Three tools for DeepSeek Harness: typeset LaTeX formulas into images, draw mathematical figures from a declarative spec, and convert a formula or SVG into an image ready to embed in a document.

0hace 21 díasVisión, voz y multimodalMIT
J

dsh-voice-input-qwen-asr

jsoncode/dsh-voice-input-qwen-asr

Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run ser

0hace 28 díasVisión, voz y multimodal
2

dsh-tts-flash

2021heei/dsh-tts-flash

Reads AI replies aloud as they stream, with voiced waiting phrases while the model thinks. Edge TTS built in, any OpenAI-compatible cloud engine supported.

0hace 23 díasVisión, voz y multimodalMIT
W

dsh-image-generation

whites18/dsh-image-generation

Configure multiple image providers in Settings and call image_generate with the one selected model; images save under generate/image and show inline in the conversation.

0hace 5 díasVisión, voz y multimodalMIT
A

dsh-reelsmaker

aayan-cloud/dsh-reelsmaker

DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.

0hace 28 díasVisión, voz y multimodal
M

dsh-voice-input-en

mohith-das/dsh-voice-input-en

Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow

0el mes pasadoVisión, voz y multimodalMIT
W

dsh-plugins (dsh-ding-sound)

wwweljf/dsh-plugins

Turn-end notification sound for DSH: built-in internet meme voices (ni gan ma ai yo, ji ni tai mei, shen ying ge, etc.), Settings panel with preview/switch/random, custom audio folder support.

0el mes pasadoVisión, voz y multimodalMIT
K

dsh-kitt-voice

kittcat-lab/dsh-kitt-voice

Voice for the DeepSeek Harness: speak to the agent, hear it back, and see what it is doing from a floating window that stays on top of whatever you are running.

0hace 28 díasVisión, voz y multimodalMIT
J

dsh-plugin-show-image

justhalfbit/dsh-plugin-show-image

DeepSeek Harness (DSH) 会话内图片渲染插件:全局 show_image 工具 + 点击放大 lightbox。 | Inline image rendering plugin for DSH: global show_image tool + click-to-enlarge lightbox.

0el mes pasadoVisión, voz y multimodalMIT
A

dsh-plugins

aetheri-ai/dsh-plugins

DeepSeek Harness plugin: a model-facing show_image tool that presents a local image to the human viewer in the conversation.

0el mes pasadoVisión, voz y multimodalMIT
S

dsh-voice-control

sucriss/dsh-voice-control

Voice control for the DSH web composer: push-to-talk speech-to-text into the input box (with optional auto-send), spoken playback of assistant replies via the Web Speech API, right-click settings popover, and a global Ctrl+M hotkey.

0el mes pasadoVisión, voz y multimodalMIT
S

dsh-vision-pro-bridge

shainedemo/dsh-vision-pro-bridge

Vision bridge for text-only DeepSeek models: transcribes attached images with deepseek-v4-flash-vision-exp before they reach deepseek-v4-pro, with no third-party dependencies.

0el mes pasadoVisión, voz y multimodalMIT
H

dsh-maclens

harzva/dsh-maclens

Apple on-device Vision tools for text-only dsh models: local OCR (zh-Hans + 30 langs), image classification, face detection, document layout, and a combined describe — 100% offline, no API key, tall-screenshot slicing.

0el mes pasadoVisión, voz y multimodalMIT
A

dsh-image-generation (tool-image-generation)

ankye/dsh-image-generation

Model-facing image-generation tool with a configurable channel and normalized image parameters.

0el mes pasadoVisión, voz y multimodalMIT
A

dsh-vision-bridge

alaxrpg/dsh-vision-bridge

Adds image input and recognition through configured DSH providers or an OpenAI-compatible endpoint.

0hace 5 díasVisión, voz y multimodalMIT