- Inicio
- Categorías
- Visión, voz y multimodal
Visión, voz y multimodal
Los plugins de Visión, voz y multimodal dan ojos, oídos y voz a los modelos DeepSeek de solo texto en DeepSeek-Harness (dsh). Encuentra herramientas de visión y rutas de proveedor que hacen OCR y describen capturas pegadas mediante Zhipu GLM, Gemini, Doubao u Ollama local, entrada de voz por micrófono a través de la Web Speech API del navegador o APIs compatibles con Whisper, lectura en voz alta con Edge TTS o voces personalizadas, modos de voz full-duplex, y generadores de imágenes, vídeo y música.
312 plugins encontrados
dsh-agnes-media
zlforward/dsh-agnes-media
dsh-funasr-voice
fenglin-ai/dsh-funasr-voice
Offline voice input for the DSH Web UI: mic to local FunASR (SenseVoiceSmall), one-click install, no cloud.
dshtools-sensevoice-input
ilovedyou6666-hub/dshtools-sensevoice-input
基于 SenseVoiceSmall(iic/SenseVoiceSmall)多语言语音理解模型的 DSH Desktop 本地语音输入插件。
dsh-file-convert
zzy-12345678/dsh-file-convert
Local-first file conversion: 26 conversions across images, PDF (with OCR and experimental PDF→DOCX), data, audio/video and office docs; 7 tools, all local, no API keys.
dsh-voice-input-npm
difimim/dsh-voice-input-npm
语音输入插件 for Deepseek Harness
dsh-live-voice
jstn-1g/dsh-live-voice
Consent-bound one-turn voice preview for DSH Web with a credential-free local synthetic demo, optional Qwen Audio, exact Session isolation, and explicit transcript-to-draft handoff without automatic submission.
dsh-video-gen
yang-wudi/dsh-video-gen
Text-to-video and image-to-video generation via DashScope Wanx, Volcengine Seedance, Google Veo and OpenAI Sora, with tool results saved to the session workspace and a video gallery (grid, lightbox, download, delete) that survives DSH restarts.
dsh-image-viewer
wsl043/dsh-image-viewer
Upgrades DSH image viewing with pointer-centered zoom, pan, galleries, downloads, keyboard support, and inline region notes.
dsh-freecanvas
justinqiuck/dsh-freecanvas
Install DSH FreeCanvas as a bundled DeepSeek Harness app with split layouts and managed local Agent connectivity. · 将 DSH FreeCanvas 作为内置应用安装到 DSH,支持分屏布局与本地 Agent 自动连接。
dsh-bilibili
moxingovo/dsh-bilibili
DeepSeek Harness plugin: Bilibili keyword video search, video metadata, subtitle transcripts, direct play URLs, and multimodal frame viewing (bilibili_search / bilibili_video / bilibili_subtitles / bilibili_playurl / bilibili_frames). Anonymous by default
dsh-pianist
laplace-bit/dsh-pianist
Piano performance plugin: ask the agent to play a piece and it renders on a Canvas2D grand piano with real Salamander Grand samples, an immersive stage, and an interactive 88-key keyboard.
dsh-image-preview
algerkong/dsh-image-preview
Image preview for DSH (DeepSeek Harness) web sessions: read_image results render as a thumbnail, click for full size in the built-in lightbox.
dsh-mmroute
jmxsxwyzjdwl/dsh-mmroute
Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.
aura-vision
ck-epsilon/aura-vision
Free vision OCR with adaptive tile recognition for long documents and Markdown/Word/PNG/Excel export.
taxue-dsh-artisan
taxueseek/taxue-dsh-artisan
Integrated visual creation toolchain for DSH: prompt reverse-engineering and audit optimization plus multi-provider image generation with async background rendering.
dsh-vision-analysis
harvey-will/dsh-vision-analysis
DeepSeek Harness vision plugin: 8 analysis modes (describe, OCR, chart data, UI review, object detection, compare, code-gen, debug), any OpenAI- or Anthropic-compatible vision API, with a built-in free vision model and automatic rate-limit failover.
remotion-video-plugin
chenjie1129/remotion-video-plugin
Remotion video creation and verified rendering plugin for DeepSeek Harness
dsh-screenshot
alain-prot0s5/dsh-screenshot
Screenshot-to-input for DeepSeek Harness: composer camera button + global hotkey (Alt+A) + listener bound to the app lifecycle, configurable in settings. 截图自动粘贴到 DSH 输入框:相机按钮 + 全局快捷键 + 生命周期绑定 + 设置页配置。
dsh-composer-image-tools
ai-yucheng/dsh-composer-image-tools
聊天输入框图片工具(自研):上传图片(≤10MB 防烧 token) + 自定义区域截图(Electron desktopCapturer),注入 DSH 草稿图片轨。零外部依赖。
dsh-speech-input
liznee/dsh-speech-input
A microphone button for DeepSeek Harness that writes browser speech recognition into the composer draft.
dsh-voice-input-cn
schumchanvi/dsh-voice-input-cn
Entrada de voz lista para China para el compositor. Requiere un puente local en Python (pip install dashscope websockets, ejecutar bridge/voice-bridge.py); el plugin por sí solo no funciona. El micrófono del navegador transmite PCM de 16 kHz al puente, que ejecuta Alibaba Cloud DashScope ASR (paraformer-realtime-v2); el texto intermedio rellena el borrador en el cursor, parada automática por silencio y envío automático opcional.
dsh-client-vision (tool-vision)
ankye/dsh-client-vision
Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.
dsh-narrate
stuarthu/dsh-narrate
DeepSeek Harness (dsh) plugin: turn one idea into a narrated video cut from your own asset folder, stopping four times to ask you first.
dsh-audio-copilot
ai-yucheng/dsh-audio-copilot
Audio Copilot for DeepSeek Harness: transcribe audio (ASR) and synthesize speech (TTS) — gives text-only agents ears and a voice. Windows-local SAPI TTS out of the box; OpenAI-compatible ASR/TTS endpoints configurable. Includes an in-composer voice-input