Saltar al contenido principal

Visión, voz y multimodal

Los plugins de Visión, voz y multimodal dan ojos, oídos y voz a los modelos DeepSeek de solo texto en DeepSeek-Harness (dsh). Encuentra herramientas de visión y rutas de proveedor que hacen OCR y describen capturas pegadas mediante Zhipu GLM, Gemini, Doubao u Ollama local, entrada de voz por micrófono a través de la Web Speech API del navegador o APIs compatibles con Whisper, lectura en voz alta con Edge TTS o voces personalizadas, modos de voz full-duplex, y generadores de imágenes, vídeo y música.

312 plugins encontrados

F

dsh-voice-announcer

flashyiyi/dsh-voice-announcer

Voice announcements when a conversation ends (session title, round, outcome) plus live sentence-by-sentence reading of replies, via built-in edge-tts with zero third-party dependencies.

0hace 3 díasVisión, voz y multimodal
B

dsh-soundscape

berserk0501/dsh-soundscape

Dual-MediaPlayer daemon ($think loop + $fx one-shots) with volume ducking, mood-aware sounds (smooth/struggle/deep_think triggered by consecutive success/failure streaks and prolonged agent runtime), 15 built-in PCM WAV synthesised from sine/square waveforms, per-tool sound mapping with 51 categorised tools, custom WAV/MP3 drop-in with auto-conversion, and a full settings UI with dark-mode fix.

0el mes pasadoVisión, voz y multimodalMIT
K

deepseek-plugin

klingai-dev/deepseek-plugin

Turn every idea into an image or video with Kling AI. In DeepSeek Harness, use natural language for text-to-image, image-to-image, text-to-video, image-to-video, reference-image creation, task tracking, and result previews.

0el mes pasadoVisión, voz y multimodal
K

dsh-auto-vision

k2d5rqjpkg-art/dsh-auto-vision

Auto-switch the DeepSeek route to the vision model on demand: flash main session switches (A), pro keeps deep reasoning and delegates image reading to a vision subagent (B), subagents always switch, with fatal-failure fallback. No manual model switching.

0el mes pasadoVisión, voz y multimodal
C

dsh-evidence

cooberped/dsh-evidence

Turns attached files into versioned evidence: `search_documents` builds a private local index (SQLite FTS5 after a startup capability probe, dependency-free JS fallback otherwise) and returns compact evidence blocks carrying an exact coordinate — PDF page, PPTX slide, text/DOCX line range, or quoted XLSX `Sheet!Range` — which `read_document` expands only while the content version still matches. Contiguous CJK runs are indexed as overlapping bigrams and queried as phrases, so word order is preserved; uploaded raster images take the native vision attachment path instead.

0el mes pasadoVisión, voz y multimodalMIT
Z

dsh-tu4-inline-images

zehenk/dsh-tu4-inline-images

对话内联图片 DSH 插件 — 在 DeepSeek Harness (DSH) Web GUI 的对话中,出现本地图片路径即直接渲染为图片。 A DSH plugin that renders local image paths as inline images in DeepSeek Harness (DSH) web conversations. Security-first: loopback-only route, strong per-process token, multi-root realpath whitelist.

0el mes pasadoVisión, voz y multimodalMIT
B

dsh-llm-capabilities

bamboostrip/dsh-llm-capabilities

DSH plugin: auto-detect and configure model capabilities (reasoningEfforts + input modalities) for llm-pi-ai. Successor to dsh-reasoning-efforts.

0el mes pasadoVisión, voz y multimodalMIT
L

dsh-vision-toggle

lijian-ui/dsh-vision-toggle

Per-model vision (image input) toggle for DeepSeek Harness (dsh): list every configured model and flip a switch to enable/disable image support without hand-editing settings.yaml. 为 DeepSeek Harness 提供按模型的「支持图片」开关:无需手改 settings.yaml。

0el mes pasadoVisión, voz y multimodalMIT
A

dsh-speech

allmodels-io/dsh-speech

Streaming speech-to-text for DeepSeek Harness using AllModels.io.

0el mes pasadoVisión, voz y multimodalMIT
C

dsh-agnes

chaoliu615/dsh-agnes

Agnes AI image & video generation tools for DSH (agnes_image_generate / agnes_video_generate)

0el mes pasadoVisión, voz y multimodalMIT
D

dsh-remote-deliver

demacia1314/dsh-remote-deliver

🚀 告别繁琐 SCP!远程部署 DSH 一键下载修改后的文件与图片预览交付插件

0el mes pasadoVisión, voz y multimodal
W

dsh-dictation

wsl043/dsh-dictation

Editable local and desktop dictation for DeepSeek Harness

0el mes pasadoVisión, voz y multimodalMIT
L

dsh-file-attach

lucasxingg/dsh-file-attach

Drag-and-drop PDF, Office, images, and text/code files into DSH conversations. The host extracts (and OCRs) them into the prompt; attach_* tools cover notebook cells, PDF-page OCR, image describe, and save.

0el mes pasadoVisión, voz y multimodalMIT
J

dsh-csv-and-image-preview

jetecho/dsh-csv-and-image-preview

Preview images / SVG (and prepare CSV) in the DeepSeek Harness chat, rendered as real browser <img> elements. Preview-first workflow: show the user the asset, wait for approval, then apply the real change.

0el mes pasadoVisión, voz y multimodalMIT
X

dsh-voice-input-space

xsakura666/dsh-voice-input-space

Voice input for DeepSeek Harness: hold Space to speak, release to insert. Zero dependencies, Web Speech API. / 语音输入:长按空格说话,松开上屏,零依赖。

0el mes pasadoVisión, voz y multimodalMIT
N

dsh-voice

navid-kianfar/dsh-voice

Dictate prompts into the DeepSeek Harness Web Client — a microphone in the composer, with swappable transcription: hosted Whisper API, self-hosted server, or a fully offline whisper.cpp binary.

0el mes pasadoVisión, voz y multimodalMIT
Z

dsh-video-understand

zeshuochen/dsh-video-understand

Subtitle-first video transcription and deterministic extractive Markdown summaries, with a faster-whisper large-v3 fallback when subtitles are unavailable.

0el mes pasadoVisión, voz y multimodal
S

dsh-auto-vision

soarguo/dsh-auto-vision

Bridges images into text for DeepSeek Harness: when the session's selected model cannot see images, a configured vision model describes them and the descriptions enter the durable session history as folded context rows — your message stays untouched.

0el mes pasadoVisión, voz y multimodalMIT
C

dsh-mmx

crazyma99/dsh-mmx

MiniMax CLI bridge for DeepSeek Harness (dsh): mmx-backed web search fallback and transparent image understanding, with first-run onboarding and a settings card.

0el mes pasadoVisión, voz y multimodalMIT
Q

dsh-voice-mode

qishuilalala/dsh-voice-mode

Full-duplex voice mode for DeepSeek Harness: zipformer2 streaming ASR → editable draft, Edge TTS sentence-by-sentence read-aloud with live captions, true barge-in — on-device ASR, no API key. · DSH 语音双工对话:流式识别入草稿、按句朗读+实时字幕、开口即打断,识别本地推理、无需 API Key

0el mes pasadoVisión, voz y multimodalMIT
Y

dsh-mimo-plugin

yyfather/dsh-mimo-plugin

MiMo (Xiaomi) tools as a DSH profile plugin: web search, image/audio/video understanding, ASR transcription, TTS, voice design and voice cloning — native agent tools with a Settings page for the API key.

0el mes pasadoVisión, voz y multimodalMIT
B

dsh-plugin-88api-image

blackdm666/dsh-plugin-88api-image

88API Image Studio for DSH: four Image2 and Nano Banana models for text-to-image, multi-reference editing, 2K/4K output, and sequential batches.

0el mes pasadoVisión, voz y multimodalMIT
X

computer-use-vision

xuanyuanluoxue/computer-use-vision

Windows computer-use capability for DeepSeek Harness: screenshot → vision model → simulated mouse/keyboard input, with self-evolving knowledge base.

0el mes pasadoVisión, voz y multimodal
F

dsh-glm-vision

fightingfirefox/dsh-glm-vision

GLM 视觉模型插件:注册 glm-vision 供应商路由(glm-4.6v 系列,image+text 输入声明)并提供 glm_vision 工具,让 DeepSeek 等文本主模型直接调用智谱视觉模型看图。

0el mes pasadoVisión, voz y multimodalMIT