Vai al contenuto principale

Visione, voce e multimodale

I plugin Visione, voce e multimodale danno occhi, orecchie e voce ai modelli DeepSeek solo testo di DeepSeek-Harness (dsh). Trova strumenti di visione e route di provider che eseguono OCR e descrivono gli screenshot incollati tramite Zhipu GLM, Gemini, Doubao o Ollama in locale, input vocale dal microfono tramite la Web Speech API del browser o API compatibili con Whisper, lettura ad alta voce con Edge TTS o voci personalizzate, modalità vocali full-duplex, e generatori di immagini, video e musica.

312 plugin trovati

Z

dsh-agnes-media

zlforward/dsh-agnes-media

123 giorni faVisione, voce e multimodaleMIT
F

dsh-funasr-voice

fenglin-ai/dsh-funasr-voice

Offline voice input for the DSH Web UI: mic to local FunASR (SenseVoiceSmall), one-click install, no cloud.

1mese scorsoVisione, voce e multimodaleMIT
I

dshtools-sensevoice-input

ilovedyou6666-hub/dshtools-sensevoice-input

基于 SenseVoiceSmall(iic/SenseVoiceSmall)多语言语音理解模型的 DSH Desktop 本地语音输入插件。

129 giorni faVisione, voce e multimodaleMIT
Z

dsh-file-convert

zzy-12345678/dsh-file-convert

Local-first file conversion: 26 conversions across images, PDF (with OCR and experimental PDF→DOCX), data, audio/video and office docs; 7 tools, all local, no API keys.

1mese scorsoVisione, voce e multimodale
D

dsh-voice-input-npm

difimim/dsh-voice-input-npm

语音输入插件 for Deepseek Harness

1mese scorsoVisione, voce e multimodaleMIT
J

dsh-live-voice

jstn-1g/dsh-live-voice

Consent-bound one-turn voice preview for DSH Web with a credential-free local synthetic demo, optional Qwen Audio, exact Session isolation, and explicit transcript-to-draft handoff without automatic submission.

127 giorni faVisione, voce e multimodaleMIT
Y

dsh-video-gen

yang-wudi/dsh-video-gen

Text-to-video and image-to-video generation via DashScope Wanx, Volcengine Seedance, Google Veo and OpenAI Sora, with tool results saved to the session workspace and a video gallery (grid, lightbox, download, delete) that survives DSH restarts.

1mese scorsoVisione, voce e multimodale
W

dsh-image-viewer

wsl043/dsh-image-viewer

Upgrades DSH image viewing with pointer-centered zoom, pan, galleries, downloads, keyboard support, and inline region notes.

1l’altro ieriVisione, voce e multimodaleMIT
J

dsh-freecanvas

justinqiuck/dsh-freecanvas

Install DSH FreeCanvas as a bundled DeepSeek Harness app with split layouts and managed local Agent connectivity. · 将 DSH FreeCanvas 作为内置应用安装到 DSH,支持分屏布局与本地 Agent 自动连接。

1mese scorsoVisione, voce e multimodaleMIT
M

dsh-bilibili

moxingovo/dsh-bilibili

DeepSeek Harness plugin: Bilibili keyword video search, video metadata, subtitle transcripts, direct play URLs, and multimodal frame viewing (bilibili_search / bilibili_video / bilibili_subtitles / bilibili_playurl / bilibili_frames). Anonymous by default

1mese scorsoVisione, voce e multimodaleMIT
L

dsh-pianist

laplace-bit/dsh-pianist

Piano performance plugin: ask the agent to play a piece and it renders on a Canvas2D grand piano with real Salamander Grand samples, an immersive stage, and an interactive 88-key keyboard.

123 giorni faVisione, voce e multimodaleMIT
A

dsh-image-preview

algerkong/dsh-image-preview

Image preview for DSH (DeepSeek Harness) web sessions: read_image results render as a thumbnail, click for full size in the built-in lightbox.

1mese scorsoVisione, voce e multimodale
J

dsh-mmroute

jmxsxwyzjdwl/dsh-mmroute

Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.

126 giorni faVisione, voce e multimodaleMIT
C

aura-vision

ck-epsilon/aura-vision

Free vision OCR with adaptive tile recognition for long documents and Markdown/Word/PNG/Excel export.

1mese scorsoVisione, voce e multimodaleMIT
T

taxue-dsh-artisan

taxueseek/taxue-dsh-artisan

Integrated visual creation toolchain for DSH: prompt reverse-engineering and audit optimization plus multi-provider image generation with async background rendering.

1mese scorsoVisione, voce e multimodaleMIT
H

dsh-vision-analysis

harvey-will/dsh-vision-analysis

DeepSeek Harness vision plugin: 8 analysis modes (describe, OCR, chart data, UI review, object detection, compare, code-gen, debug), any OpenAI- or Anthropic-compatible vision API, with a built-in free vision model and automatic rate-limit failover.

13 giorni faVisione, voce e multimodaleMIT
C

remotion-video-plugin

chenjie1129/remotion-video-plugin

Remotion video creation and verified rendering plugin for DeepSeek Harness

1mese scorsoVisione, voce e multimodaleMIT
A

dsh-screenshot

alain-prot0s5/dsh-screenshot

Screenshot-to-input for DeepSeek Harness: composer camera button + global hotkey (Alt+A) + listener bound to the app lifecycle, configurable in settings. 截图自动粘贴到 DSH 输入框:相机按钮 + 全局快捷键 + 生命周期绑定 + 设置页配置。

1mese scorsoVisione, voce e multimodaleMIT
A

dsh-composer-image-tools

ai-yucheng/dsh-composer-image-tools

聊天输入框图片工具(自研):上传图片(≤10MB 防烧 token) + 自定义区域截图(Electron desktopCapturer),注入 DSH 草稿图片轨。零外部依赖。

1mese scorsoVisione, voce e multimodaleMIT
L

dsh-speech-input

liznee/dsh-speech-input

A microphone button for DeepSeek Harness that writes browser speech recognition into the composer draft.

1mese scorsoVisione, voce e multimodaleMIT
S

dsh-voice-input-cn

schumchanvi/dsh-voice-input-cn

Input vocale pronto per la Cina per il composer. Richiede un bridge Python locale (pip install dashscope websockets, esegui bridge/voice-bridge.py) — il plugin da solo non funziona. Il microfono del browser trasmette PCM a 16 kHz al bridge, che esegue Alibaba Cloud DashScope ASR (paraformer-realtime-v2); il testo intermedio riempie la bozza al cursore, arresto automatico al silenzio, invio automatico opzionale.

116 giorni faVisione, voce e multimodaleMIT
A

dsh-client-vision (tool-vision)

ankye/dsh-client-vision

Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.

123 giorni faVisione, voce e multimodaleMIT
S

dsh-narrate

stuarthu/dsh-narrate

DeepSeek Harness (dsh) plugin: turn one idea into a narrated video cut from your own asset folder, stopping four times to ask you first.

1mese scorsoVisione, voce e multimodaleMIT
A

dsh-audio-copilot

ai-yucheng/dsh-audio-copilot

Audio Copilot for DeepSeek Harness: transcribe audio (ASR) and synthesize speech (TTS) — gives text-only agents ears and a voice. Windows-local SAPI TTS out of the box; OpenAI-compatible ASR/TTS endpoints configurable. Includes an in-composer voice-input

1mese scorsoVisione, voce e multimodaleMIT