- Accueil
- Catégories
- Vision, voix et multimodal
Vision, voix et multimodal
Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.
312 plugins trouvés
dsh-agnes-media
zlforward/dsh-agnes-media
dsh-funasr-voice
fenglin-ai/dsh-funasr-voice
Offline voice input for the DSH Web UI: mic to local FunASR (SenseVoiceSmall), one-click install, no cloud.
dshtools-sensevoice-input
ilovedyou6666-hub/dshtools-sensevoice-input
基于 SenseVoiceSmall(iic/SenseVoiceSmall)多语言语音理解模型的 DSH Desktop 本地语音输入插件。
dsh-file-convert
zzy-12345678/dsh-file-convert
Local-first file conversion: 26 conversions across images, PDF (with OCR and experimental PDF→DOCX), data, audio/video and office docs; 7 tools, all local, no API keys.
dsh-voice-input-npm
difimim/dsh-voice-input-npm
语音输入插件 for Deepseek Harness
dsh-live-voice
jstn-1g/dsh-live-voice
Consent-bound one-turn voice preview for DSH Web with a credential-free local synthetic demo, optional Qwen Audio, exact Session isolation, and explicit transcript-to-draft handoff without automatic submission.
dsh-video-gen
yang-wudi/dsh-video-gen
Text-to-video and image-to-video generation via DashScope Wanx, Volcengine Seedance, Google Veo and OpenAI Sora, with tool results saved to the session workspace and a video gallery (grid, lightbox, download, delete) that survives DSH restarts.
dsh-image-viewer
wsl043/dsh-image-viewer
Upgrades DSH image viewing with pointer-centered zoom, pan, galleries, downloads, keyboard support, and inline region notes.
dsh-freecanvas
justinqiuck/dsh-freecanvas
Install DSH FreeCanvas as a bundled DeepSeek Harness app with split layouts and managed local Agent connectivity. · 将 DSH FreeCanvas 作为内置应用安装到 DSH,支持分屏布局与本地 Agent 自动连接。
dsh-bilibili
moxingovo/dsh-bilibili
DeepSeek Harness plugin: Bilibili keyword video search, video metadata, subtitle transcripts, direct play URLs, and multimodal frame viewing (bilibili_search / bilibili_video / bilibili_subtitles / bilibili_playurl / bilibili_frames). Anonymous by default
dsh-pianist
laplace-bit/dsh-pianist
Piano performance plugin: ask the agent to play a piece and it renders on a Canvas2D grand piano with real Salamander Grand samples, an immersive stage, and an interactive 88-key keyboard.
dsh-image-preview
algerkong/dsh-image-preview
Image preview for DSH (DeepSeek Harness) web sessions: read_image results render as a thumbnail, click for full size in the built-in lightbox.
dsh-mmroute
jmxsxwyzjdwl/dsh-mmroute
Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.
aura-vision
ck-epsilon/aura-vision
Free vision OCR with adaptive tile recognition for long documents and Markdown/Word/PNG/Excel export.
taxue-dsh-artisan
taxueseek/taxue-dsh-artisan
Integrated visual creation toolchain for DSH: prompt reverse-engineering and audit optimization plus multi-provider image generation with async background rendering.
dsh-vision-analysis
harvey-will/dsh-vision-analysis
DeepSeek Harness vision plugin: 8 analysis modes (describe, OCR, chart data, UI review, object detection, compare, code-gen, debug), any OpenAI- or Anthropic-compatible vision API, with a built-in free vision model and automatic rate-limit failover.
remotion-video-plugin
chenjie1129/remotion-video-plugin
Remotion video creation and verified rendering plugin for DeepSeek Harness
dsh-screenshot
alain-prot0s5/dsh-screenshot
Screenshot-to-input for DeepSeek Harness: composer camera button + global hotkey (Alt+A) + listener bound to the app lifecycle, configurable in settings. 截图自动粘贴到 DSH 输入框:相机按钮 + 全局快捷键 + 生命周期绑定 + 设置页配置。
dsh-composer-image-tools
ai-yucheng/dsh-composer-image-tools
聊天输入框图片工具(自研):上传图片(≤10MB 防烧 token) + 自定义区域截图(Electron desktopCapturer),注入 DSH 草稿图片轨。零外部依赖。
dsh-speech-input
liznee/dsh-speech-input
A microphone button for DeepSeek Harness that writes browser speech recognition into the composer draft.
dsh-voice-input-cn
schumchanvi/dsh-voice-input-cn
Saisie vocale prête pour la Chine pour le composer. Nécessite un pont Python local (pip install dashscope websockets, exécuter bridge/voice-bridge.py) — le plugin seul ne fonctionne pas. Le micro du navigateur diffuse du PCM 16 kHz au pont, qui exécute Alibaba Cloud DashScope ASR (paraformer-realtime-v2) ; le texte intermédiaire remplit le brouillon au curseur, arrêt automatique au silence, envoi automatique optionnel.
dsh-client-vision (tool-vision)
ankye/dsh-client-vision
Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.
dsh-narrate
stuarthu/dsh-narrate
DeepSeek Harness (dsh) plugin: turn one idea into a narrated video cut from your own asset folder, stopping four times to ask you first.
dsh-audio-copilot
ai-yucheng/dsh-audio-copilot
Audio Copilot for DeepSeek Harness: transcribe audio (ASR) and synthesize speech (TTS) — gives text-only agents ears and a voice. Windows-local SAPI TTS out of the box; OpenAI-compatible ASR/TTS endpoints configurable. Includes an in-composer voice-input