- Accueil
- Catégories
- Vision, voix et multimodal
Vision, voix et multimodal
Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.
312 plugins trouvés
dsh-plugin-vision-toolkit
yytbit/dsh-plugin-vision-toolkit
Boîte à outils de vision pour DeepSeek Harness -- outils CLI glance, ground, detect, crop pour que les agents texte seul comprennent les images.
dsh-yali-image-generator
pptt121212/dsh-yali-image-generator
Plugin de génération d'images DeepSeek-Harness. Demandez une clé API Yali AI : https://api.yaliai.com/
deepsee
chang416/deepsee
DeepSee : vision DeepSeek Harness, routage multi-modèles, et auto-vérifications visuelles Gemini avant livraison.
dsh-sight
fu3rte/dsh-sight
Vision en plugin pour les modèles DeepSeek Harness (dsh) texte seul : un outil `vision` avec des préréglages VLM gratuits/économiques intégrés, analyse par lot multi-images, admission d'image par indice de collage, et une page de réglages web avec rechargement à chaud.
dsh-friend
wanghehe123/dsh-friend
Plugin de compagnon personnifié pour DeepSeek Harness : fiches de personnage, voix, Live2D, mémoire locale et accompagnement pendant le travail.
dsh-voice-webspeech
anweat/dsh-voice-webspeech
Saisie vocale via la Web Speech API du navigateur : zéro serveur, zéro clé, zéro téléchargement de modèle (Edge=Azure, Chrome=Google speech).
dsh-plugins
retiredphysicist/dsh-plugins
DeepSeek Harness plugin: web browsing tools (markdown / screenshot / pdf / crawl) powered by Cloudflare Browser Run — real headless Chrome with JS rendering, login sessions and WebMCP support.
dsh-kite
weibaohui/dsh-kite
Animated kite that responds to agent activity, with configurable kite frames, patterns and colors, plus user-supplied image textures.
dsh-zcode-cli-proxy
kyle123740/dsh-zcode-cli-proxy
ZCode CLI app-server relay for DeepSeek Harness, with GLM streaming, image input and session reuse using an existing Start Plan login; requires the ZCode client.
dsh-fal-imagegen
enchanted0911/dsh-fal-imagegen
fal.ai native image generation for DSH: a FAL_KEY settings card that follows DSH's language (zh/en, English default) plus Agent tools (fal_generate_image async / fal_edit_image / fal_get_image_task / fal_list_image_models) that call queue.fal.run directly
dsh-video-to-notes
yll-kb/dsh-video-to-notes
Opt-in DeepSeek Harness skill bundle that turns course, lecture, tutorial, documentary, meeting, and talk videos into structured study notes.
dsh-voice-danmaku
weizhida/dsh-voice-danmaku
这是dsh的插件,语音发送弹幕。在玩游戏时通过语音输入在b站发弹幕,不切出游戏可以正常操作
dsh-xiaozhi
toddpan/dsh-xiaozhi
Connects the Xiaozhi voice assistant to DSH Web over MCP: DSH is the tool provider, weaving 35 DSH Web endpoints into 16 voice-friendly tools for workspaces, sessions, chat, models, settings and files — outbound WebSocket to the Xiaozhi MCP access point by default (no public IP or port forwarding needed), several devices bound at once, plus a DSH settings page with live status.
sh-volume-knob
jianghu-lao-yao/sh-volume-knob
Speaker button beside the composer microphone — one click scrolls to the start of your newest question, marks it with a blinking caret and reads from there through the newest agent reply (dsh-tts, browser voice as fallback); press-and-drag-right picks any other reading start position on the page, press-and-drag-up opens a vertical mixer for in-page media volume and system output volume.
lookover (dsh-look)
mengxiaoxian/lookover
Scene-awareness probe for DSH on macOS: a privacy-first, pull-model `look` tool that reads the frontmost non-self window (app info, AX title/selection, gated local OCR), plus a summon-hotkey snapshot captured the instant you press.
dsh-agnes-gen
ylhow06/dsh-agnes-gen
Agnes AI image/video generation tools (agnes_image / agnes_video) for DSH (DeepSeek Harness). Built-in cross-process RPM rate limiting, 429 backoff and local ffmpeg GIF conversion.
dsh-plugin-notify
goodandready/dsh-plugin-notify
DSH plugin: audio chimes, cross-session toasts, desktop push, and IM webhooks for turn completion, errors, and approvals.
dsh-say
fangqian616/dsh-say
Speaks your agent's reports in a voice you choose, compressing long reports before speaking. Character voices need no training - a 3-10 second reference clip clones one, and community-trained models work too - and no 6.4 GB GPT-SoVITS install: the plugin installs its own runtime and voice. An existing GPT-SoVITS can be used instead, and it is the same voice model.
dsh-read-aloud
cccc12138/dsh-read-aloud
Adds a speaker button immediately right of the Like button on every finalized assistant reply; click it to hear the reply through the browser speech engine, and hover it for a speed and voice panel.
dsh-voice-alert
loyalchiiina/dsh-voice-alert
Speaks or plays a sound at the end of every turn and plays a failure cue when a tool call or a turn errors. Ships 20 built-in effects (10 alert chimes + 10 nature sounds) and defaults to effect mode, so it needs no API key and no audio files; an optional Volcengine voice-clone route generates the three announcement clips in one click. Plays through winmm/waveOut as-is, never changing the system volume or mute state. Windows only.
dsh-screen-eye
davidekingsss/dsh-screen-eye
Agent screen capture on macOS and Windows: one tool captures the screen and returns the image itself, so the model sees it without a second call. On Windows a whole burst runs in one engine call; on macOS a resident helper drops a region capture to 13ms and a change check to 23ms. A second tool, macOS only, reports whether Screen Recording is granted and opens the settings pane that fixes it.
vision-exp-tile
nicholaskin/vision-exp-tile
Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.
dsh-stt-plugin
zemanzhang809/dsh-stt-plugin
A speech-to-text (voice input) plugin for DeepSeek Harness.
dsh-asr-voice
bittersmilezzz/dsh-asr-voice
开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。