Passer au contenu principal

Vision, voix et multimodal

Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.

312 plugins trouvés

Y

dsh-plugin-vision-toolkit

yytbit/dsh-plugin-vision-toolkit

Boîte à outils de vision pour DeepSeek Harness -- outils CLI glance, ground, detect, crop pour que les agents texte seul comprennent les images.

1il y a 2 moisVision, voix et multimodalMIT
P

dsh-yali-image-generator

pptt121212/dsh-yali-image-generator

Plugin de génération d'images DeepSeek-Harness. Demandez une clé API Yali AI : https://api.yaliai.com/

1il y a 2 moisVision, voix et multimodalMIT
C

deepsee

chang416/deepsee

DeepSee : vision DeepSeek Harness, routage multi-modèles, et auto-vérifications visuelles Gemini avant livraison.

1il y a 2 moisVision, voix et multimodalMIT
F

dsh-sight

fu3rte/dsh-sight

Vision en plugin pour les modèles DeepSeek Harness (dsh) texte seul : un outil `vision` avec des préréglages VLM gratuits/économiques intégrés, analyse par lot multi-images, admission d'image par indice de collage, et une page de réglages web avec rechargement à chaud.

1il y a 2 moisVision, voix et multimodalMIT
W

dsh-friend

wanghehe123/dsh-friend

Plugin de compagnon personnifié pour DeepSeek Harness : fiches de personnage, voix, Live2D, mémoire locale et accompagnement pendant le travail.

1il y a 2 moisVision, voix et multimodalMIT
A

dsh-voice-webspeech

anweat/dsh-voice-webspeech

Saisie vocale via la Web Speech API du navigateur : zéro serveur, zéro clé, zéro téléchargement de modèle (Edge=Azure, Chrome=Google speech).

1le mois dernierVision, voix et multimodalMIT
R

dsh-plugins

retiredphysicist/dsh-plugins

DeepSeek Harness plugin: web browsing tools (markdown / screenshot / pdf / crawl) powered by Cloudflare Browser Run — real headless Chrome with JS rendering, login sessions and WebMCP support.

0avant-hierVision, voix et multimodal
W

dsh-kite

weibaohui/dsh-kite

Animated kite that responds to agent activity, with configurable kite frames, patterns and colors, plus user-supplied image textures.

0il y a 4 joursVision, voix et multimodal
K

dsh-zcode-cli-proxy

kyle123740/dsh-zcode-cli-proxy

ZCode CLI app-server relay for DeepSeek Harness, with GLM streaming, image input and session reuse using an existing Start Plan login; requires the ZCode client.

0il y a 3 joursVision, voix et multimodalAGPL-3.0
E

dsh-fal-imagegen

enchanted0911/dsh-fal-imagegen

fal.ai native image generation for DSH: a FAL_KEY settings card that follows DSH's language (zh/en, English default) plus Agent tools (fal_generate_image async / fal_edit_image / fal_get_image_task / fal_list_image_models) that call queue.fal.run directly

0il y a 9 joursVision, voix et multimodalMIT
Y

dsh-video-to-notes

yll-kb/dsh-video-to-notes

Opt-in DeepSeek Harness skill bundle that turns course, lecture, tutorial, documentary, meeting, and talk videos into structured study notes.

0il y a 12 joursVision, voix et multimodalMIT
W

dsh-voice-danmaku

weizhida/dsh-voice-danmaku

这是dsh的插件,语音发送弹幕。在玩游戏时通过语音输入在b站发弹幕,不切出游戏可以正常操作

0il y a 4 joursVision, voix et multimodalMIT
T

dsh-xiaozhi

toddpan/dsh-xiaozhi

Connects the Xiaozhi voice assistant to DSH Web over MCP: DSH is the tool provider, weaving 35 DSH Web endpoints into 16 voice-friendly tools for workspaces, sessions, chat, models, settings and files — outbound WebSocket to the Xiaozhi MCP access point by default (no public IP or port forwarding needed), several devices bound at once, plus a DSH settings page with live status.

0il y a 6 joursVision, voix et multimodal
J

sh-volume-knob

jianghu-lao-yao/sh-volume-knob

Speaker button beside the composer microphone — one click scrolls to the start of your newest question, marks it with a blinking caret and reads from there through the newest agent reply (dsh-tts, browser voice as fallback); press-and-drag-right picks any other reading start position on the page, press-and-drag-up opens a vertical mixer for in-page media volume and system output volume.

0il y a 8 joursVision, voix et multimodalMIT
M

lookover (dsh-look)

mengxiaoxian/lookover

Scene-awareness probe for DSH on macOS: a privacy-first, pull-model `look` tool that reads the frontmost non-self window (app info, AX title/selection, gated local OCR), plus a summon-hotkey snapshot captured the instant you press.

0il y a 11 joursVision, voix et multimodalMIT
Y

dsh-agnes-gen

ylhow06/dsh-agnes-gen

Agnes AI image/video generation tools (agnes_image / agnes_video) for DSH (DeepSeek Harness). Built-in cross-process RPM rate limiting, 429 backoff and local ffmpeg GIF conversion.

0il y a 11 joursVision, voix et multimodalMIT
G

dsh-plugin-notify

goodandready/dsh-plugin-notify

DSH plugin: audio chimes, cross-session toasts, desktop push, and IM webhooks for turn completion, errors, and approvals.

0il y a 11 joursVision, voix et multimodalMIT
F

dsh-say

fangqian616/dsh-say

Speaks your agent's reports in a voice you choose, compressing long reports before speaking. Character voices need no training - a 3-10 second reference clip clones one, and community-trained models work too - and no 6.4 GB GPT-SoVITS install: the plugin installs its own runtime and voice. An existing GPT-SoVITS can be used instead, and it is the same voice model.

0il y a 12 joursVision, voix et multimodal
C

dsh-read-aloud

cccc12138/dsh-read-aloud

Adds a speaker button immediately right of the Like button on every finalized assistant reply; click it to hear the reply through the browser speech engine, and hover it for a speed and voice panel.

0il y a 15 joursVision, voix et multimodalMIT
L

dsh-voice-alert

loyalchiiina/dsh-voice-alert

Speaks or plays a sound at the end of every turn and plays a failure cue when a tool call or a turn errors. Ships 20 built-in effects (10 alert chimes + 10 nature sounds) and defaults to effect mode, so it needs no API key and no audio files; an optional Volcengine voice-clone route generates the three announcement clips in one click. Plays through winmm/waveOut as-is, never changing the system volume or mute state. Windows only.

0il y a 4 joursVision, voix et multimodalMIT
D

dsh-screen-eye

davidekingsss/dsh-screen-eye

Agent screen capture on macOS and Windows: one tool captures the screen and returns the image itself, so the model sees it without a second call. On Windows a whole burst runs in one engine call; on macOS a resident helper drops a region capture to 13ms and a change check to 23ms. A second tool, macOS only, reports whether Screen Recording is granted and opens the settings pane that fixes it.

0il y a 14 joursVision, voix et multimodalMIT
N

vision-exp-tile

nicholaskin/vision-exp-tile

Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.

0il y a 22 joursVision, voix et multimodalMIT
Z

dsh-stt-plugin

zemanzhang809/dsh-stt-plugin

A speech-to-text (voice input) plugin for DeepSeek Harness.

0il y a 15 joursVision, voix et multimodalMIT
B

dsh-asr-voice

bittersmilezzz/dsh-asr-voice

开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。

0il y a 18 joursVision, voix et multimodalMIT