Vai al contenuto principale

Visione, voce e multimodale

I plugin Visione, voce e multimodale danno occhi, orecchie e voce ai modelli DeepSeek solo testo di DeepSeek-Harness (dsh). Trova strumenti di visione e route di provider che eseguono OCR e descrivono gli screenshot incollati tramite Zhipu GLM, Gemini, Doubao o Ollama in locale, input vocale dal microfono tramite la Web Speech API del browser o API compatibili con Whisper, lettura ad alta voce con Edge TTS o voci personalizzate, modalità vocali full-duplex, e generatori di immagini, video e musica.

312 plugin trovati

Y

dsh-plugin-vision-toolkit

yytbit/dsh-plugin-vision-toolkit

Toolkit di visione per DeepSeek Harness -- strumenti CLI glance, ground, detect, crop per far comprendere le immagini agli agenti solo testo

12 mesi faVisione, voce e multimodaleMIT
P

dsh-yali-image-generator

pptt121212/dsh-yali-image-generator

Plugin di generazione immagini per DeepSeek-Harness. Richiedi una API Key di Yali AI su: https://api.yaliai.com/

12 mesi faVisione, voce e multimodaleMIT
C

deepsee

chang416/deepsee

DeepSee: visione per DeepSeek Harness, instradamento multi-modello e autoverifiche visive con Gemini prima della consegna

12 mesi faVisione, voce e multimodaleMIT
F

dsh-sight

fu3rte/dsh-sight

Visione plug-in per i modelli DeepSeek Harness (dsh) solo testo: uno strumento `vision` con preset VLM economici/gratuiti integrati, analisi in batch di più immagini, ammissione delle immagini tramite incolla-per-suggerire e una pagina di impostazioni web con hot-reload

12 mesi faVisione, voce e multimodaleMIT
W

dsh-friend

wanghehe123/dsh-friend

Plugin companion personificato per DeepSeek Harness: schede personaggio, voce, Live2D, memoria locale e compagnia durante il lavoro.

12 mesi faVisione, voce e multimodaleMIT
A

dsh-voice-webspeech

anweat/dsh-voice-webspeech

Input vocale tramite Web Speech API del browser: zero server, zero chiavi, zero download di modelli (Edge=Azure, Chrome=Google speech)

1mese scorsoVisione, voce e multimodaleMIT
R

dsh-plugins

retiredphysicist/dsh-plugins

DeepSeek Harness plugin: web browsing tools (markdown / screenshot / pdf / crawl) powered by Cloudflare Browser Run — real headless Chrome with JS rendering, login sessions and WebMCP support.

0l’altro ieriVisione, voce e multimodale
W

dsh-kite

weibaohui/dsh-kite

Animated kite that responds to agent activity, with configurable kite frames, patterns and colors, plus user-supplied image textures.

04 giorni faVisione, voce e multimodale
K

dsh-zcode-cli-proxy

kyle123740/dsh-zcode-cli-proxy

ZCode CLI app-server relay for DeepSeek Harness, with GLM streaming, image input and session reuse using an existing Start Plan login; requires the ZCode client.

03 giorni faVisione, voce e multimodaleAGPL-3.0
E

dsh-fal-imagegen

enchanted0911/dsh-fal-imagegen

fal.ai native image generation for DSH: a FAL_KEY settings card that follows DSH's language (zh/en, English default) plus Agent tools (fal_generate_image async / fal_edit_image / fal_get_image_task / fal_list_image_models) that call queue.fal.run directly

09 giorni faVisione, voce e multimodaleMIT
Y

dsh-video-to-notes

yll-kb/dsh-video-to-notes

Opt-in DeepSeek Harness skill bundle that turns course, lecture, tutorial, documentary, meeting, and talk videos into structured study notes.

012 giorni faVisione, voce e multimodaleMIT
W

dsh-voice-danmaku

weizhida/dsh-voice-danmaku

这是dsh的插件,语音发送弹幕。在玩游戏时通过语音输入在b站发弹幕,不切出游戏可以正常操作

04 giorni faVisione, voce e multimodaleMIT
T

dsh-xiaozhi

toddpan/dsh-xiaozhi

Connects the Xiaozhi voice assistant to DSH Web over MCP: DSH is the tool provider, weaving 35 DSH Web endpoints into 16 voice-friendly tools for workspaces, sessions, chat, models, settings and files — outbound WebSocket to the Xiaozhi MCP access point by default (no public IP or port forwarding needed), several devices bound at once, plus a DSH settings page with live status.

06 giorni faVisione, voce e multimodale
J

sh-volume-knob

jianghu-lao-yao/sh-volume-knob

Speaker button beside the composer microphone — one click scrolls to the start of your newest question, marks it with a blinking caret and reads from there through the newest agent reply (dsh-tts, browser voice as fallback); press-and-drag-right picks any other reading start position on the page, press-and-drag-up opens a vertical mixer for in-page media volume and system output volume.

08 giorni faVisione, voce e multimodaleMIT
M

lookover (dsh-look)

mengxiaoxian/lookover

Scene-awareness probe for DSH on macOS: a privacy-first, pull-model `look` tool that reads the frontmost non-self window (app info, AX title/selection, gated local OCR), plus a summon-hotkey snapshot captured the instant you press.

011 giorni faVisione, voce e multimodaleMIT
Y

dsh-agnes-gen

ylhow06/dsh-agnes-gen

Agnes AI image/video generation tools (agnes_image / agnes_video) for DSH (DeepSeek Harness). Built-in cross-process RPM rate limiting, 429 backoff and local ffmpeg GIF conversion.

011 giorni faVisione, voce e multimodaleMIT
G

dsh-plugin-notify

goodandready/dsh-plugin-notify

DSH plugin: audio chimes, cross-session toasts, desktop push, and IM webhooks for turn completion, errors, and approvals.

011 giorni faVisione, voce e multimodaleMIT
F

dsh-say

fangqian616/dsh-say

Speaks your agent's reports in a voice you choose, compressing long reports before speaking. Character voices need no training - a 3-10 second reference clip clones one, and community-trained models work too - and no 6.4 GB GPT-SoVITS install: the plugin installs its own runtime and voice. An existing GPT-SoVITS can be used instead, and it is the same voice model.

012 giorni faVisione, voce e multimodale
C

dsh-read-aloud

cccc12138/dsh-read-aloud

Adds a speaker button immediately right of the Like button on every finalized assistant reply; click it to hear the reply through the browser speech engine, and hover it for a speed and voice panel.

015 giorni faVisione, voce e multimodaleMIT
L

dsh-voice-alert

loyalchiiina/dsh-voice-alert

Speaks or plays a sound at the end of every turn and plays a failure cue when a tool call or a turn errors. Ships 20 built-in effects (10 alert chimes + 10 nature sounds) and defaults to effect mode, so it needs no API key and no audio files; an optional Volcengine voice-clone route generates the three announcement clips in one click. Plays through winmm/waveOut as-is, never changing the system volume or mute state. Windows only.

04 giorni faVisione, voce e multimodaleMIT
D

dsh-screen-eye

davidekingsss/dsh-screen-eye

Agent screen capture on macOS and Windows: one tool captures the screen and returns the image itself, so the model sees it without a second call. On Windows a whole burst runs in one engine call; on macOS a resident helper drops a region capture to 13ms and a change check to 23ms. A second tool, macOS only, reports whether Screen Recording is granted and opens the settings pane that fixes it.

014 giorni faVisione, voce e multimodaleMIT
N

vision-exp-tile

nicholaskin/vision-exp-tile

Large-image recognition for vision-exp models: lossless 800×800 tile recognition (smart/pipeline/full), local OCR with preprocessing & handwriting routing, optional multi-vendor GPU (DirectML/CUDA/OpenVINO) with auto CPU fallback.

022 giorni faVisione, voce e multimodaleMIT
Z

dsh-stt-plugin

zemanzhang809/dsh-stt-plugin

A speech-to-text (voice input) plugin for DeepSeek Harness.

015 giorni faVisione, voce e multimodaleMIT
B

dsh-asr-voice

bittersmilezzz/dsh-asr-voice

开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。

018 giorni faVisione, voce e multimodaleMIT