Passer au contenu principal

Vision, voix et multimodal

Les plugins Vision, voix et multimodal donnent des yeux, des oreilles et une voix aux modèles DeepSeek texte seul de DeepSeek-Harness (dsh). Trouvez des outils de vision et des routes de fournisseur qui font l'OCR et décrivent les captures collées via Zhipu GLM, Gemini, Doubao ou Ollama en local, une saisie vocale au micro via la Web Speech API du navigateur ou des API compatibles Whisper, une lecture à voix haute avec Edge TTS ou des voix personnalisées, des modes vocaux full-duplex, et des générateurs d'images, de vidéos et de musique.

312 plugins trouvés

G

dsh-fal-image-gen

goodandready/dsh-fal-image-gen

FAL image generation for DeepSeek Harness: a generate_image tool backed by the FAL queue API (default model fal-ai/flux-2/klein/9b). Images render directly in the conversation, are saved to the workspace, and all settings are editable from the Web GUI.

0il y a 2 moisVision, voix et multimodalMIT
S

dsh-bundle-vision

skillre/dsh-bundle-vision

Vision bundle + plugin for DeepSeek Harness: the describe_image tool reads local images and asks any configured multimodal route, with zero core changes

0il y a 2 moisVision, voix et multimodalMIT
1

dsh-mindsee

123cdxcc/dsh-mindsee

DeepSeek Harness 插件:以 MindSee 为后端,为 DeepSeek 提供图片相关能力

0il y a 2 moisVision, voix et multimodalMIT
O

dsh-voice-input

opensquad-ai/dsh-voice-input

SenseVoice 语音输入插件 for DeepSeek Harness:在对话输入框旁添加麦克风按钮,录音后调用本地 SenseVoice 服务转成文本填入输入框。首次使用自动下载模型并显示进度,后端由插件自动启动。

0il y a 2 moisVision, voix et multimodalMIT
R

dsh-gemini-multimodal

realalexandreai/dsh-gemini-multimodal

DeepSeek Harness plugin: multimodal tools (image/audio/video/document understanding, transcription, image generation) via Gemini API or the local Antigravity CLI.

0il y a 2 moisVision, voix et multimodalMIT
Z

dsh-tool-image-gen

zhangjunjesse/dsh-tool-image-gen

DSH tool plugin: generate images through ToAPIs async GPT-Image-2 API (submit task, poll, download).

0il y a 2 moisVision, voix et multimodalMIT
O

soyo

ottohere-mourn/soyo

DSH-native video understanding with configurable multimodal providers

0il y a 2 moisVision, voix et multimodalMIT
T

dsh-voice-live

tangzheng202202/dsh-voice-live

Real-time duplex voice over Volcengine streaming ASR/TTS: agent reply narration, barge-in, wake word, live captions, 30 Chinese voices and a reply-first acknowledgment; builds in the DSH monorepo.

0avant-hierVision, voix et multimodalMIT
N

dsh-voice-input

newdanew/dsh-voice-input

Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.

0il y a 2 moisVision, voix et multimodalMIT
S

dsh-easyvision

s3yf1337/dsh-easyvision

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

0il y a 2 moisVision, voix et multimodalMIT
N

dsh-subagent-vision

niuniuaba/dsh-subagent-vision

Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.

0il y a 17 joursVision, voix et multimodalMIT
M

dsh-unsloth-hands

microherox/dsh-unsloth-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.

0le mois dernierVision, voix et multimodalMIT
M

dsh-koboldcpp-hands

microherox/dsh-koboldcpp-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.

0le mois dernierVision, voix et multimodalMIT
M

dsh-windows-ocr

maxwell-feng/dsh-windows-ocr

OCR local pour les images jointes via le moteur intégré de Windows (Windows.Media.Ocr) : seul le texte reconnu est envoyé au modèle, jamais les octets de l'image ; le passthrough de vision est optionnel (opt-in).

0Vision, voix et multimodalPeut-être obsolète
M

dsh-tesseract-ocr

maxwell-feng/dsh-tesseract-ocr

OCR locale pour les images jointes via Tesseract : seul le texte reconnu est envoyé au modèle, jamais les octets de l’image ; le passage direct de la vision est facultatif.

0Vision, voix et multimodalPeut-être obsolète
L

dsh-vision-plugin

ld-1101/dsh-vision-plugin

Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).

0il y a 2 moisVision, voix et multimodalMIT
I

dsh-quicksight

isanti2016/dsh-quicksight

Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).

0il y a 2 moisVision, voix et multimodalMIT
G

dsh-vision-guard

good-boy4069/dsh-vision-guard

Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.

0il y a 2 moisVision, voix et multimodalMIT
E

dsh-plugin-image-input

elohia/dsh-plugin-image-input

Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).

0il y a 2 moisVision, voix et multimodalMIT
B

dsh-vision-solution

br1nosense/dsh-vision-solution

Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.

0il y a 2 moisVision, voix et multimodal
T

dsh-plugin-vision

tdf1995/dsh-plugin-vision

Vision pour les LLM texte seul : description d'image / OCR / VQA via les API de vision gratuites Gemini et GLM.

0il y a 2 moisVision, voix et multimodalMIT
H

vision-tool

haowencang/vision-tool

Plugin de revue visuelle UI/UX conversationnel pour DeepSeek Harness : outils vision_review / vision_ask adossés à n'importe quel endpoint multimodal compatible OpenAI (par ex. agnes-2.5-flash).

0il y a 2 moisVision, voix et multimodalMIT
S

dsh-nanobananapro

synmindai/dsh-nanobananapro

Générez des images et des vidéos dans DeepSeek Harness via l'API NanoBananaPro.

0il y a 2 moisVision, voix et multimodalMIT
S

dsh-seedance2

synmindai/dsh-seedance2

Générez des images et des vidéos Seedance dans DeepSeek Harness via l'API Seedance 2 AI.

0il y a 2 moisVision, voix et multimodalMIT