Зрение, голос и мультимодальность
Плагины «Зрение, голос и мультимодальность» дают текстовым моделям DeepSeek в DeepSeek-Harness (dsh) глаза, уши и голос. Здесь есть инструменты зрения и провайдерские маршруты, которые распознают и описывают вставленные скриншоты через Zhipu GLM, Gemini, Doubao или локальный Ollama, голосовой ввод с микрофона через браузерный Web Speech API или Whisper-совместимые API, озвучивание ответов с помощью Edge TTS или собственных голосов, полнодуплексные голосовые режимы, а также генераторы изображений, видео и музыки.
Найдено плагинов: 312
dsh-fal-image-gen
goodandready/dsh-fal-image-gen
FAL image generation for DeepSeek Harness: a generate_image tool backed by the FAL queue API (default model fal-ai/flux-2/klein/9b). Images render directly in the conversation, are saved to the workspace, and all settings are editable from the Web GUI.
dsh-bundle-vision
skillre/dsh-bundle-vision
Vision bundle + plugin for DeepSeek Harness: the describe_image tool reads local images and asks any configured multimodal route, with zero core changes
dsh-mindsee
123cdxcc/dsh-mindsee
DeepSeek Harness 插件:以 MindSee 为后端,为 DeepSeek 提供图片相关能力
dsh-voice-input
opensquad-ai/dsh-voice-input
SenseVoice 语音输入插件 for DeepSeek Harness:在对话输入框旁添加麦克风按钮,录音后调用本地 SenseVoice 服务转成文本填入输入框。首次使用自动下载模型并显示进度,后端由插件自动启动。
dsh-gemini-multimodal
realalexandreai/dsh-gemini-multimodal
DeepSeek Harness plugin: multimodal tools (image/audio/video/document understanding, transcription, image generation) via Gemini API or the local Antigravity CLI.
dsh-tool-image-gen
zhangjunjesse/dsh-tool-image-gen
DSH tool plugin: generate images through ToAPIs async GPT-Image-2 API (submit task, poll, download).
soyo
ottohere-mourn/soyo
DSH-native video understanding with configurable multimodal providers
dsh-voice-live
tangzheng202202/dsh-voice-live
Real-time duplex voice over Volcengine streaming ASR/TTS: agent reply narration, barge-in, wake word, live captions, 30 Chinese voices and a reply-first acknowledgment; builds in the DSH monorepo.
dsh-voice-input
newdanew/dsh-voice-input
Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.
dsh-easyvision
s3yf1337/dsh-easyvision
Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.
dsh-subagent-vision
niuniuaba/dsh-subagent-vision
Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.
dsh-unsloth-hands
microherox/dsh-unsloth-hands
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.
dsh-koboldcpp-hands
microherox/dsh-koboldcpp-hands
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.
dsh-windows-ocr
maxwell-feng/dsh-windows-ocr
Локальное распознавание текста для прикреплённых изображений через встроенный движок Windows (Windows.Media.Ocr): модели отправляется только распознанный текст, а не байты изображения; сквозная передача изображения включается по желанию (opt-in).
dsh-tesseract-ocr
maxwell-feng/dsh-tesseract-ocr
Локальный OCR для прикреплённых изображений через Tesseract: в модель отправляется только распознанный текст, а не байты изображения; vision passthrough включается только по желанию.
dsh-vision-plugin
ld-1101/dsh-vision-plugin
Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).
dsh-quicksight
isanti2016/dsh-quicksight
Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).
dsh-vision-guard
good-boy4069/dsh-vision-guard
Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.
dsh-plugin-image-input
elohia/dsh-plugin-image-input
Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).
dsh-vision-solution
br1nosense/dsh-vision-solution
Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.
dsh-plugin-vision
tdf1995/dsh-plugin-vision
Зрение для текстовых LLM: описание изображений / OCR / VQA через бесплатные vision API Gemini и GLM.
vision-tool
haowencang/vision-tool
Плагин визуального ревью UI/UX в формате беседы для DeepSeek Harness: инструменты vision_review / vision_ask на основе любого совместимого с OpenAI мультимодального эндпоинта (например, agnes-2.5-flash)
dsh-nanobananapro
synmindai/dsh-nanobananapro
Генерация изображений и видео в DeepSeek Harness через API NanoBananaPro
dsh-seedance2
synmindai/dsh-seedance2
Генерация изображений и видео Seedance в DeepSeek Harness через API Seedance 2 AI