Перейти к основному содержимому

Зрение, голос и мультимодальность

Плагины «Зрение, голос и мультимодальность» дают текстовым моделям DeepSeek в DeepSeek-Harness (dsh) глаза, уши и голос. Здесь есть инструменты зрения и провайдерские маршруты, которые распознают и описывают вставленные скриншоты через Zhipu GLM, Gemini, Doubao или локальный Ollama, голосовой ввод с микрофона через браузерный Web Speech API или Whisper-совместимые API, озвучивание ответов с помощью Edge TTS или собственных голосов, полнодуплексные голосовые режимы, а также генераторы изображений, видео и музыки.

Найдено плагинов: 86

B

dsh-voice-ai-girlfriend-plugin

beiyege-01/dsh-voice-ai-girlfriend-plugin

Voice AI girlfriend for the Web UI: FunASR mic input, Qwen3-TTS spoken replies, companion animation window, and two-way QQ chat (text/voice/image push) via NapCat.

0позавчераЗрение, голос и мультимодальность
T

dsh-plugin-appshot

tauruswood/dsh-plugin-appshot

Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the composer for agent queries.

0вчераЗрение, голос и мультимодальностьMIT
S

dsh-easyvision

s3yf1337/dsh-easyvision

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

0позавчераЗрение, голос и мультимодальностьMIT
R

dsh-vision-subagent

ruby1304/dsh-vision-subagent

Vision for any DSH route: paste images in the Web composer with intent-aware auto-analysis, delegate workspace image reads to a Kimi/MiniMax vision subagent, and materialize pasted originals for editing.

011 часов назадЗрение, голос и мультимодальностьMIT
N

dsh-subagent-vision

niuniuaba/dsh-subagent-vision

Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.

015 часов назадЗрение, голос и мультимодальностьMIT
M

dsh-unsloth-hands

microherox/dsh-unsloth-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.

016 часов назадЗрение, голос и мультимодальностьMIT
L

dsh-vision-plugin

ld-1101/dsh-vision-plugin

Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).

0позавчераЗрение, голос и мультимодальностьMIT
K

dsh-vision-recognizer

kaixinbaba/dsh-vision-recognizer

Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.

03 дня назадЗрение, голос и мультимодальностьMIT
I

dsh-quicksight

isanti2016/dsh-quicksight

Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).

0позавчераЗрение, голос и мультимодальностьMIT
G

dsh-vision-guard

good-boy4069/dsh-vision-guard

Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.

03 дня назадЗрение, голос и мультимодальностьMIT
E

dsh-plugin-image-input

elohia/dsh-plugin-image-input

Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).

04 дня назадЗрение, голос и мультимодальностьMIT
B

dsh-vision-solution

br1nosense/dsh-vision-solution

Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.

0вчераЗрение, голос и мультимодальность
S

dsh-nanobananapro

synmindai/dsh-nanobananapro

Генерация изображений и видео в DeepSeek Harness через API NanoBananaPro

04 дня назадЗрение, голос и мультимодальностьMIT
S

dsh-seedance2

synmindai/dsh-seedance2

Генерация изображений и видео Seedance в DeepSeek Harness через API Seedance 2 AI

04 дня назадЗрение, голос и мультимодальностьMIT