Перейти к основному содержимому

Зрение, голос и мультимодальность

Плагины «Зрение, голос и мультимодальность» дают текстовым моделям DeepSeek в DeepSeek-Harness (dsh) глаза, уши и голос. Здесь есть инструменты зрения и провайдерские маршруты, которые распознают и описывают вставленные скриншоты через Zhipu GLM, Gemini, Doubao или локальный Ollama, голосовой ввод с микрофона через браузерный Web Speech API или Whisper-совместимые API, озвучивание ответов с помощью Edge TTS или собственных голосов, полнодуплексные голосовые режимы, а также генераторы изображений, видео и музыки.

Найдено плагинов: 86

D

dsh-web-ui (dsh-tool-describe-image)

damonkoy/dsh-web-ui

Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.

33 дня назадЗрение, голос и мультимодальностьApache-2.0
G

dsh-vision-bridge

gxx182/dsh-vision-bridge

Плагин DeepSeek Harness, связывающий изображения сессии с подключаемыми vision API, сохраняя DeepSeek в роли основной модели.

311 часов назадЗрение, голос и мультимодальностьMIT
E

dsh-llm-vision-bridge

einskyle/dsh-llm-vision-bridge

Нативный vision-мост провайдера LLM: изображения, вставленные в чат, описываются vision-моделью (Qwen3-VL через pi-ai/llama.cpp), а текстовое описание передаётся текстовой модели DeepSeek для ответа — приём изображений, маршрутизация и сжатие идут через нативные механизмы harness, с LRU-кэшем описаний и повтором при 503.

34 дня назадЗрение, голос и мультимодальностьMIT
X

dsh-vision-bridge

ximengxiaolan/dsh-vision-bridge

Изображения, прикреплённые в composer, распознаются в текст vision-моделью, совместимой с OpenAI, прежде чем попасть к текстовым моделям DeepSeek.

34 дня назадЗрение, голос и мультимодальностьMIT
Q

dsh-mic-input

qt-chen/dsh-mic-input

Голосовой ввод с микрофона для composer: транскрипция в реальном времени через Web Speech API браузера, дедупликация/автопродолжение, умная пунктуация, настройки языка и автоотправки.

3позавчераЗрение, голос и мультимодальностьMIT
S

dsh-voice

stardustlc666/dsh-voice

Voice pair: free edge-tts neural speech synthesis + OpenAI-compatible ASR transcription.

26 часов назадЗрение, голос и мультимодальностьMIT
N

voco-input-sh

nothree-code/voco-input-sh

Voice input for the Web UI: a mic button that drives local VocoType offline speech recognition and auto-inserts recognized text into the composer (auto-deploy, dedupe, continuous dictation).

220 часов назадЗрение, голос и мультимодальностьMIT
0

dsh-voice-input

0nt-one/dsh-voice-input

Mic button in the composer tool row: Web Speech API speech-to-text (Chrome/Edge), language switching, and optional auto-send, zero dependencies.

23 дня назадЗрение, голос и мультимодальностьMIT
X

dsh-draw-router

xiaozhe7772222/dsh-draw-router

Universal image generation for DeepSeek Harness: auto-discovers image models from any OpenAI-compatible endpoint (SenseNova, StepFun, Agnes, Qwen, Flux, SD, Imagen and more), with agent tools and REST API.

24 часа назадЗрение, голос и мультимодальностьMIT
X

dsh-vision

xiaoshihou514/dsh-vision

Native vision capability extension, using either Zhipu (free) or Qwen-VL (local).

2позавчераЗрение, голос и мультимодальностьMIT
W

visual-review

wang-bool/visual-review

Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.

216 часов назадЗрение, голос и мультимодальностьMIT
N

dsh-llm-deepseek-vision

nagasakisoyo-ui/dsh-llm-deepseek-vision

Vision-augmented DeepSeek adapter: a vision-capable model describes image input, then a text-only DeepSeek model reasons over the description.

24 дня назадЗрение, голос и мультимодальностьMIT
M

dsh-vision-tools

moon09300731/dsh-vision-tools

Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.

2вчераЗрение, голос и мультимодальностьMIT
M

dsh-tesseract-ocr

maxwell-feng/dsh-tesseract-ocr

Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

2вчераЗрение, голос и мультимодальностьMIT
L

dsh-image-gen

leemancheung/dsh-image-gen

GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.

23 дня назадЗрение, голос и мультимодальностьMIT
H

dsh-open-eyes

hyp6666/dsh-open-eyes

Vision bridge for text-only DeepSeek routes that analyzes attached and local images through configurable OpenAI Responses, Chat Completions, or Anthropic Messages endpoints while leaving image-capable routes native.

2позавчераЗрение, голос и мультимодальностьMIT
H

dsh-vision-mix

haiziyao/dsh-vision-mix

Combine text, vision, and image-generation APIs into one Mix model with automatic routing: text-only requests go to the chat model, user images and agent screenshots go to the vision model, follow-ups keep using the same session image, and agents can generate or edit images with session-scoped call history.

23 дня назадЗрение, голос и мультимодальностьMIT
F

dsh-free-vision

fuzzysoul/dsh-free-vision

Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.

27 часов назадЗрение, голос и мультимодальностьMIT
P

dsh-voice-call

pandapolo/dsh-voice-call

Agent-initiated voice calls: `offer_call` rings the human (接听/拒接/稍后再说); accepted calls synthesize and play locally via CrispASR + Qwen3-TTS (9 speakers, 2 Chinese dialects), rejected calls return the decision to the agent.

1вчераЗрение, голос и мультимодальностьMIT
F

dsh-chatvoice

fuzzysoul/dsh-chatvoice

Free voice closed loop for the Web UI: browser SpeechRecognition mic input with live interim results plus read-aloud speaker buttons and auto-read for assistant replies, zero configuration and no API key.

13 дня назадЗрение, голос и мультимодальностьMIT
B

dsh-stt-input

baisama-cloud/dsh-stt-input

Speech-to-text voice input for the web UI: a mic button in the composer transcribes speech into the draft via the browser Web Speech API (zero-config) or an OpenAI-compatible Whisper API (OpenAI / Groq), with a selectable model and language in Settings.

1вчераЗрение, голос и мультимодальностьMIT
X

dsh-image-vision

xsoc1/dsh-image-vision

Chat image-attachment bridge with a `view_image` tool for any OpenAI-compatible VLM (local Ollama or cloud): pasted/dropped images become `view_image` path markers before reaching text-only DeepSeek models.

13 дня назадЗрение, голос и мультимодальность
W

mimo-vision

wulusai2333/mimo-vision

`describe_image` tool: a vision bridge that sends images to mimo-v2.5 through the opencode Zen API (credential `OPENCODE_GO_API_KEY`, free route first with paid fallback) and returns text descriptions for text-only models, with native passthrough and ImageMagick transcoding of SVG/TIFF/HEIC formats.

1позавчераЗрение, голос и мультимодальностьMIT
W

dsh-tool-vision

wanshichenguang/dsh-tool-vision

Adds an image_describe tool backed by the DashScope OpenAI-compatible vision API; on text-only sessions, pasted images are stored as local paths for the model and rendered inline in the chat transcript.

1вчераЗрение, голос и мультимодальностьMIT