- Início
- Categorias
- Visão, voz e multimodal
Visão, voz e multimodal
Os plugins de Visão, voz e multimodal dão olhos, ouvidos e voz aos modelos DeepSeek somente texto no DeepSeek-Harness (dsh). Encontre ferramentas de visão e rotas de provedor que fazem OCR e descrevem capturas coladas via Zhipu GLM, Gemini, Doubao ou Ollama local, entrada de voz por microfone via Web Speech API do navegador ou APIs compatíveis com Whisper, leitura em voz alta com Edge TTS ou vozes personalizadas, modos de voz full-duplex, e geradores de imagens, vídeo e música.
86 plugins encontrados
modlens
liustack/modlens
Ponte de visão para modelos somente-texto: cole uma imagem e receba evidências estruturadas em JSON (OCR, layout, semântica).
dsh-vision-router
ysr666/dsh-vision-router
Visão gratuita para agentes somente-texto: cadeia de visão integrada sem chave, além de ferramentas de pixel (perguntas e respostas, grounding, recorte, diff de pixels, cores, OCR, rastreamento SVG, recorte de objetos, capturas de tela); cole uma imagem para usar.
dsh-vision-toolkit
anionex/dsh-vision-toolkit
Tarefas de visão para modelos somente-texto: perguntas e respostas sobre imagens com reconhecimento de intenção, OCR de capturas de tela longas, reprodução de UI, grounding e diff de pixels.
picturereader
jing-hy/picturereader
Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.
dsh-vision
linenxi-ctrl/dsh-vision
Plugin de visão externa para o DeepSeek Harness: painel de configuração via botão da baleia, reconhecimento de imagens com resposta automática e ferramentas de captura de tela/reconhecimento para o agente.
dsh-media-skills
mjorgin/dsh-media-skills
Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.
dsh-vision-proxy
flyvhidbwo/dsh-vision-proxy
Cérebro DeepSeek + transcrição automática de imagens: anexe imagens na GUI e cada uma é transcrita em texto por qualquer VLM compatível com OpenAI antes de chegar ao DeepSeek, que é somente texto — um caminho rápido com chave própria (padrão qwen3.7-flash; DashScope/Zhipu/OpenRouter ou qualquer endpoint compatível com OpenAI), ou Ollama local detectado automaticamente sem configuração.
dsh-vision-opencode
poiuyjie/dsh-vision-opencode
Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.
dsh-visual-plugin
jyh20030112/dsh-visual-plugin
Dsh-visual-plugin. Dê olhos ao seu modelo somente-texto: encaminhe imagens do usuário para qualquer modelo de visão compatível com OpenAI e veja os resultados em um painel à direita da Web UI
Gemini-Eyes
consolesun/gemini-eyes
MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.
dsh-plugin-tts
1624318455/dsh-plugin-tts
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
dsh-imagegen
dickpy/dsh-imagegen
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
dsh-chat-imagine
corrinehu/dsh-chat-imagine
Automatically generates and displays images in the DSH chat via API channels or local CLIs (supports mmx / codex / agy).
dsh-vision
54xkeee/dsh-vision
Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.
dsh-voice-input-plugin
zhangbo-cn/dsh-voice-input-plugin
Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.
dsh-deepseek-vision
siegfly/dsh-deepseek-vision
A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.
dsh-windows-ocr
maxwell-feng/dsh-windows-ocr
Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
dsh-voice
3274375092/dsh-voice
Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.
deepseek-vision (dsh-plugin-deepseek-vision)
gou-gee/deepseek-vision
Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.
dsh-guide-dog
atropinoltt/dsh-guide-dog
MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.
dsh-her-eyes
huashenglian/dsh-her-eyes
Um plugin dsh que permite à IA invocar automaticamente VLMs (modelos multimodais) para análise visual.
dsh-speak
alan2z/dsh-speak
Voice-announce the final reply on Windows (SAPI5 natural voices) and macOS (system voice); skips reasoning and tool calls, one-line npm install.
dsh-plugin-multimodal
shinjiyu/dsh-plugin-multimodal
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.
free-vision-skill
niyongsheng/free-vision-skill
Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.