Pular para o conteúdo principal

Visão, voz e multimodal

Os plugins de Visão, voz e multimodal dão olhos, ouvidos e voz aos modelos DeepSeek somente texto no DeepSeek-Harness (dsh). Encontre ferramentas de visão e rotas de provedor que fazem OCR e descrevem capturas coladas via Zhipu GLM, Gemini, Doubao ou Ollama local, entrada de voz por microfone via Web Speech API do navegador ou APIs compatíveis com Whisper, leitura em voz alta com Edge TTS ou vozes personalizadas, modos de voz full-duplex, e geradores de imagens, vídeo e música.

86 plugins encontrados

T

dsh-plugins (dsh-vision)

tzhr-invest/dsh-plugins

Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.

1ontemVisão, voz e multimodalMIT
P

dsh-screenshot

paicat1/dsh-screenshot

Standalone screen capture for DeepSeek Harness (dsh). Browser hotkeys for instant capture plus an agent-facing capture and read tool that lets the agent see and analyze any screen region.

1ontemVisão, voz e multimodalMIT
N

dsh-auto-vision

normanfxxkingrockwell/dsh-auto-vision

Auto-discovery vision bridge for text-only DeepSeek Harness agents: automatically finds an image-capable model from your configured providers and returns picture descriptions as plain text via a vision tool.

1há 12 horasVisão, voz e multimodalMIT
M

dsh-koboldcpp-hands

microherox/dsh-koboldcpp-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.

1há 17 horasVisão, voz e multimodalMIT
L

dsh-plugin-grok2api-media-tool

lsjspl/dsh-plugin-grok2api-media-tool

Gives dsh the ability to generate images and videos through the grok2api API.

1há 10 horasVisão, voz e multimodal
L

dsh-eyes

leeminjing/dsh-eyes

On-demand vision for text-only DeepSeek models: upload images, and the model calls a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).

1anteontemVisão, voz e multimodalMIT
K

dsh-mindseye

kanchengw/dsh-mindseye

Vision plugin for text-only DeepSeek Harness models: native image paste, layered evidence memory and cache, and intent-driven tool selection.

1há 5 horasVisão, voz e multimodalMIT
G

dsh-tool-vision

gloryxpnv/dsh-tool-vision

Local-first structured vision for text-only agents: images go to a local OpenAI-compatible VLM and come back as JSON evidence (summary, verbatim OCR, layout regions, entities/relations, colors, explicit uncertainty), with anti-hallucination fallback and an optional paste/upload bridge; zero cloud cost, images never leave the machine.

1há 3 diasVisão, voz e multimodalMIT
E

dsh-plugin-mm-vision

elohia/dsh-plugin-mm-vision

Synesthesia Encoder for DSH: a vision model translates images into compact structured spatial text (canvas/elements/percentage coordinates), giving text-only LLMs pixel-level image understanding via the `mm_vision` tool.

1há 5 diasVisão, voz e multimodalMIT
C

dsh-deepseek-vision

cheng-cheng9669/dsh-deepseek-vision

Reuses DeepSeek web's built-in vision mode for text-only models: the deepseek_vision tool drives the local deepseek-vision-cli browser automation (manual login helper, deep-think enabled, auto-closes browser) and returns image descriptions as text.

1anteontemVisão, voz e multimodalMIT
5

dsh-youreyes

54xkeee/dsh-youreyes

Vision toolkit for text-only DeepSeek: model-invokable `vision` tool, wrapper adapters for deepseek/opencode-go (v4 flash/pro), Antigravity IDE quota (default, flash/pro) / any OpenAI-compatible VLM / Gemini / local Ollama channels, evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

1há 13 horasVisão, voz e multimodalMIT
3

dsh-vision (vision-tool)

314857493/dsh-vision

Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.

1ontemVisão, voz e multimodalMIT
3

dsh-vision (vision-route)

314857493/dsh-vision

Registers a `deepseek-vision` provider route: the Web GUI accepts pasted images and transcribes them to text via the free Zhipu GLM vision API before delegating to the DeepSeek adapter.

1ontemVisão, voz e multimodalMIT
1

dsh-vision-fallback

1helloman1/dsh-vision-fallback

Routes chat images to a fixed OpenAI-compatible vision model, returns factual observations to the selected main model, and reuses session-scoped observations across replay, compaction, and restarts.

1há 12 horasVisão, voz e multimodalMIT
Z

dsh-voice

zhuiyueya/dsh-voice

Voz para o DeepSeek Harness (dsh) — entrada por fala para texto + leitura em voz alta via TTS para o DeepSeek somente de texto, sem necessidade de chave de API

1há 5 diasVisão, voz e multimodalMIT
S

multimodal-bridge

spirit4471/multimodal-bridge

Pacote de plugins do DeepSeek Harness: ferramentas qwen_vision (compreensão de imagens via Qwen-VL) e qwen_generate (texto para imagem via Qwen-Image) para modelos somente de texto

1há 5 diasVisão, voz e multimodalMIT
P

dsh-yali-image-generator

pptt121212/dsh-yali-image-generator

Plugin de geração de imagens para o DeepSeek-Harness. Solicite uma API Key da Yali AI em: https://api.yaliai.com/

1há 5 diasVisão, voz e multimodalMIT
C

deepsee

chang416/deepsee

DeepSee: visão para o DeepSeek Harness, roteamento multimodelo e autoverificações visuais com Gemini antes da entrega

1há 4 diasVisão, voz e multimodalMIT
A

dsh-voice-webspeech

anweat/dsh-voice-webspeech

Entrada de voz via Web Speech API do navegador: sem servidor, sem chaves, sem downloads de modelo (Edge=Azure, Chrome=Google speech)

1há 5 diasVisão, voz e multimodalMIT
J

dsh-voice

jesse-njx/dsh-voice

Notas de voz na entrada, respostas faladas na saída: dite áudio que se torna mensagens de usuário (transcrever), faça o agente ler as respostas em voz alta (falar), local-first em ~/.dsh/voice

1há 5 diasVisão, voz e multimodalMIT
T

dsh-voice-live

tangzheng202202/dsh-voice-live

Real-time duplex voice over Volcengine streaming ASR/TTS: agent reply narration, barge-in, wake word, live captions, 30 Chinese voices and a reply-first acknowledgment; builds in the DSH monorepo.

0anteontemVisão, voz e multimodalMIT
N

dsh-voice-input

newdanew/dsh-voice-input

Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.

0há 4 diasVisão, voz e multimodalMIT
M

dsh-fish-tts

mari23333/dsh-fish-tts

Reads assistant replies aloud via Fish Audio API only (bring your own key): per-message read-aloud, auto-read toggle, and a settings page for model, voice reference_id, encrypted API key, and proxy.

0ontemVisão, voz e multimodalMIT
H

dsh-voice

haoku123/dsh-voice

Full-duplex voice mode for the Web UI: a composer mic (RMS endpoint detection) transcribes speech with whisper running locally in the browser, assistant replies stream back as spoken audio sentence-by-sentence, and speaking interrupts playback and the running turn (true barge-in). No API key.

0anteontemVisão, voz e multimodalMIT