Pular para o conteúdo principal

Visão, voz e multimodal

Os plugins de Visão, voz e multimodal dão olhos, ouvidos e voz aos modelos DeepSeek somente texto no DeepSeek-Harness (dsh). Encontre ferramentas de visão e rotas de provedor que fazem OCR e descrevem capturas coladas via Zhipu GLM, Gemini, Doubao ou Ollama local, entrada de voz por microfone via Web Speech API do navegador ou APIs compatíveis com Whisper, leitura em voz alta com Edge TTS ou vozes personalizadas, modos de voz full-duplex, e geradores de imagens, vídeo e música.

86 plugins encontrados

B

dsh-voice-ai-girlfriend-plugin

beiyege-01/dsh-voice-ai-girlfriend-plugin

Voice AI girlfriend for the Web UI: FunASR mic input, Qwen3-TTS spoken replies, companion animation window, and two-way QQ chat (text/voice/image push) via NapCat.

0anteontemVisão, voz e multimodal
T

dsh-plugin-appshot

tauruswood/dsh-plugin-appshot

Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the composer for agent queries.

0ontemVisão, voz e multimodalMIT
S

dsh-easyvision

s3yf1337/dsh-easyvision

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

0anteontemVisão, voz e multimodalMIT
R

dsh-vision-subagent

ruby1304/dsh-vision-subagent

Vision for any DSH route: paste images in the Web composer with intent-aware auto-analysis, delegate workspace image reads to a Kimi/MiniMax vision subagent, and materialize pasted originals for editing.

0há 11 horasVisão, voz e multimodalMIT
N

dsh-subagent-vision

niuniuaba/dsh-subagent-vision

Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.

0há 15 horasVisão, voz e multimodalMIT
M

dsh-unsloth-hands

microherox/dsh-unsloth-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.

0há 16 horasVisão, voz e multimodalMIT
L

dsh-vision-plugin

ld-1101/dsh-vision-plugin

Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).

0anteontemVisão, voz e multimodalMIT
K

dsh-vision-recognizer

kaixinbaba/dsh-vision-recognizer

Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.

0há 3 diasVisão, voz e multimodalMIT
I

dsh-quicksight

isanti2016/dsh-quicksight

Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).

0anteontemVisão, voz e multimodalMIT
G

dsh-vision-guard

good-boy4069/dsh-vision-guard

Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.

0há 3 diasVisão, voz e multimodalMIT
E

dsh-plugin-image-input

elohia/dsh-plugin-image-input

Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).

0há 4 diasVisão, voz e multimodalMIT
B

dsh-vision-solution

br1nosense/dsh-vision-solution

Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.

0ontemVisão, voz e multimodal
S

dsh-nanobananapro

synmindai/dsh-nanobananapro

Gere imagens e vídeos no DeepSeek Harness por meio da API NanoBananaPro

0há 4 diasVisão, voz e multimodalMIT
S

dsh-seedance2

synmindai/dsh-seedance2

Gere imagens e vídeos Seedance no DeepSeek Harness por meio da API de IA Seedance 2

0há 4 diasVisão, voz e multimodalMIT