Skip to main content

Vision, Voice & Multimodal

Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.

312 plugins found

1

dsh-vision-sidecar

121103qwq/dsh-vision-sidecar

Hosted free vision sidecar for DeepSeek Harness with durable session evidence

42 months agoVision, Voice & MultimodalMIT
Y

dsh-image-subagent

yuqingsh/dsh-image-subagent

42 months agoVision, Voice & MultimodalMIT
G

dsh-vision-bridge

gxx182/dsh-vision-bridge

Bridges session images to configurable vision providers and returns text-only analysis to eligible DeepSeek Harness model routes.

42 months agoVision, Voice & MultimodalMIT
H

dsh-omni-workstation

huashenglian/dsh-omni-workstation

Omni-modal workstation for DSH: analyze_image over an ordered multi-card VLM failover chain, a six-tool vision toolkit (zoom, colour sampling, pixel diff, OCR, element detection, inline display) sharing the same image resolver, generate_image over OpenAI/DashScope/ComfyUI protocols, multi-card async generate_video with a /build-video-tool builder, and speak/clone_voice TTS over seven providers, all driven by one auto-saving settings page.

314 days agoVision, Voice & MultimodalMIT
W

dsh-hold-to-talk

wangzhanchao883/dsh-hold-to-talk

One-handed, keyboard-free input for the composer: press and hold the mouse on the input box, speak, release — the text lands in the draft, and sliding up cancels without leaving a half sentence behind. No mic button to aim at and no shortcut to remember; the hand never has to leave the input area, which is the point when the other hand is busy. Recognition runs fully offline (SenseVoice via sherpa-onnx, no API key, audio never leaves the machine).

3yesterdayVision, Voice & MultimodalMIT
N

dsh-plugin-speech

nakamuraia/dsh-plugin-speech

Read assistant replies aloud in DeepSeek Harness: text-to-speech providers with streaming playback.

318 days agoVision, Voice & MultimodalMIT
S

dsh-voice-assistant

supersyh-sss/dsh-voice-assistant

Voice assistant for dsh web: say the wake phrase (e.g. "小鲸") to activate hands-free dictation — what you say is transcribed and typed into the chat box automatically. Supports spoken edit commands (send, clear, new line, stop reading) and reads assistant replies aloud in Chinese. Speech recognition runs locally in-browser via sherpa-onnx WASM, so it works offline without an API key.

3last monthVision, Voice & MultimodalMIT
S

dsh-audiogen

shimingming520/dsh-audiogen

AI audio generation for the DeepSeek Harness web GUI — multi-vendor TTS, music, sound effects and voice design with a sidebar panel, model comparison, resource library and Agent tools.

324 days agoVision, Voice & MultimodalApache-2.0
L

dsh-voco

lgquan/dsh-voco

Continuous voice conversations for DSH with hands-free listening, push-to-talk, speech recognition, TTS replies, and background Agent delegation.

3last monthVision, Voice & MultimodalMIT
T

dsh-gsv

taoruiliu19/dsh-gsv

Real-time local TTS for DeepSeek Harness: voice presets, auto-read, engine setup assistant, a read-aloud button, and a settings panel for the GSV-TTS-Lite engine.

3last monthVision, Voice & MultimodalMIT
J

dsh-ximalaya

jerryqx/dsh-ximalaya

Listen to Ximalaya podcasts and audiobooks inside the DSH web UI: search albums, browse paginated tracks, and play through a now-playing bar (prev/play/next, seek, volume, rate). After QR login the "Mine" page lists your liked tracks (click to play), subscribed albums with latest-episode hints, and followed anchors with their public albums; the host relays audio streams with Range support and registers an `ximalaya_play` tool so the agent can search and play on request.

3last monthVision, Voice & MultimodalMIT
B

phone-eye

boheastill/phone-eye

Let your AI agent see and operate a real Android phone: phone_look (vision + UI-tree fusion), tap/swipe/type, screenshot — over adb, for any MCP client.

3last monthVision, Voice & MultimodalMIT
X

dsh-image-vision

xiaoyuink/dsh-image-vision

Image understanding for any DSH model: vision, OCR, grounding, and crop tools with domain presets for histopathology, cell biology, anatomy, clinical images and scientific figures.

329 days agoVision, Voice & Multimodal
L

Deepseek-Continuity

linxuhao/deepseek-continuity

Local image, voice, music and SFX generation plus transcription, with pinned identity: characters, animals, objects and actor voices are defined once and reused on every later call, degenerate output (a near-flat image, silent audio) is rejected instead of returned as success, and a generated line can be read back as text so a clone that swallowed its ending becomes visible. The engines unload when idle, and image generation and transcription can each be pointed at an OpenAI-shaped API instead of the local Vulkan backend.

312 days agoVision, Voice & MultimodalMIT
Y

dsh-ui-spec

yumimanji/dsh-ui-spec

DeepSeek Harness plugin that turns UI screenshots into implementation-grade web specs using OCR, deterministic geometry, scene graphs, assets, and render comparison.

32 months agoVision, Voice & MultimodalMIT
D

deepseekeyes

dttxorg/deepseekeyes

Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.

32 months agoVision, Voice & MultimodalMIT
N

voco-input-sh

nothree-code/voco-input-sh

Voice input for the Web UI: a mic button that drives local VocoType offline speech recognition and auto-inserts recognized text into the composer (auto-deploy, dedupe, continuous dictation).

320 days agoVision, Voice & MultimodalMIT
B

dsh-stt-input

baisama-cloud/dsh-stt-input

Speech-to-text voice input for the web UI: a mic button in the composer transcribes speech into the draft via the browser Web Speech API (zero-config) or an OpenAI-compatible Whisper API (OpenAI / Groq), with a selectable model and language in Settings.

3last monthVision, Voice & MultimodalMIT
W

visual-review

wang-bool/visual-review

Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.

32 months agoVision, Voice & MultimodalMIT
T

dsh-plugins (dsh-vision)

tzhr-invest/dsh-plugins

Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.

35 days agoVision, Voice & MultimodalMIT
S

dsh-plugin-multimodal

shinjiyu/dsh-plugin-multimodal

Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.

32 months agoVision, Voice & MultimodalMIT
M

dsh-vision-tools

moon09300731/dsh-vision-tools

Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.

32 months agoVision, Voice & MultimodalMIT
L

dsh-plugin-grok2api-media-tool

lsjspl/dsh-plugin-grok2api-media-tool

Gives dsh the ability to generate images and videos through the grok2api API.

3last monthVision, Voice & Multimodal
L

dsh-image-gen

leemancheung/dsh-image-gen

GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.

35 days agoVision, Voice & MultimodalMIT