- Home
- Categories
- Vision, Voice & Multimodal
Vision, Voice & Multimodal
Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.
312 plugins found
dsh-vision-sidecar
121103qwq/dsh-vision-sidecar
Hosted free vision sidecar for DeepSeek Harness with durable session evidence
dsh-image-subagent
yuqingsh/dsh-image-subagent
dsh-vision-bridge
gxx182/dsh-vision-bridge
Bridges session images to configurable vision providers and returns text-only analysis to eligible DeepSeek Harness model routes.
dsh-omni-workstation
huashenglian/dsh-omni-workstation
Omni-modal workstation for DSH: analyze_image over an ordered multi-card VLM failover chain, a six-tool vision toolkit (zoom, colour sampling, pixel diff, OCR, element detection, inline display) sharing the same image resolver, generate_image over OpenAI/DashScope/ComfyUI protocols, multi-card async generate_video with a /build-video-tool builder, and speak/clone_voice TTS over seven providers, all driven by one auto-saving settings page.
dsh-hold-to-talk
wangzhanchao883/dsh-hold-to-talk
One-handed, keyboard-free input for the composer: press and hold the mouse on the input box, speak, release — the text lands in the draft, and sliding up cancels without leaving a half sentence behind. No mic button to aim at and no shortcut to remember; the hand never has to leave the input area, which is the point when the other hand is busy. Recognition runs fully offline (SenseVoice via sherpa-onnx, no API key, audio never leaves the machine).
dsh-plugin-speech
nakamuraia/dsh-plugin-speech
Read assistant replies aloud in DeepSeek Harness: text-to-speech providers with streaming playback.
dsh-voice-assistant
supersyh-sss/dsh-voice-assistant
Voice assistant for dsh web: say the wake phrase (e.g. "小鲸") to activate hands-free dictation — what you say is transcribed and typed into the chat box automatically. Supports spoken edit commands (send, clear, new line, stop reading) and reads assistant replies aloud in Chinese. Speech recognition runs locally in-browser via sherpa-onnx WASM, so it works offline without an API key.
dsh-audiogen
shimingming520/dsh-audiogen
AI audio generation for the DeepSeek Harness web GUI — multi-vendor TTS, music, sound effects and voice design with a sidebar panel, model comparison, resource library and Agent tools.
dsh-voco
lgquan/dsh-voco
Continuous voice conversations for DSH with hands-free listening, push-to-talk, speech recognition, TTS replies, and background Agent delegation.
dsh-gsv
taoruiliu19/dsh-gsv
Real-time local TTS for DeepSeek Harness: voice presets, auto-read, engine setup assistant, a read-aloud button, and a settings panel for the GSV-TTS-Lite engine.
dsh-ximalaya
jerryqx/dsh-ximalaya
Listen to Ximalaya podcasts and audiobooks inside the DSH web UI: search albums, browse paginated tracks, and play through a now-playing bar (prev/play/next, seek, volume, rate). After QR login the "Mine" page lists your liked tracks (click to play), subscribed albums with latest-episode hints, and followed anchors with their public albums; the host relays audio streams with Range support and registers an `ximalaya_play` tool so the agent can search and play on request.
phone-eye
boheastill/phone-eye
Let your AI agent see and operate a real Android phone: phone_look (vision + UI-tree fusion), tap/swipe/type, screenshot — over adb, for any MCP client.
dsh-image-vision
xiaoyuink/dsh-image-vision
Image understanding for any DSH model: vision, OCR, grounding, and crop tools with domain presets for histopathology, cell biology, anatomy, clinical images and scientific figures.
Deepseek-Continuity
linxuhao/deepseek-continuity
Local image, voice, music and SFX generation plus transcription, with pinned identity: characters, animals, objects and actor voices are defined once and reused on every later call, degenerate output (a near-flat image, silent audio) is rejected instead of returned as success, and a generated line can be read back as text so a clone that swallowed its ending becomes visible. The engines unload when idle, and image generation and transcription can each be pointed at an OpenAI-shaped API instead of the local Vulkan backend.
dsh-ui-spec
yumimanji/dsh-ui-spec
DeepSeek Harness plugin that turns UI screenshots into implementation-grade web specs using OCR, deterministic geometry, scene graphs, assets, and render comparison.
deepseekeyes
dttxorg/deepseekeyes
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
voco-input-sh
nothree-code/voco-input-sh
Voice input for the Web UI: a mic button that drives local VocoType offline speech recognition and auto-inserts recognized text into the composer (auto-deploy, dedupe, continuous dictation).
dsh-stt-input
baisama-cloud/dsh-stt-input
Speech-to-text voice input for the web UI: a mic button in the composer transcribes speech into the draft via the browser Web Speech API (zero-config) or an OpenAI-compatible Whisper API (OpenAI / Groq), with a selectable model and language in Settings.
visual-review
wang-bool/visual-review
Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.
dsh-plugins (dsh-vision)
tzhr-invest/dsh-plugins
Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.
dsh-plugin-multimodal
shinjiyu/dsh-plugin-multimodal
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.
dsh-vision-tools
moon09300731/dsh-vision-tools
Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.
dsh-plugin-grok2api-media-tool
lsjspl/dsh-plugin-grok2api-media-tool
Gives dsh the ability to generate images and videos through the grok2api API.
dsh-image-gen
leemancheung/dsh-image-gen
GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.