- Home
- Categories
- Vision, Voice & Multimodal
Vision, Voice & Multimodal
Vision, Voice & Multimodal plugins give the text-only DeepSeek models in DeepSeek-Harness (dsh) eyes, ears, and a voice. Find vision tools and provider routes that OCR and describe pasted screenshots via Zhipu GLM, Gemini, Doubao, or local Ollama, microphone voice input via the browser Web Speech API or Whisper-compatible APIs, read-aloud TTS with Edge TTS or custom voices, full-duplex voice modes, and generators for images, video, and music.
312 plugins found
dsh-vision-mix
haiziyao/dsh-vision-mix
Combine text, vision, and image-generation APIs into one Mix model with automatic routing: text-only requests go to the chat model, user images and agent screenshots go to the vision model, follow-ups keep using the same session image, and agents can generate or edit images with session-scoped call history.
dsh-ernie-image
omdsh-dev/dsh-ernie-image
openflowframes
zerohackz/openflowframes
dsh-paddle-ocr
omdsh-dev/dsh-paddle-ocr
dsh-plugin-aigc-canvas
huanlinoto/dsh-plugin-aigc-canvas
provider-agnostic AIGC HTTP 桥 + 无限画布 + ffmpeg 后处理,13 个工具含画布连边/reroll/媒体编辑 | Provider-agnostic AIGC HTTP bridge + infinite canvas + ffmpeg post-processing; 13 tools incl. canvas linking/reroll/media-edit
dsh-voice
jesse-njx/dsh-voice
Voice notes in, spoken answers out: dictate audio that becomes user messages (transcribe), have the agent read replies aloud (speak), local-first under ~/.dsh/voice.
dsh-llm-vision-bridge
einskyle/dsh-llm-vision-bridge
Native LLM-provider vision bridge: images pasted in the chat are described by a vision model (Qwen3-VL via pi-ai/llama.cpp) and the text description is fed to text-only DeepSeek for the reply — image admission, routing and compaction all run through harness-native mechanisms, with an LRU description cache and 503 retry.
dsh-mic-input
qt-chen/dsh-mic-input
Microphone voice input for the composer: browser Web Speech API live transcription, dedupe/auto-continue, smart punctuation, language and auto-send settings.
dsh-open-eyes
hyper-dsh-plugins/dsh-open-eyes
Vision bridge for text-only DeepSeek routes that analyzes attached and local images through configurable OpenAI Responses, Chat Completions, or Anthropic Messages endpoints while leaving image-capable routes native.
dsh-gpt-image
cai-mh/dsh-gpt-image
Generate images through a logged-in web ChatGPT session and download them to a local directory.
dsh-ui-mockup
mackwan84/dsh-ui-mockup
ui_mockup 工具 Consumer:研讨阶段生成 UI 线框图/高保真草图,落盘设计稿并提示确认后锁定设计规格;含客户端卡片、图片路由与 i18n
dsh-voice-call
biliye/dsh-voice-call
Voice call assistant for the DSH Web GUI: a draggable floating call ball, browser-side VAD that sends each utterance after a pause, FunASR HTTP or streaming speech recognition, MiniMax or OpenAI-compatible TTS replies, an optional wake-word mode with auto-sleep, and task dispatch to separate subagent sessions with progress, stop, and spoken completion reports.
dsh-tool-lipsync
yu-wenchao/dsh-tool-lipsync
Free lip-sync video generation plugin for DSH with 3000+ voices and 500+ languages.
dsh-vision-plugin
bug-huntter/dsh-vision-plugin
Configurable image recognition for text-only DSH models: image messages are first transcribed by an OpenAI-compatible vision model (Base URL, model ID and API key set in a Settings section) and then passed to the main model as text, while image-input support is advertised. The API key auth scheme is selectable — OpenAI, Anthropic, Gemini or Azure style request headers — and a missing key is reported before any request is sent.
dseyesopen
balue8246-maker/dseyesopen
Vision bridge for text-only DeepSeek: send images (alone or mixed with text) and a fast tiny vision model describes them behind the scenes — the chat keeps the picture, DeepSeek sees only text. Pure plugin: uninstall restores everything.
dsh-visibridge
lhbsaa/dsh-visibridge
Structured vision evidence (OCR/layout/semantics) plus a USB camera capture tool for a "shoot-look-adjust" debug loop; backends: Ollama, DeepSeek, Xiaomi.
dsh-short-video-studio
fengyungithub/dsh-short-video-studio
基于deepseek harness的类MiniMax-Design的AI视频创作工作台
dsh-watch-video
zeshuochen/dsh-watch-video
Subtitle-first video transcription with SRT export, cancellable job controls, and a faster-whisper large-v3 fallback when subtitles are unavailable.
dsh-voice-agent (voice-app)
wayneyu430/dsh-voice-agent
A conversational voice frontend Agent for dsh: speak naturally over ByteDance Duplex, delegate requests to background tasks, and hear their asynchronous results reported by voice.
dsh-voice-chat
maoyuching/dsh-voice-chat
Doubao-style voice chat for DSH: press-and-hold mic in the composer converts speech to text and auto-sends, and AI replies are read aloud. Optional LLM condensing (long replies summarized into short spoken lines, following the current conversation model), TTS-friendly text cleaning, selectable Edge TTS voices, adjustable silence auto-stop, and settings embedded in DSH settings dialog.
phone-lens (phone-lens)
yxqfg/phone-lens
Turn a phone camera into a live viewfinder and photo input for your dsh session, over LAN or USB.
dsh-capture
max-null/dsh-capture
Dual-engine screen capture for DSH: inside the SSiD desktop shell, a global hotkey or tray opens a fullscreen box-select overlay (all monitors, per-display pixel-perfect frames) with in-place red-box annotation and a WeChat-style toolbar; in plain DSH (browser), the composer camera button captures the display via getDisplayMedia and reuses the same single-phase box-select + annotation overlay in-page. The cropped image lands in the current conversation composer through the official attachment intake.
phone-lens (lens-mate)
yxqfg/phone-lens
Turn a phone camera into a live viewfinder and photo input for your dsh session, over LAN or USB.
dsh-labnana
exoticknight/dsh-labnana
Labnana image generation for DeepSeek Harness: text-to-image / image-to-image / precise editing with credits estimation, subscription balance and web settings UI.