ビジョン、音声とマルチモーダル
ビジョン、音声とマルチモーダルプラグインは、DeepSeek-Harness(dsh)のテキスト専用 DeepSeek モデルに目、耳、そして声を与えます。貼り付けたスクリーンショットを智譜 GLM、Gemini、Doubao、あるいはローカルの Ollama で OCR・説明するビジョンツールとプロバイダールート、ブラウザの Web Speech API や Whisper 互換 API によるマイク音声入力、Edge TTS やカスタム音声による読み上げ、全二重の音声モード、画像・動画・音楽の生成ツールを探せます。
86 件のプラグインが見つかりました
modlens
liustack/modlens
テキスト専用モデル向けのビジョンブリッジ: 画像を貼り付けると、構造化された JSON エビデンス(OCR、レイアウト、セマンティクス)を取得。
dsh-vision-router
ysr666/dsh-vision-router
テキスト専用エージェント向けの無料ビジョン機能: キー不要の内蔵ビジョンチェーンとピクセルツール群(Q&A、グラウンディング、クロップ、ピクセル差分、カラー、OCR、SVG トレース、切り抜き、スクリーンショット)。画像を貼り付けるだけで使用可能。
dsh-vision-toolkit
anionex/dsh-vision-toolkit
テキスト専用モデル向けのビジョンタスク: 意図を汲む画像 Q&A、長尺スクリーンショット OCR、UI 再現、グラウンディング、ピクセル差分。
picturereader
jing-hy/picturereader
Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.
dsh-vision
linenxi-ctrl/dsh-vision
DeepSeek Harness 向け外部ビジョンプラグイン: クジラボタンの設定パネル、自動返信付き画像認識、エージェント用スクリーンショット/認識ツール。
dsh-media-skills
mjorgin/dsh-media-skills
Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.
dsh-vision-proxy
flyvhidbwo/dsh-vision-proxy
DeepSeek の頭脳 + 自動画像文字起こし: GUI で画像を添付すると、テキスト専用の DeepSeek に届く前に任意の OpenAI 互換 VLM でテキストに変換される——自分の API キーを使うキー方式の高速パス(デフォルト qwen3.7-flash、DashScope/Zhipu/OpenRouter や任意の OpenAI 互換エンドポイントに対応)、または設定不要で自動検出されるローカル Ollama。
dsh-vision-opencode
poiuyjie/dsh-vision-opencode
Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.
dsh-visual-plugin
jyh20030112/dsh-visual-plugin
dsh-visual-plugin。テキスト専用モデルに目を与える: ユーザーの画像を任意の OpenAI 互換ビジョンモデルに転送し、結果を Web UI の右パネルで確認。
Gemini-Eyes
consolesun/gemini-eyes
MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.
dsh-plugin-tts
1624318455/dsh-plugin-tts
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
dsh-imagegen
dickpy/dsh-imagegen
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
dsh-chat-imagine
corrinehu/dsh-chat-imagine
Automatically generates and displays images in the DSH chat via API channels or local CLIs (supports mmx / codex / agy).
dsh-vision
54xkeee/dsh-vision
Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.
dsh-voice-input-plugin
zhangbo-cn/dsh-voice-input-plugin
Composer mic for the Web UI: tap-to-monitor live transcription and hold-to-talk, with host Edge TTS reply reading that streams while the model generates, echo-pause during reading, and tap-to-stop.
dsh-deepseek-vision
siegfly/dsh-deepseek-vision
A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.
dsh-windows-ocr
maxwell-feng/dsh-windows-ocr
Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
dsh-voice
3274375092/dsh-voice
Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.
deepseek-vision (dsh-plugin-deepseek-vision)
gou-gee/deepseek-vision
Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.
dsh-guide-dog
atropinoltt/dsh-guide-dog
MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.
dsh-her-eyes
huashenglian/dsh-her-eyes
AI が VLM(マルチモーダルモデル)を自動呼び出しして視覚分析を行えるようにする dsh プラグイン。
dsh-speak
alan2z/dsh-speak
Voice-announce the final reply on Windows (SAPI5 natural voices) and macOS (system voice); skips reasoning and tool calls, one-line npm install.
dsh-plugin-multimodal
shinjiyu/dsh-plugin-multimodal
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.
free-vision-skill
niyongsheng/free-vision-skill
Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.