メインコンテンツへスキップ

ビジョン、音声とマルチモーダル

ビジョン、音声とマルチモーダルプラグインは、DeepSeek-Harness(dsh)のテキスト専用 DeepSeek モデルに目、耳、そして声を与えます。貼り付けたスクリーンショットを智譜 GLM、Gemini、Doubao、あるいはローカルの Ollama で OCR・説明するビジョンツールとプロバイダールート、ブラウザの Web Speech API や Whisper 互換 API によるマイク音声入力、Edge TTS やカスタム音声による読み上げ、全二重の音声モード、画像・動画・音楽の生成ツールを探せます。

312 件のプラグインが見つかりました

G

dsh-fal-image-gen

goodandready/dsh-fal-image-gen

FAL image generation for DeepSeek Harness: a generate_image tool backed by the FAL queue API (default model fal-ai/flux-2/klein/9b). Images render directly in the conversation, are saved to the workspace, and all settings are editable from the Web GUI.

02 か月前ビジョン、音声とマルチモーダルMIT
S

dsh-bundle-vision

skillre/dsh-bundle-vision

Vision bundle + plugin for DeepSeek Harness: the describe_image tool reads local images and asks any configured multimodal route, with zero core changes

02 か月前ビジョン、音声とマルチモーダルMIT
1

dsh-mindsee

123cdxcc/dsh-mindsee

DeepSeek Harness 插件:以 MindSee 为后端,为 DeepSeek 提供图片相关能力

02 か月前ビジョン、音声とマルチモーダルMIT
O

dsh-voice-input

opensquad-ai/dsh-voice-input

SenseVoice 语音输入插件 for DeepSeek Harness:在对话输入框旁添加麦克风按钮,录音后调用本地 SenseVoice 服务转成文本填入输入框。首次使用自动下载模型并显示进度,后端由插件自动启动。

02 か月前ビジョン、音声とマルチモーダルMIT
R

dsh-gemini-multimodal

realalexandreai/dsh-gemini-multimodal

DeepSeek Harness plugin: multimodal tools (image/audio/video/document understanding, transcription, image generation) via Gemini API or the local Antigravity CLI.

02 か月前ビジョン、音声とマルチモーダルMIT
Z

dsh-tool-image-gen

zhangjunjesse/dsh-tool-image-gen

DSH tool plugin: generate images through ToAPIs async GPT-Image-2 API (submit task, poll, download).

02 か月前ビジョン、音声とマルチモーダルMIT
O

soyo

ottohere-mourn/soyo

DSH-native video understanding with configurable multimodal providers

02 か月前ビジョン、音声とマルチモーダルMIT
T

dsh-voice-live

tangzheng202202/dsh-voice-live

Real-time duplex voice over Volcengine streaming ASR/TTS: agent reply narration, barge-in, wake word, live captions, 30 Chinese voices and a reply-first acknowledgment; builds in the DSH monorepo.

03 日前ビジョン、音声とマルチモーダルMIT
N

dsh-voice-input

newdanew/dsh-voice-input

Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.

02 か月前ビジョン、音声とマルチモーダルMIT
S

dsh-easyvision

s3yf1337/dsh-easyvision

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

02 か月前ビジョン、音声とマルチモーダルMIT
N

dsh-subagent-vision

niuniuaba/dsh-subagent-vision

Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.

017 日前ビジョン、音声とマルチモーダルMIT
M

dsh-unsloth-hands

microherox/dsh-unsloth-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.

0先月ビジョン、音声とマルチモーダルMIT
M

dsh-koboldcpp-hands

microherox/dsh-koboldcpp-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.

0先月ビジョン、音声とマルチモーダルMIT
M

dsh-windows-ocr

maxwell-feng/dsh-windows-ocr

Windows組み込みエンジン(Windows.Media.Ocr)による添付画像のローカルOCR:モデルに送信されるのは認識されたテキストのみで、画像バイトは送信されません。ビジョンパススルーはオプトインです。

0ビジョン、音声とマルチモーダル更新が停滞している可能性
M

dsh-tesseract-ocr

maxwell-feng/dsh-tesseract-ocr

Tesseract による添付画像のローカル OCR。認識されたテキストのみを model に送り、画像バイト列は送らない。vision passthrough は任意。

0ビジョン、音声とマルチモーダル更新が停滞している可能性
L

dsh-vision-plugin

ld-1101/dsh-vision-plugin

Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).

02 か月前ビジョン、音声とマルチモーダルMIT
I

dsh-quicksight

isanti2016/dsh-quicksight

Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).

02 か月前ビジョン、音声とマルチモーダルMIT
G

dsh-vision-guard

good-boy4069/dsh-vision-guard

Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.

02 か月前ビジョン、音声とマルチモーダルMIT
E

dsh-plugin-image-input

elohia/dsh-plugin-image-input

Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).

02 か月前ビジョン、音声とマルチモーダルMIT
B

dsh-vision-solution

br1nosense/dsh-vision-solution

Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.

02 か月前ビジョン、音声とマルチモーダル
T

dsh-plugin-vision

tdf1995/dsh-plugin-vision

テキスト専用 LLM 向けのビジョン機能。無料の Gemini と GLM のビジョン API を利用した、画像説明/OCR/視覚質問応答(VQA)に対応します。

02 か月前ビジョン、音声とマルチモーダルMIT
H

vision-tool

haowencang/vision-tool

DeepSeek Harness 向けの、会話形式による UI/UX ビジュアルレビュープラグイン。vision_review / vision_ask ツールは、任意の OpenAI 互換マルチモーダルエンドポイント(例: agnes-2.5-flash)を利用します。

02 か月前ビジョン、音声とマルチモーダルMIT
S

dsh-nanobananapro

synmindai/dsh-nanobananapro

NanoBananaPro API を通じて、DeepSeek Harness で画像と動画を生成できます。

02 か月前ビジョン、音声とマルチモーダルMIT
S

dsh-seedance2

synmindai/dsh-seedance2

Seedance 2 AI API を通じて、DeepSeek Harness で画像と Seedance 動画を生成できます。

02 か月前ビジョン、音声とマルチモーダルMIT