メインコンテンツへスキップ

プラグイン

DeepSeek-Harness プラグインを閲覧・絞り込み・インストールできます。

18 件のプラグインが見つかりました

F

dsh-vision-proxy

flyvhidbwo/dsh-vision-proxy

DeepSeek ブレイン + 自動画像文字起こし:GUI で画像を添付すると、デフォルトで公式の deepseek-v4-flash-vision-exp により文字起こしされます(純粋テキストの V4-Pro ブレインも画像を見ることができます)。代替として、任意の OpenAI 互換 VLM またはローカル Ollama も使用可能。

16先月ビジョン、音声とマルチモーダルMIT
R

dsh-plugin-call-me

radres/dsh-plugin-call-me

CallKit 経由でスマートフォンに着信: `call_me` と `text_me` ツールに加え、ターン終了時や承認時にオプションで電話をかけ、話した回答をセッションに文字起こしして戻す。

7一昨日統合とリモートMIT
S

dsh-voice-assistant

supersyh-sss/dsh-voice-assistant

dsh web 向けの音声アシスタント:ウェイクフレーズ(例:「小鲸」)を言うとハンズフリーの音声入力が有効になり、話した内容が文字起こしされて自動的にチャットボックスへ入力されます。音声による編集コマンド(送信、クリア、改行、読み上げ停止)に対応し、アシスタントの返答を中国語で読み上げます。音声認識は sherpa-onnx WASM によりブラウザ内でローカルに実行されるため、API キーなしでオフラインでも動作します。

3先月ビジョン、音声とマルチモーダルMIT
B

dsh-stt-input

baisama-cloud/dsh-stt-input

Web UI 用音声入力:コンポーザーのマイクボタンがブラウザ Web Speech API(ゼロコンフィグ)または OpenAI 互換 Whisper API(OpenAI / Groq)で音声をドラフトに転写。Settings でモデルと言語を選択可能。

3先月ビジョン、音声とマルチモーダルMIT
J

dsh-voice

jesse-njx/dsh-voice

音声メモを入力し、音声で回答を受け取れます。話した内容がユーザーメッセージになり(文字起こし)、エージェントの返信を読み上げさせることもできます(発話)。~/.dsh/voice 配下でローカル優先に動作します。

32 か月前ビジョン、音声とマルチモーダルMIT
B

dsh-vision-plugin

bug-huntter/dsh-vision-plugin

テキスト専用DSHモデル向けの設定可能な画像認識:画像メッセージはまずOpenAI互換のビジョンモデル(Base URL、モデルID、APIキーは設定セクションで指定)によって書き起こされ、画像入力サポートが告知されつつ、テキストとしてメインモデルに渡されます。APIキー認証方式は選択可能 — OpenAI、Anthropic、Gemini、Azureスタイルのリクエストヘッダー — キーが欠落している場合、リクエスト送信前に報告されます。

27 日前ビジョン、音声とマルチモーダルMIT
K

dsh-vision-recognizer

kaixinbaba/dsh-vision-recognizer

添付された画像を設定可能なモデル(15以上のOpenAI互換およびAnthropicベンダー)経由でテキストに書き起こし、DeepSeekが応答し続ける間に処理するビジョンプロバイダールート。

22 か月前ビジョン、音声とマルチモーダルMIT
3

dsh-vision (vision-route)

314857493/dsh-vision

`deepseek-vision` プロバイダルートを登録する。Web GUI は貼り付けた画像を受け付け、無料の Zhipu GLM vision API で文字起こししてから DeepSeek アダプタへ渡す。

2先月ビジョン、音声とマルチモーダルMIT
X

dsh-vision-bridge

ximengxiaolan/dsh-vision-bridge

コンポーザーに添付された画像は、テキスト専用の DeepSeek モデルに渡る前に、OpenAI 互換のビジョンモデルによってテキストに変換されます。

22 か月前ビジョン、音声とマルチモーダルMIT
J

dsh-mmroute

jmxsxwyzjdwl/dsh-mmroute

テキスト専用モデル向けの透過的なマルチモーダルルーティング:すべてのモデル呼び出しにおけるすべての画像が、自身のマルチモーダル理解器によって完全に文字起こしされます(逐語 OCR、データ、不確実領域、インジェクション対策済み)。vision_relook による集中的な再確認と、画像関連の失敗時の自動リトライを備えます。バンドルされたエンドポイントなし、借用ログインなし。

128 日前ビジョン、音声とマルチモーダルMIT
A

dsh-audio-copilot

ai-yucheng/dsh-audio-copilot

DeepSeek Harness 用 Audio Copilot:音声の文字起こし(ASR)と音声合成(TTS)。テキストのみのエージェントに耳と声を与える。Windows ローカル SAPI TTS を即座に利用可能。OpenAI 互換の ASR/TTS エンドポイント設定可能。コンポーザー内音声入力も含む。

12 か月前ビジョン、音声とマルチモーダルMIT
3

dsh-vision

314857493/dsh-vision

DeepSeek Harness プラグイン。画像入力を宣言する `deepseek-vision` ルートで、無料の Zhipu GLM ビジョンモデルを使って貼り付けられた画像を文字起こしし、その後 DeepSeek アダプターに委譲します。

12 か月前ビジョン、音声とマルチモーダルMIT
C

dsh-voice-mimo

ch1bug/dsh-voice-mimo

Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).

021 日前ビジョン、音声とマルチモーダルMIT
M

dsh-voice-input-en

mohith-das/dsh-voice-input-en

Minimal English-only voice input for the DeepSeek Harness Web UI: a mic button in the composer that transcribes speech into the draft via the browser's native SpeechRecognition API. No dependencies, no subprocess, no network calls beyond whatever the brow

0先月ビジョン、音声とマルチモーダルMIT
S

dsh-vision-pro-bridge

shainedemo/dsh-vision-pro-bridge

Vision bridge for text-only DeepSeek models: transcribes attached images with deepseek-v4-flash-vision-exp before they reach deepseek-v4-pro, with no third-party dependencies.

0先月ビジョン、音声とマルチモーダルMIT
J

dsh-autovision

junkrat9527/dsh-autovision

Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.

02 か月前ビジョン、音声とマルチモーダルMIT
N

dsh-voice-input

newdanew/dsh-voice-input

Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.

02 か月前ビジョン、音声とマルチモーダルMIT
E

dsh-plugin-image-input

elohia/dsh-plugin-image-input

Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).

02 か月前ビジョン、音声とマルチモーダルMIT