본문으로 건너뛰기

비전, 음성 및 멀티모달

비전, 음성 및 멀티모달 플러그인은 DeepSeek-Harness(dsh)의 텍스트 전용 DeepSeek 모델에 눈, 귀, 그리고 목소리를 부여합니다. 붙여넣은 스크린샷을 Zhipu GLM, Gemini, Doubao 또는 로컬 Ollama로 OCR·설명하는 비전 도구와 제공자 라우트, 브라우저 Web Speech API 또는 Whisper 호환 API를 이용한 마이크 음성 입력, Edge TTS나 커스텀 음성으로 답변을 읽어 주는 TTS, 전이중 음성 모드, 이미지·영상·음악 생성기를 살펴보세요.

플러그인 312개 찾음

G

dsh-fal-image-gen

goodandready/dsh-fal-image-gen

FAL image generation for DeepSeek Harness: a generate_image tool backed by the FAL queue API (default model fal-ai/flux-2/klein/9b). Images render directly in the conversation, are saved to the workspace, and all settings are editable from the Web GUI.

02개월 전비전, 음성 및 멀티모달MIT
S

dsh-bundle-vision

skillre/dsh-bundle-vision

Vision bundle + plugin for DeepSeek Harness: the describe_image tool reads local images and asks any configured multimodal route, with zero core changes

02개월 전비전, 음성 및 멀티모달MIT
1

dsh-mindsee

123cdxcc/dsh-mindsee

DeepSeek Harness 插件:以 MindSee 为后端,为 DeepSeek 提供图片相关能力

02개월 전비전, 음성 및 멀티모달MIT
O

dsh-voice-input

opensquad-ai/dsh-voice-input

SenseVoice 语音输入插件 for DeepSeek Harness:在对话输入框旁添加麦克风按钮,录音后调用本地 SenseVoice 服务转成文本填入输入框。首次使用自动下载模型并显示进度,后端由插件自动启动。

02개월 전비전, 음성 및 멀티모달MIT
R

dsh-gemini-multimodal

realalexandreai/dsh-gemini-multimodal

DeepSeek Harness plugin: multimodal tools (image/audio/video/document understanding, transcription, image generation) via Gemini API or the local Antigravity CLI.

02개월 전비전, 음성 및 멀티모달MIT
Z

dsh-tool-image-gen

zhangjunjesse/dsh-tool-image-gen

DSH tool plugin: generate images through ToAPIs async GPT-Image-2 API (submit task, poll, download).

02개월 전비전, 음성 및 멀티모달MIT
O

soyo

ottohere-mourn/soyo

DSH-native video understanding with configurable multimodal providers

02개월 전비전, 음성 및 멀티모달MIT
T

dsh-voice-live

tangzheng202202/dsh-voice-live

Real-time duplex voice over Volcengine streaming ASR/TTS: agent reply narration, barge-in, wake word, live captions, 30 Chinese voices and a reply-first acknowledgment; builds in the DSH monorepo.

0그저께비전, 음성 및 멀티모달MIT
N

dsh-voice-input

newdanew/dsh-voice-input

Voice input for the web UI: a mic button in the composer that transcribes speech into the draft via the Web Speech API, with an optional auto-send toggle.

02개월 전비전, 음성 및 멀티모달MIT
S

dsh-easyvision

s3yf1337/dsh-easyvision

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

02개월 전비전, 음성 및 멀티모달MIT
N

dsh-subagent-vision

niuniuaba/dsh-subagent-vision

Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.

017일 전비전, 음성 및 멀티모달MIT
M

dsh-unsloth-hands

microherox/dsh-unsloth-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.

0지난달비전, 음성 및 멀티모달MIT
M

dsh-koboldcpp-hands

microherox/dsh-koboldcpp-hands

Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.

0지난달비전, 음성 및 멀티모달MIT
M

dsh-windows-ocr

maxwell-feng/dsh-windows-ocr

Windows 내장 엔진(Windows.Media.Ocr)을 통한 첨부 이미지의 로컬 OCR: 모델에는 인식된 텍스트만 전송되며 이미지 바이트는 전송되지 않습니다. 비전 패스스루는 선택적으로 활성화합니다.

0비전, 음성 및 멀티모달업데이트 중단 추정
M

dsh-tesseract-ocr

maxwell-feng/dsh-tesseract-ocr

Tesseract를 통한 첨부 이미지의 로컬 OCR: 인식된 텍스트만 모델로 전송하고 이미지 바이트는 전송하지 않으며, vision passthrough는 선택 사항이다.

0비전, 음성 및 멀티모달업데이트 중단 추정
L

dsh-vision-plugin

ld-1101/dsh-vision-plugin

Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).

02개월 전비전, 음성 및 멀티모달MIT
I

dsh-quicksight

isanti2016/dsh-quicksight

Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).

02개월 전비전, 음성 및 멀티모달MIT
G

dsh-vision-guard

good-boy4069/dsh-vision-guard

Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.

02개월 전비전, 음성 및 멀티모달MIT
E

dsh-plugin-image-input

elohia/dsh-plugin-image-input

Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).

02개월 전비전, 음성 및 멀티모달MIT
B

dsh-vision-solution

br1nosense/dsh-vision-solution

Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.

02개월 전비전, 음성 및 멀티모달
T

dsh-plugin-vision

tdf1995/dsh-plugin-vision

텍스트 전용 LLM을 위한 비전 기능: 무료 Gemini 및 GLM 비전 API를 통한 이미지 설명 / OCR / VQA

02개월 전비전, 음성 및 멀티모달MIT
H

vision-tool

haowencang/vision-tool

DeepSeek Harness용 대화형 UI/UX 시각 리뷰 플러그인: 모든 OpenAI 호환 멀티모달 엔드포인트(예: agnes-2.5-flash) 기반의 vision_review / vision_ask 도구

02개월 전비전, 음성 및 멀티모달MIT
S

dsh-nanobananapro

synmindai/dsh-nanobananapro

NanoBananaPro API를 통해 DeepSeek Harness에서 이미지와 동영상 생성

02개월 전비전, 음성 및 멀티모달MIT
S

dsh-seedance2

synmindai/dsh-seedance2

Seedance 2 AI API를 통해 DeepSeek Harness에서 이미지와 Seedance 동영상 생성

02개월 전비전, 음성 및 멀티모달MIT