Skip to main content

Plugins

Browse, filter, and install DeepSeek-Harness plugins.

83 plugins found

L

dsh-attachment-formats

linkingoscar/dsh-attachment-formats

DeepSeek Harness Web plugin — PDF text-layer and scanned-PDF OCR, Office docx/xlsx/pptx to Markdown, TIFF/epub and long-document index cards.

42 days agoTools & CapabilitiesApache-2.0
M

dsh-nuphus-mcp

mrpulor-gh/dsh-nuphus-mcp

Official DSH plugin for nuphus-mcp (by the nuphus-mcp maintainers): desktop and browser automation (38 tools) with PaddleOCR element perception and Chrome CDP browsing, via a persistent stdio MCP child process.

4last monthTools & CapabilitiesMIT
G

deepseek-vision

gou-gee/deepseek-vision

DeepSeek Harness 原生视觉 Bundle:粘贴或拖入图片,通过托管的 deepseek-vision-mcp 调用 OpenAI 兼容视觉模型。

4last monthMCP & ConnectorsMIT
N

free-vision-skill

niyongsheng/free-vision-skill

Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.

4last monthVision, Voice & MultimodalMIT
G

deepseek-vision (dsh-plugin-deepseek-vision)

gou-gee/deepseek-vision

Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.

423 days agoVision, Voice & MultimodalMIT
C

dsh-learning-mode

chplus0/dsh-learning-mode

Learning Mode agent preset: a coding agent that teaches while coding — concrete scenario-grounded explanations, Socratic guidance, and TODO(你) practice blanks, modeled on Claude Code's Learning output style; installable via dsh plugin add (dsh-learning-mode on npm).

4last monthTools & Capabilities
F

dsh-plugin-deepeye

favio8/dsh-plugin-deepeye

DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.

42 months agoVision, Voice & Multimodal
H

dsh-omni-workstation

huashenglian/dsh-omni-workstation

Omni-modal workstation for DSH: analyze_image over an ordered multi-card VLM failover chain, a six-tool vision toolkit (zoom, colour sampling, pixel diff, OCR, element detection, inline display) sharing the same image resolver, generate_image over OpenAI/DashScope/ComfyUI protocols, multi-card async generate_video with a /build-video-tool builder, and speak/clone_voice TTS over seven providers, all driven by one auto-saving settings page.

311 days agoVision, Voice & MultimodalMIT
X

dsh-image-vision

xiaoyuink/dsh-image-vision

Image understanding for any DSH model: vision, OCR, grounding, and crop tools with domain presets for histopathology, cell biology, anatomy, clinical images and scientific figures.

326 days agoVision, Voice & Multimodal
Y

dsh-ui-spec

yumimanji/dsh-ui-spec

DeepSeek Harness plugin that turns UI screenshots into implementation-grade web specs using OCR, deterministic geometry, scene graphs, assets, and render comparison.

3last monthVision, Voice & MultimodalMIT
F

dsh-wechat-mp-studio

funcwei/dsh-wechat-mp-studio

WeChat Official Account content studio: anti-homogenization rotation writing, low-creativity remediation playbook, blessing-image visual baseline, gpt-image pipeline with OCR acceptance, and the measured xiaolvshu draft web API.

3last monthSkillsMIT
O

dsh-paddle-ocr

omdsh-dev/dsh-paddle-ocr

32 months agoVision, Voice & MultimodalBSD-3-Clause
Y

dsh-md-convert

yakoylp/dsh-md-convert

Convert Office documents and PDFs (including scanned ones) to structurally formatted Markdown via a CPU-first routing OCR pipeline (RapidOCR/SLANet/FormulaNet).

219 days agoWorkflow & AutomationMIT
L

dsh-visibridge

lhbsaa/dsh-visibridge

Structured vision evidence (OCR/layout/semantics) plus a USB camera capture tool for a "shoot-look-adjust" debug loop; backends: Ollama, DeepSeek, Xiaomi.

2last monthVision, Voice & MultimodalMIT
F

auto-mouse

fish121380/auto-mouse

Windows desktop UI context picker for Codex, DeepSeek Harness, and MCP clients: select windows, UI elements, or screen regions with hover highlighting, UI Automation, screenshots, and local OCR, then review and approve Markdown/JSON context before output.

2last monthTools & CapabilitiesMIT
G

dsh-vision-bridge

goodandready/dsh-vision-bridge

Routes images to a vision model of your choice - auto-rewrite, explicit tools, or hybrid - so a text-only chat model does not fail a turn that contains a picture.

23 days agoVision, Voice & MultimodalMIT
3

dsh-vision (vision-tool)

314857493/dsh-vision

Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.

225 days agoVision, Voice & MultimodalMIT
1

dsh-wsl-media

173787247/dsh-wsl-media

Local media/doc pipeline: ffprobe, extract, thumbnail, PDF, ASR, pandoc, OCR, exif (allowRoots include IM inbox).

110 hours agoDev & Plugin ToolsMIT
T

Mimir (mimir-skin)

tommyhedgerow/mimir

Publishes a lesson into the DSH conversation: the dependency spine, the question and the vault's drawings, drawn in the vault's own palette.

14 days agoUI EnhancementsMIT
X

computer-user-vision

xie129716/computer-user-vision

Windows computer use forked from computer-user: 13 computer_* tools that read the screen and drive the mouse and keyboard. Image-capable routes get the screenshot as a real image with an exact image-to-screen mapping, so no external OCR; elements return as UI Automation refs so a click lands on the exact control rectangle; Ctrl+Alt+Esc stops every call.

115 days agoTools & CapabilitiesMIT
W

dsh-screenshot-capture

wangzhanchao883/dsh-screenshot-capture

Point-and-shoot screenshot capture: a system floating window turns a new screenshot (or copied image) into an Obsidian note with comments and key-point marking, instant Tongyi Qianwen OCR, per-day merging, and an optional evening AI organization pass.

118 hours agoTools & CapabilitiesMIT
Z

dsh-file-convert

zzy-12345678/dsh-file-convert

Local-first file conversion: 26 conversions across images, PDF (with OCR and experimental PDF→DOCX), data, audio/video and office docs; 7 tools, all local, no API keys.

1last monthVision, Voice & Multimodal
J

dsh-mmroute

jmxsxwyzjdwl/dsh-mmroute

Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.

123 days agoVision, Voice & MultimodalMIT
C

aura-vision

ck-epsilon/aura-vision

Free vision OCR with adaptive tile recognition for long documents and Markdown/Word/PNG/Excel export.

1last monthVision, Voice & MultimodalMIT