Passer au contenu principal
H

dsh-vision-analysis

harvey-will/dsh-vision-analysis

Plugin de vision DeepSeek Harness : 8 modes d'analyse (décrire, OCR, données de graphique, revue d'interface, détection d'objets, comparaison, génération de code, débogage), toute API de vision compatible OpenAI ou Anthropic, avec un modèle de vision gratuit intégré et un basculement automatique en cas de limite de débit.

Installer

dsh plugin --profile web add github:harvey-will/dsh-vision-analysis

README

DSH Vision Analysis — image understanding for the DeepSeek Harness

npm version vision source: FREE License: MIT DSH 0.1.x dsh-plugin awesome-dsh-plugin

English · 中文

中文: DeepSeek Harness 图像理解插件 · 8 种分析模式(描述 / OCR / 图表取数 / UI 评审 / 目标检测 / 对比 / 代码生成 / 端点诊断)· 兼容任意 OpenAI / Anthropic 视觉端点 · 支持本地图片、链接与截图 · 密钥掩码、隐私优先。


✨ Why DSH Vision Analysis?

Your text-only agent can finally "see" — with a free vision source built in: install the plugin, paste an image, ask. No API key, no model swap, no local-file dance.

  • 🆓 Built-in FREE vision source — ships pointed at OVHcloud AI Endpoints' anonymous tier (Qwen2.5-VL-72B). Zero cost, zero key, zero config.
  • 🖼️ Image bridge for text-only models — paste or send images directly in conversation; the plugin routes them to vision automatically (native multimodal routes stay untouched).
  • 🔁 Rate-limit failover — when one vision model is throttled, the next in the chain answers; if everything is exhausted you get clear recovery guidance instead of a failure.
  • 🧾 Structured output — chart-data and ocr return machine-readable JSON (rows, lines, …) your agent can consume directly.
  • 8 analysis modes out of the box — describe, ocr, ui-review, chart-data, object-detect, compare, code-gen, debug — each with a tuned instruction template.
  • Any vision endpoint — OpenAI chat/completions or Anthropic messages wire formats. MiMo, Step, SiliconFlow, OpenRouter, Gemini (OpenAI-compat), GPT-4o, Claude, Qwen-VL, or a local Ollama / LM Studio / vLLM.
  • Any input — absolute local path, http(s) URL, or base64 data: URL; up to 4 images per call with built-in comparison.
  • Privacy-first by design — image bytes never enter the session log or reach your main model; only the vision model's text comes back. The debug report never reveals your API key (fully masked).
  • Production-grade plumbing — result caching, retry with exponential backoff, live configuration from Settings → 插件配置.
  • Dependency-light — just @deepseek-ai/schemastery + @deepseek-ai/dsh-settings at runtime.

🖼️ Demo

Paste an image, ask a question, get a real answer — even on a text-only model. The image is routed to your configured vision endpoint and the analysis lands straight in the conversation:

dsh-vision-analysis in action: a pasted image of a DeepSeek fan-art character is identified with full reasoning in a DSH conversation

In the screenshot: a pasted image plus the question "这是谁?" — the vision endpoint identifies the DeepSeek fan-art character and walks through its reasoning, all without switching models or saving files locally.

More scenarios — real outputs from the free vision models

Three everyday capabilities, each answered by a different free vision model automatically (when one is rate limited, the plugin fails over to the next).

1. OCR — pull text out of documents and screenshots

A short ops-report document to transcribe

Weekly Ops Report — 2026-W33 Item 01 · Pending action: review queue / escalate blocker Item 02 · Pending action: review queue / escalate blocker … (all lines transcribed verbatim)

2. Charts → structured data your agent can use

A monthly revenue bar chart

{ "title": "Monthly Revenue — Q1–Q3", "rows": [["Jan","82"],["Feb","95"],…] }

3. UI review — a designer's eye on your interface

A simple e-commerce product page mockup

• Inconsistent button styling across "Add to cart" and "Checkout" (High) • Product name and price lack visual hierarchy (Medium) • Cart items unstructured; subtotal not visually distinct (Medium)


🚀 Quick start

# From GitHub (no npm needed)
dsh plugin --profile web add github:Harvey-Will/dsh-vision-analysis

# Or one-click from the plugin market inside the Harness

Restart the web profile and ask your agent to analyze an image by path or URL:

"Use analyze_image to OCR /tmp/screenshot.png and tell me what it says."

That works with zero configuration: the plugin ships pointed at a free anonymous vision endpoint (OVHcloud AI Endpoints, Qwen2.5-VL-72B) — no API key required.

Two ways to use it

1. analyze_image tool (zero config) — the agent reads a local path, an http(s) URL, or a data URL. Works immediately after install.

2. Paste images straight into the conversation (image bridge) — requires two setup steps:

  • add the model to bridgeModels in the plugin config;
  • declare image in that model's inputModalities in settings.yaml (this is what lets the Harness admit image prompts for it).
# ① ~/.dsh/settings.yaml — under llm-deepseek.models, for each text-only model:
#    inputModalities: [text, image]
# ② plugin config:
bridgeModels: [deepseek-v4-flash]
Bring your own endpoint (optional)
config:
  apiFormat: openai          # or anthropic
  baseURL: https://api.siliconflow.cn/v1
  apiKey: your-key           # leave empty for anonymous/local endpoints
  model: Qwen/Qwen2.5-VL-72B-Instruct
  fallbackModels: [Qwen3.5-9B]   # same-endpoint alternates tried on HTTP 429

🧭 Choose the right mode

ModeWhat it doesBuilt-in tokens / temp
describeGeneral understanding (default)4096 / 0.7
ocrExact text extraction4096 / 0.0
ui-reviewDesign review with score4096 / 0.5
chart-dataTables + trend from charts4096 / 0.0
object-detectObjects, people, activities4096 / 0.5
compareTwo+ images side by side4096 / 0.5
code-genHTML+CSS from a UI shot4096 / 0.3
debugEndpoint connectivity report4096 / 0.7

🔧 The tool

analyze_image(image?, images?, mode?, prompt?)
  • image — absolute path, http(s) URL, or data:image/...;base64, URL
  • images — up to maxImages (default 2, max 4) for multi-image calls
  • mode — one of the eight above; describe by default
  • prompt — your precise instruction overrides the mode template

A targeted prompt beats a generic description: prompt: "Extract the table as CSV" >> prompt: "Describe this".

⚙️ Configuration

- id: vision-analysis
  name: dsh-vision-analysis
  config:
    apiFormat: openai          # openai | anthropic
    baseURL: https://api.siliconflow.cn/v1
    apiKey: ''                # empty → UNIVERSAL_VISION_API_KEY → local model
    model: Qwen/Qwen2.5-VL-72B-Instruct
    defaultMode: describe
    maxImages: 2              # 1-4
    maxBytes: 10485760        # per-image cap (10 MB)
    timeoutMs: 120000
    maxTokens: 4096
    temperature: 0.7
    modes:                    # per-mode overrides
      ocr:
        temperature: 0.0

All fields are editable live from Settings → 插件配置 (API key field is masked).


🔒 Security & privacy

  • Your images stay private: local files are read by the tool and sent base64-embedded only to your configured endpoint; the raw bytes never enter the session log or reach the main model.
  • Your key stays secret: never embedded in requests to the main model; the debug report only says configured / not configured — no prefix, no characters.
  • Prefer the environment: keep keys out of cordis.yml — use UNIVERSAL_VISION_API_KEY or the masked secret field in Settings.
  • Endpoints are not sandboxed by tool approvals — only point the tool at endpoints you control, and only reference http(s) image URLs you trust the endpoint to fetch.
  • Installing a plugin runs its code with your permissions — review the source before installing.

🧩 Compatibility

Supported
DeepSeek Harness0.1.0-rc.x – 0.2.0-rc.x (verified on 0.2.0-rc.1)
Node.js^22.19 || >=24
Vision wire formatsOpenAI chat/completions, Anthropic messages
Image formatsPNG, JPEG, GIF, WebP, BMP (local / URL / data URL)

⚠️ Community plugin — not an official DeepSeek product. The Harness API is in developer preview and may break between versions.


Built for the DeepSeek Harness community · dsh-plugin topic · awesome-dsh-plugin

Found a bug or have an idea? Open an issue — PRs welcome.

MIT License

Plugins associés