dsh-vision-bridge
goodandready/dsh-vision-bridge
Self-owned vision bridge for DeepSeek Harness: pick the vision model yourself from the available ones. Image content blocks are rewritten into text markers at the agent boundary so text-only models never fail a turn, and a describe_image tool answers via
インストール
dsh plugin --profile web add github:goodandready/dsh-vision-bridgeREADME
dsh-vision-bridge
Vision bridge for DeepSeek Harness (dsh) — a self-contained replacement for dsh-vision-router.
When the chat model has no vision (e.g. deepseek-v4-flash) and a message contains an image, the image never reaches the text-only model. Instead, the plugin substitutes an automatic text description from the vision model you choose (Hermes-style). The text model just reads text and keeps the conversation going; the image stays in the session log and in the UI.
describe_image— a tool for when you need more precise details on an image (ask a model).- Description cache by
contentHashof the bytes — an image is described once, then the previous description is reused on later turns. - You pick the vision model in Settings → Vision (a top-level settings section alongside General / Models / Plugins).
Install
# From npm after publishing:
dsh plugin --profile web add @goodandready/dsh-vision-bridge
# From GitHub:
dsh plugin --profile web add github:GooDAnDReaDY/dsh-vision-bridge
# Locally from a checkout:
dsh plugin --profile web add /path/to/dsh-vision-bridge
Restart the Web UI, open Settings → Vision, pick a provider and model (or leave empty — auto-picks the first vision-capable model).
Replacing dsh-vision-router
This does the same job but under your control and with your own vision model. If dsh-vision-router is installed, remove it:
dsh plugin --profile web remove dsh-vision-router
How it works (Hermes-style)
- An image (from you or from a tool like
generate_image) enters the chat and is shown by the UI. - On every LLM request, in
agent/pre-step:- if the chat model is vision-capable (
inputModalitiesincludesimage) — images go through as-is; - if the model is text-only — for each image block:
- already cached description (by
attachmentId) → reuse it; - otherwise automatically call the vision model via
ctx.llm.stream, get a description, cache it, and inject it as[The user attached an image. Here is what it contains: ...].
- already cached description (by
- if the chat model is vision-capable (
- The text model reads text, not pixels. Turns never fail with
UNSUPPORTED_CONTENT. describe_imagestays available for when you need more precise details on an image (ask a model directly).
Settings
Settings → Vision (top-level section):
- Provider — the vision model's provider (from the LLM catalog).
- Model — the model (only vision-capable ones are shown).
- Empty → auto-pick the first vision model.
In settings.yaml:
dsh-vision-bridge:
visionProvider: "" # empty = auto-pick
visionModel: "" # empty = auto-pick
sanitizeImages: true
maxImageBytes: 20971520
timeoutMs: 120000
Structure
dsh-vision-bridge/
├── package.json # dsh bundle/plugin metadata + peerDependencies
├── cordis.patch.yml # bundle layer: inserts the plugin row
├── lib/index.js # host: agent/pre-step sanitizer + describe_image + cache + model list
├── lib/client.js # browser: top-level Settings → Vision section
├── README.md
└── LICENSE # MIT
License
MIT