Passer au contenu principal
P

dsh-defend

perrylink/dsh-defend

Détecte les motifs d'injection de prompt, de jailbreak et de fuite de secrets au niveau des points de jonction agent/pre-step, tools/pre-execute et tools/post-execute, avec des niveaux allow/ask/block, des événements d'audit defend/detection assainis, un outil defend_report et une protection contre les commandes de suppression destructive.

Installer

dsh plugin --profile web add github:perrylink/dsh-defend

README

🛡️ dsh-defend

  • 1024 store channel: npm i -g dsh1024 once, then dsh1024 plugin --profile web add dsh-defend (counts toward the deepseek1024.com install ranking). Gitee dshfind OpenSSF Scorecard

Prompt-injection, jailbreak, and secret-leak defense for DeepSeek Harness.

Rules decide the known. Interception decides the rest — and everything is audited.

License DSH plugin dsh-doctor DSH Market Node CI Version npm version npm downloads

English · 简体中文 · Español · Português · हिन्दी


⭐ 如果它帮到了你

这个插件是 DSH 插件家族的一员(40+ 个,全部 Apache-2.0)。如果你在用,给个 star —— 它不会解锁任何功能,但会让下一个人在搜索里更容易找到它。

English: part of a 40+ plugin family for DeepSeek Harness. If it is useful, a star helps the next person find it — nothing is gated behind it.

Compatibility

SurfaceStatus
HarnessDeepSeek Harness dsh-v0.1.7-rc.2 (verified 2026-09-25; peer ranges >=0.1.2-rc.1 <0.2.0 || >=0.1.5-alpha.1 <0.2.0 || >=0.1.6-0 <0.2.0 || >=0.1.7-0 <0.2.0). On this line Session.append's third argument exists only for surface-eligible event types and is a SurfaceIntent, so the non-surface defend/detection type still cannot stamp the ignorable marker: session-log audit stays fail-closed-disabled and /defend renders that state explicitly. Session format V4 has no tool-result content block — this plugin never produced one, and its two content walkers keep a read-only fallback for the retired V3 wrapper so sessions written before the upgrade still scan. Verified 2026-09-25 (dual typecheck rulers + full test suite + build + self-contained/artifacts gates + pack; exactly one copy of the host type graph).
Node^22.19.0 || >=24.0.0
PlatformsAll (pure host; no native code, no network)
ModelAny (detection runs before content reaches the model)

What you get

dsh-defend puts two independent layers in front of the agent:

  1. Destructive-delete guard — the executable form of the 8·14/8·16 postmortem lesson. On tools/pre-execute, recursively deleting shell commands are refused unless every target is an explicit absolute path inside the session workspace and outside the protected prefixes (home config, .dsh/.claude, system directories). Dry-run markers (-WhatIf, --dry-run, git clean -n) pass, because they are exactly the check the lesson demands.
  2. Detection layer — ported from four upstream assets (all Apache-2.0, see THIRD_PARTY_NOTICES.md): 25 Prompt-Injection-Payloads rules, 25 Jailbreak-Detector patterns through a pure-TypeScript Aho-Corasick automaton, 12 secret grammars from Secret-Key-Leaker-Detect plus the issuers' public references, and the Prompt-Attack-Dataset kept verbatim as the regression benchmark.

Three interception points, one decision model each:

PointScannedDecision
agent/pre-stepinbound user messagesallow → next(); ask → approval; block → reject the step
tools/pre-executetool argumentsallow → next(); ask → approval; block → deny
tools/post-executetool resultsallow → next(); ask → approval; block → corrective feedback

Defaults: ask for every family, block for critical secrets (the upstream interrupt-on-sight semantics). No approval answerer = fail closed. Every pass-through calls next() — downstream policy plugins are never short-circuited.

inbound message ── agent/pre-step ── scan ── clean → next()/enter
tool arguments ── tools/pre-execute ── scan ── allow → next()
tool results   ── tools/post-execute ── scan ── block → feedback
                                  │
                                  └─ defend/detection audit (rule id, family,
                                     severity, decision — never matched text)

Quick start

# 1. install the bundle into your profile
dsh plugin --profile web add "github:PerryLink/dsh-defend#main"

# or from npm (published releases)
dsh plugin --profile web add dsh-defend

# 2. restart and verify the row
dsh --profile web --dump-config | grep -A3 'id: dsh-defend'

Install & uninstall

  • git channel (latest main): dsh plugin --profile web add "github:PerryLink/dsh-defend#main" — the prepare script builds with production dependencies only.
  • npm channel (published releases): dsh plugin --profile web add dsh-defend.
  • tarball channel: pnpm pack in this repo, then dsh plugin --profile web add ./dsh-defend-<version>.tgz.
  • uninstall: dsh plugin --profile web remove dsh-defend (or remove the row from the profile patch).

Configuration

All tunables are Schemastery Config fields (changeable from cordis.yml). An id-targeted override replaces the whole row — restate every key you need. cordis.patch.yml documents each key inline.

KeyDefaultMeaning
enabledtrueMaster switch for both layers
actiondenyDestructive-delete guard action (deny / ask)
toolNames['bash','persistent-bash','terminal-bash']Tool names whose command arguments the guard reviews
detection.enabledtrueDetection-layer switch
detection.maxScanChars10000Scan cap per interception (head only)
detection.normalizeUnicodetrueNFKC-normalize text before scanning (blocks lookalike-Unicode bypass)
detection.secretMinEntropy3.0Minimum Shannon entropy (bits/char) to admit a secret regex hit; 0 disables
detection.injectionActionaskInjection family: allow / ask / block
detection.jailbreakActionaskJailbreak family: allow / ask / block
detection.secretActionaskSecret family: allow / ask / block
detection.secretBlockCriticaltrueCritical secrets always block regardless of secretAction
detection.audittrueWrite defend/detection session audit events
detection.allowUnmarkedAuditfalseKeep writing session audit on hosts whose Session.append predates the ignorable marker (every released line so far) or that fail-closed on unknown event types (host 0.1.2-rc.1+), accepting the unresumable-session hazard
detection.maxReportEntries200In-memory report ring-buffer cap
registerCommandtrueRegister the /defend command
registerTooltrueRegister the defend_report tool

Tools & surfaces

SurfaceKindNotes
defend_reporttoolTotals (recorded/blocked/asked), per-family counts, and the 20 most recent matches — never matched text
/defendcommandThe same summary as text
agent/pre-steplistenerInbound message scanning (enter/reject)
tools/pre-executelistenerTool-argument scanning (deny/ask) + the destructive-delete guard
tools/post-executelistenerTool-result scanning (block feedback)

Permissions & data

  • Permissions: ask decisions ride the official approval seam; nothing is re-implemented or bypassed. The plugin declares session:append and network:none in its workshop manifest.
  • Data: nothing is stored on disk; the report ring buffer is in-memory and bounded. No network requests, no subprocesses.
  • Session log: defend/detection events carry rule id, family, category, severity, secret type, decision, and scan facts — matched text never reaches the log, and secret matches are type-only by construction.

Security boundaries

  • Detection, not enforcement. The guard and the detection layer only produce deny/ask/block decisions on official seams; the sandbox and approval systems remain the enforcement authorities.
  • Fail closed. Missing approval answerer, missing session, or a missing services surface degrades to the strictest decision — never to silent pass-through.
  • No content leaves the process. Scanning is local; audit events are sanitized; secrets are never logged, displayed, or reported.
  • Bounded work. Scan caps, one match per rule, and ring-buffer bounds keep hostile inputs from consuming unbounded resources.

Known limitations

  • Detection gaps. The rule library catches the ported vocabularies and their tolerant variants; novel phrasing, lookalike-Unicode encodings (NFKC normalization is tracked as future work), and multi-step attacks can evade it. The benchmark pins the measured floor (27/28 on the upstream dataset) so regressions are visible.
  • No model-level verdicts. dsh-defend is deterministic; it never calls a model and cannot judge novel intent.
  • Message rejection is silent. agent/pre-step reject carries no reason to the model (the seam has no reason field); the audit event records the rule facts.
  • Session audit and the ignorable marker. Audit appends request the envelope's ignorable: true marker so any harness build can load the log. Every released harness line so far (0.1.0-rc.1–0.1.0-rc.8, 0.1.1-rc.1–0.1.1-rc.2) silently drops it — the event lands unmarked and makes the session unresumable on stricter builds; host 0.1.2-rc.1 retains the envelope field for stored-log read compatibility only, but Session.append still cannot stamp it and the read path rejects unmarked unknown event types (defend/detection is not registered), so writing there also makes the session unloadable. dsh-defend therefore decides BEFORE the first append (peer-version pre-check; unresolvable versions fail closed) and disables session-log audit with a one-time warning. Set detection.allowUnmarkedAudit: true to opt back in. See issue #2.

Development

pnpm install        # node ^22.19 || >=24
pnpm run typecheck  # tsc: src + tests against the local harness checkout
pnpm run typecheck:ci  # tsc against the published 0.1.7-rc.2 types (no paths)
pnpm test           # vitest: 96 tests, 9 suites (detection benchmark incl.)
pnpm run build      # tsdown bundle + tsc declarations (lib/)
pnpm run verify:self-contained  # dependency specs resolve from the registry
pnpm run verify:artifacts       # built ESM face + shipped files present
pnpm pack           # the published tarball

Benchmark

The red-team benchmark (per-category P/R/F1 over 105 samples, plus the 27/28 fixture floor) is published in benchmark/RESULTS.md; regenerate it with node --experimental-strip-types benchmark/run.mjs (zero new dependencies, no build step).

Topics

dsh, dsh-plugin, deepseek-harness, deepseek, cordis, security, prompt-injection, jailbreak, secret-scanning, ai-safety

Contributors

  • @PerryLink — creator and maintainer: destructive-delete guard, the four-asset detection port, interception wiring, audit surface, and the five-language docs.
  • @cuohua — the precise report on defend/detection events landing unmarked and making sessions unresumable on stricter builds (#2); the runtime host-capability detection and the ignorable-marker discipline derive directly from that analysis.

This project is one of the 45 DeepSeek Harness plugins maintained by PerryLink. If this one helps you, the others likely will too:

PluginOne-liner
dsh-auto-reviewSecond-model auto-review on the approval chain, fail-closed by default
dsh-autotierAutomatic strong/cheap model-tier routing with deterministic risk guards and a /tier command
dsh-background-agentsDurable background child agents with a Web UI sidebar, messaging and interrupt
dsh-budgetCost governance for DeepSeek Harness: budgets, carbon, and latency in one panel.
dsh-catalogDSH Desktop Market standard catalog source for the PerryLink family
dsh-cert-mcpRead-only MCP server exposing the certification registry: grades, snapshots and five-dimension evidence
dsh-checkpoint-rewindClaude Code /rewind-equivalent: snapshots, session forks, one-shot restore
dsh-claude-moveMigrate Claude Code sessions, memory, skills and CLAUDE.md into DSH
dsh-clickCross-platform native desktop control for DeepSeek Harness — Windows first.
dsh-composer-historyTerminal-style input history for the web composer: arrows, Ctrl+R search
dsh-data-qualityDataset quality checks and citation cross-checks (the optional numeric bridge consumed here)
dsh-defendPrompt-injection, jailbreak, and secret-leak defense for DeepSeek Harness.
dsh-doublecheckEngineering-discipline guard: requirements grill, test gates, adversary review
dsh-drawUnified static-image generation routing for DeepSeek Harness.
dsh-fastRead-only performance diagnostics for DeepSeek Harness.
dsh-fund-researchDeterministic research reports for Chinese public mutual funds
dsh-githubGitHub PR/issues integration for DSH, every write gated by approval
dsh-industry-researchIndustry research orchestration that seals its deliverables through this plugin's ctx.researchReport.assemble
dsh-layaLaya typed decisions (noul/choice/score) as a first-class Cordis service and model-visible tools
dsh-libraryLocal document knowledge base for DeepSeek Harness.
dsh-local-aiLocal-model (Ollama) integration for DeepSeek Harness.
dsh-lsp-actionsLSP diagnostics, formatting, completion, code actions and rename over language servers
dsh-maskPII masking middleware: anonymize at the model boundary, restore at the display layer
dsh-mcp-panelRead-only MCP runtime panel: /mcp command + Settings tab with status, tools and errors
dsh-mementoApproval-gated cross-session memory: ctx.memory seam + SQLite + memory tool
dsh-observeOpenTelemetry and Langfuse observability exporter for DeepSeek Harness.
dsh-output-stylesClaude Code outputStyles-equivalent runtime style switching
dsh-permission-rulesClaude Code-style declarative allow/deny/ask permission rules with audit
dsh-plugin-certificationCommunity certification registry with repro-checkable grades and badges
dsh-plugin-doctorZero-dependency static + sandbox smoke detector for DSH plugins
dsh-plugin-guidePlugin-development knowledge base as an on-demand agent skill
dsh-plugin-kitShared zero-runtime-dependency toolkit for the PerryLink DSH plugins
dsh-plugin-upgradeOne-package, one-corridor-index plugin upgrade skill: routes a repository to the matching closed corridor card
dsh-plugin-upgrade-015Merged 0.1.3-alpha.1 → 0.1.5-rc.1 upgrade corridor card plus a zero-dependency seam scanner
dsh-reachMulti-channel approval/question bridge: WeChat/Telegram/Feishu, session console
dsh-research-reportVerifiable research-report engine: content-addressed evidence ledger and sealed versions
dsh-scoreMulti-dimensional quality scoring for DeepSeek Harness plugins.
dsh-session-pinPin sessions in the Web sidebar with durable ordering
dsh-session-syncCross-device session sync for DeepSeek Harness — a dedicated git mirror of your session store.
dsh-skill-pack-securitySecurity-audit skill pack: secret scan, dependency and supply-chain review
dsh-talkVoice-first session loop for DeepSeek Harness: talk to it, hear it answer.
dsh-team-roomsCross-session team rooms: shared message bus, task board and timeline
dsh-test-driveIsolated install-and-smoke test drives for DeepSeek Harness plugins.
dsh-ticktickTickTick/Dida365 task bridge: session-header panel + 11 tools
dsh-translateVendor parameter translation and deterministic JSON repair for DeepSeek Harness.

Install from the DSH Desktop Market

All PerryLink plugins are browsable in the built-in DSH Desktop Market: Market → Sources → add source → paste https://perrylink-dsh-catalog.perrylink.workers.dev/catalog-source.json → select it. Installation still goes through the Market's npm-identity verification and your confirmation.

License

Apache License 2.0 © 2026 dsh-defend contributors

Plugins associés