Skip to main content
Z

dsh-continual-evolve

zk-andy/dsh-continual-evolve

Continual self-evolution: versioned, auditable, rollback-safe harness state (prompts, memory, skills, subagent specs) refined from session trajectories, with review gates and hot-reloaded skills.

Install

dsh plugin --profile web add github:zk-andy/dsh-continual-evolve

README

dsh-continual-evolve

中文 | English

awesome · DSH plugin npm CI License: MIT Node Tests Coverage · statements Coverage · branches Coverage · functions

Continual self-evolution for DeepSeek Harness: a versioned, auditable, rollback-safe harness state layer — prompt notes, memories, skills, subagent specs — refined from session trajectories.

The model proposes, the code guarantees. Every mechanical safety property — schema validation, atomic writes, snapshots, versioning, audit trail, acceptance decisions — is enforced in code, never by prompt discipline.

Why

Agents accumulate reusable experience (repeated failures, durable facts, reusable procedures) and forget it next session. This plugin turns that experience into first-class state:

  • Three scopes with merge semantics (global < project < local): local per-session staging, project per-workspace cross-session store, global cross-project — plus mechanical promotion guards so only portable, substantial, non-duplicate knowledge reaches global
  • Typed one-fact memories: every memory entry carries a recall type (user | feedback | project | reference); pitfalls (feedback) must include Why + How to apply
  • Dedicated background memory agent: eligible successful turns feed a bounded ZCode-style loop that searches the frozen memory manifest and proposes memory-only edits through a closed tool set; it cannot call agents, MCP, the network, or write source files. The generic review/planner/fate path is not run by this listener
  • Memory recall, projection, and receipts: evolve_recall reads back full memory content by query/kind/scope/type; every memory apply also materializes a readable MEMORY.md index plus one fact file per entry; each extraction lands a unified audit receipt (no-op/applied/declined with duration and turn stats) and only applied outcomes notify the session
  • Deterministic rollback: inverse edits generated from applied results — no LLM re-guessing
  • Benchmark loop: candidate refinements are evaluated against frozen cases by a separate scorer before acceptance (rubric encrypted at rest)
  • Store hygiene: /evolve consolidate turns write-time conflict hints and zero-use staleness into one approved, fully reversible batch of archives — with merge, near-duplicate content folds into the surviving original

How it works

  1. Sediment — the model creates entries via evolve_add, or the automatic Memory Agent consumes incremental snapshots after successful turns. Generic review/planner remains manual unless separately invoked.
  2. Capability-aware auxiliary calls — the memory loop, review, planner, wrapup, and fate resolve exact provider/model metadata through src/llm-text.ts, use the lowest advertised enabled reasoning effort (falling back to a closing effort only when no enabled level exists), and forward the host session id for provider routing; models without reasoning metadata use their provider default.
  3. Guard — code-enforced validation: edit schema, blast-radius/scope coherence, and the promotion policy (project-scoped markers, thin content, near-duplicate detection, credential screening keep the global store clean — secrets are rejected at every write sink, including mount materialization). Global creates that near-duplicate an existing entry are rejected at write time (≥0.8 similarity); moderate overlaps carry a conflictHint for later consolidation.
  4. Approve — global and project writes require explicit human approval; the dialog shows the bounded structured edit diff and conflict warnings, while malformed/lost responses remain retryable rather than counting as rejection.
  5. Apply & inject — memory batches preflight every persistent approval, recheck abort before writes, and compensate earlier scope writes if a later batch fails; every successful scope still passes through snapshot + audit. Prompt notes and delegation specs inject into the system prompt (capped, relevance-ranked, contradicted entries demoted, zero tokens when empty); memories/skills appear as a relevance-ordered capped directory index (- [memory:type:id] title hooks, full text one evolve_list away).
  6. Validate & roll back — benchmarks score candidates against frozen cases; rejected candidates roll back deterministically and are captured as draft regression cases (auto_regression benchmark).

Install

# from npm (installs and activates — ships its own bundle patch)
dsh plugin add dsh-continual-evolve

# or from source (first GitHub installs require approving the allowBuilds step)
dsh plugin add ZK-Andy/dsh-continual-evolve

Restart the DSH profile you use (dsh web or the desktop host) after installing or updating.

Usage

Commands (in-session):

CommandEffect
/evolvehelp + current local store
/evolve list · history · rollback <id>inspect and revert (add project for this project's store, global for the cross-project store)
/evolve plan [msg]run the LLM planner against the store
/evolve wrapupassess this session's local entries: promote / archive / keep
/evolve archive · unarchive · demote <id>hide from injection (data kept, restorable) — demote targets global noise
/evolve recall [scope] <query…>targeted memory recall: full content with version, source, and staleness
/evolve remember <type> [scope] <text…>immediately persist one typed memory (`user
/evolve forget [scope] <query…>locate one memory and archive it (restorable); ambiguous queries only list
/evolve consolidate [apply] [merge]report (or apply) one batch archive of conflict-hinted + stale zero-use global entries; merge folds near-duplicate content into the survivors
/evolve failuresaggregated failure classes (gate + benchmark)
/evolve log [tail N] [session <id>]plugin log
/evolve export · import <path>backup / restore a store
/evolve mount · unmount <skillId>hot-mount an executable skill as a live plugin
/evolve goal [objective · done · block]round-driven auto-review goal
/evolve benchmark …case lifecycle, runs, acceptance
/evolve pause · resume · statuspause/resume the auto-review gate (manual tools keep working), gate state
/evolve usageper-entry injection counts + exact provider-reported tokens for direct memory/review/planner/wrapup/fate calls (benchmark host subagents excluded)

Model tools: evolve_list / add / update / delete / rollback / recall (evolve_delete takes id or a batch ids array — one refinement, one approval; evolve_recall filters by query, kinds, scopes, memory types, and limit, and returns full content with version, source, and staleness).

For third-party consumers: every applied evolution (gate or manual) appends a structured evolve_complete event to reviews.jsonl (src/evolve-event.ts defines the shape) alongside the human-readable audit records.

/evolve usage also reads evolve/token-usage.jsonl: exact provider-reported input/cache/output/total tokens for the plugin's direct memory-agent, review, planner, manual-wrapup, and automatic-fate calls. The report covers a retained tail rather than lifetime usage, distinguishes missing provider samples, and explicitly excludes host benchmark subagents, their agent-loop calls, and per-entry injection attribution.

Injection shape: prompt notes and delegation specs inject with content (≤6/kind × 180 chars, relevance-ranked). Memories and skills appear as a relevance-ordered directory index ([memory:type:id] title hooks, capped at 15 lines with a fold counter) — full text via evolve_recall (targeted) or evolve_list. Every memory apply also refreshes a readable MEMORY.md index plus one fact file per entry in the store directory. Empty store = zero injected tokens.

Configuration

KeyDefaultMeaning
baseDirresolved DSH homeroot for the evolve/ stores
autoReviewfalseinitial Memory Agent default when no evolve/runtime.json exists yet; the listener is always registered, so this is not a registration gate
memoryMinUserWords3ZCode-style minimum lexical words in one direct user text part; uses CJK-aware segmentation
sessionCloseDrainMs15000session-close bounded drain for in-flight extraction in ms (0 aborts immediately)
reviewIntervalTurns6legacy local-fate cadence fallback; successful-turn review no longer waits for this interval
maxReviewInputChars40000trajectory slice handed to the gate
reviewBudgetTokens4096output budget for the gate call
notifyOnAutoReviewtruevisible follow-up notice after an applied gate run
requireGlobalApprovaltrueglobal and project edits ask for explicit approval
localFatefalseoptional local-entry promote/archive fate assessment; unreachable while the listener runs memory-only, so it only affects direct/full callers
fateIntervalTurnsfollows reviewIntervalTurnsminimum turns between fate assessments
goalBlockedWrapupTurns3consecutive blocked-goal gate runs trigger one fate assessment (0 disables)
promotionBlockPatternsPOSIX paths, session ids, ~/.dshcontent matching these is project-scoped and never promoted to global
promotionMinChars100whole promotions below this length stay local
injectionDirectoryLines15entry-directory lines per build before folding into a counter
sectionOrder118system-prompt section order
skillsDir<dshHome>/skillswhere skill entries materialize as SKILL.md bundles
rubricKeyauto-generated key fileAES-256-GCM passphrase for benchmark rubrics (DSH_EVOLVE_RUBRIC_KEY overrides)
logToFile / logLevel / logMaxBytestrue / 1 / 5 MiBplugin-owned JSONL file log with rotation
autoRollbackOnRejecttruedeterministic rollback after a benchmark rejection
autoCasetruefailed evolution attempts are captured as draft regression cases (auto_regression benchmark)
reviewModelagent's ownoptional cheaper model for the dedicated memory agent and review gate ("provider/model")
plannerPrefixCacheautoRoute A session-prefix input when cache evidence exists (session always, off legacy flat text)
plannerPrefixMaxChars12000session-prefix budget for Route A planning inputs (chars)
historyRetain{snapshots: 20, refinements: 500, reviews: 500, tokenUsage: 500}storage hygiene: snapshots per store, tail lines per store history, shared reviews.jsonl tail, and direct-call token-usage.jsonl tail

Example profile patch:

- id: continual-evolve
  config:
    autoReview: true
    reviewIntervalTurns: 6

The Memory Agent listener is registered even when autoReview is false — autoReview only supplies the initial default, so an install works without editing the profile. Use /evolve resume to enable successful-turn snapshots immediately, /evolve pause to suppress new snapshots and model work, and /evolve status to inspect the configured default plus the current runtime state. The runtime switch is stored in evolve/runtime.json; manual evolve_* tools and /evolve commands are not paused. Only the Memory Agent runs automatically: the generic review/planner, prompt/skill writes, and local-fate phases are not reachable from this listener. The memory trigger follows ZCode's lightweight eligibility: direct user text must contain at least memoryMinUserWords lexical words (CJK-aware segmentation), while empty/internal/direct-memory-write snapshots are skipped; compaction does not add a separate memory-only trigger. Every extraction writes a unified audit receipt (noop/applied/declined with duration and turn/search stats) to reviews.jsonl; only applied outcomes queue a visible follow-up, and a closing session lets in-flight extraction settle up to sessionCloseDrainMs before aborting.

Development

pnpm install && pnpm build   # deps + tsc -> lib/
pnpm test                    # vitest (1075 tests)
pnpm test:coverage           # v8 coverage, thresholds enforced in CI
pnpm coverage:gaps           # locate uncovered lines per file (read-only)
pnpm lint                    # oxlint src test

Project layout:

├── src/                   # engine, tools, commands, memory agent, recall, projection, gate, fate, benchmark, injection + token usage…
├── test/                  # vitest suites (58 files)
├── lib/                   # build output (tsc)
├── docs/
│   ├── design.md          # full design doc (hardening matrix)
│   ├── FAQ.md             # real failure/fix records
│   ├── gap-analysis.md    # vs prime-agent /refine + penguin-harness
│   ├── research/pi-dsh-competitor-gap-analysis.md  # pi/dsh ecosystem competitors
│   ├── experiment-bootstrap.md
│   ├── archive/           # closed point-in-time reports
│   └── research/          # penguin report + prime-agent annotated source
├── examples/README.md     # seed benchmark cases
└── .agents/               # AI collaboration layer (AGENTS.md, skills, ADR notes)

Docs & provenance

License

MIT

Related plugins