Skip to main content
K

dsh-session-eval

kittimzhe/dsh-session-eval

Retrospective agent-session evaluation for DeepSeek Harness: deterministic grade cards from persisted session logs, cross-session regression diffs, era comparison — no benchmark runs, no LLM judges.

Install

dsh plugin --profile web add github:kittimzhe/dsh-session-eval

README

dsh-session-eval

English | 中文

npm npm downloads tests license

Retrospective evaluation for DeepSeek Harness agent sessions — deterministic grade cards computed from sessions that already happened, plus cross-session regression diffs that answer "did my agent get better or worse after that plugin / prompt / model change?"

No benchmark suites to author. No LLM judges. No new runs. Every grade is a fixed, documented threshold over persisted session logs — fully reproducible.

Install

Requirements: Node.js 20 or 22 · a DeepSeek Harness profile that mounts the commands and sessionQuery services (the shipped web / agent profiles qualify).

dsh plugin --profile web add dsh-session-eval

Or from GitHub:

dsh plugin --profile web add github:kittimzhe/dsh-session-eval

The bundle overlay mounts /eval, /eval-diff, and /eval-history. If you maintain your own cordis.patch.yml, keep this row:

- insert:
    - id: session-eval
      name: 'dsh-session-eval'

Try it once

In DeepSeek Harness:

/eval
Session eval — a1b2c3d4… (dsh-session-eval v0.2.2)
Overall: A
  Reliability   A    2/87 tool results errored (2.3%)
  Re-ask        A    1 re-ask signal(s) over 23 turns (0.4/10 turns)
  Tool load     B    3.8 tool calls per turn
Notes:
  • Wall time 41.2 min over 23 turns.

The measurement layer of the session toolchain

PluginLayerAnswers
dsh-session-exportEvidence"What exactly happened in this session?"
dsh-session-recallMemory"What did I do before, and where is it?"
dsh-session-evalMeasurement"Was that session good? Is the trend improving?"

All three read through the same trusted ctx.sessionQuery seam, so any persistence backend (JSONL or SQLite) works without touching raw artifacts.

Contributing

  • Local dev: npm ci && npm run typecheck && npm test && npm run bundle (Node 20 or 22).
  • Start in the source: src/grade.ts (thresholds, grade cards), src/metrics.ts (metric extraction), src/trend.ts (/eval-history trend logic). See CONTRIBUTING.md for the source map and first-PR suggestions.
  • Open gaps: #1 (configurable thresholds), #2 (/eval-history --since), #3 (/eval-history --out) — or browse issues labeled good first issue.
  • Rules: behavior changes need tests; doc changes must update README.md and README.zh.md in sync; version releases belong to the maintainer. Details: CONTRIBUTING.md.

How it differs from benchmark-style evaluation

dsh-eval (benchmark)dsh-session-eval (this)
EvaluatesRuns you launch ahead of timeSessions that already happened
SetupAuthor benchmark YAML + casesNone — read existing logs
JudgeMetric folding over trialsFixed thresholds, deterministic
Best atPre-deploy benchmarksRegression tracking over real work

Both are useful; they don't compete. Benchmarks test what you predicted; retrospective evaluation watches what actually happened.

Commands

CommandEffect
/evalGrade card for the current session
/eval --id <sessionId>Grade a specific session
/eval --jsonMachine-readable metrics + card
/eval-diff <beforeId> <afterId>Regression comparison, noise-tolerant
/eval-diff … --jsonMachine-readable trend report
/eval-history [N]Grade the last N sessions in this workspace and show the trend
/eval-history … --jsonMachine-readable history report

/eval-history finds the sessions for you via the same sessionQuery seam /archive uses (scoped to the current workspace's cwd), grades each, and diffs the first against the last. /eval-diff remains the tool when you already know the two ids that matter.

Typical use: grade the ten sessions before and after a plugin or prompt change; a persistent Reliability regression is an early warning that the change hurt real work, not just benchmarks.

Dimensions & thresholds

All grades derive from three deterministic metrics. n/a when the session lacks the relevant activity.

DimensionMetricABCDE
Reliabilitytool error rate≤5%≤10%≤20%≤35%>35%
Re-askcorrections per 10 turns¹0<1<2<4≥4
Tool loadtool calls per turn0.5–80.1–0.5 or just above 8<0.1 or 8–20>20—

¹ A correction signal is each extra user message inside a consecutive-user run — the human re-asking before any assistant reply, a behavioral proxy for "the agent didn't get it right the first time."

Overall is the worst graded dimension — conservative by design. Thresholds are constants in src/grade.ts; making them configurable is on the roadmap (see issues).

Diff verdicts include a noise band (±2pp error rate, ±0.5 re-ask, ±1 tool load) so trivial jitter doesn't cry regression.

Determinism guarantee

Same session logs in → same grade card out, every time, forever. There is no model in the loop, no sampling, no clock dependence in the metrics (wall time is reported, never graded). This is what makes /eval-diff trustworthy as a regression signal.

Development

npm ci
npm run typecheck   # tsc --noEmit
npm test            # vitest run
npm run bundle      # tsdown → lib/

License

MIT

Community

Related plugins