Skip to main content
M

agent-guard

mokuyoaxis/agent-guard

Harness-neutral reliability infrastructure for AI coding agents: reversible destructive actions, pre-emission redaction, recovery evidence, and a shared Decision Protocol with optional native adapters.

Install

dsh plugin --profile web add github:mokuyoaxis/agent-guard

README

AGENT-GUARD

CI npm version License Python 3.9+ v0.2.3-rc2 source

Make destructive agent actions reversible by default. · 简体中文

Agent Guard is a reliability layer for coding agents. It makes supported high-impact actions recoverable instead of permanently destructive, while keeping routine work automatic.

  • Destructive file operations can be relocated to .agent-trash/ with a recovery manifest instead of being permanently deleted.
  • Destructive Git operations can snapshot recoverable state before they overwrite the working tree.
  • Accidental outbound disclosure of known credentials or host-identifying absolute paths can be checked through a cooperative text CLI. A caller that owns the emission can apply its redaction plan, escalate, or block.
  • Synthetic honeytoken experiments can exercise declared local channels with zero-token controls and fail-inconclusive evidence health checks.

The Core is harness-neutral and supports Python 3.9+ and Git. Automatic interception still depends on whether the host exposes a compatible hook; a Skill by itself does not intercept tool calls. The shared Core, Decision Protocol, and Skills define the product; harness adapters are replaceable integration bridges rather than the product boundary.

Agent Guard keeps reversible actions automatic and escalates only when it cannot safely automate them. It is reliability infrastructure, not a security sandbox: it protects against mistakes, not a malicious agent with equal OS privileges.

What it looks like

rm -rf build/       → RELOCATE   # an in-scope tree moves to quarantine
rm -rf .            → BLOCK      # the workspace root is protected
git reset --hard    → SNAPSHOT   # snapshot first when Git state supports it
git push --force    → BLOCK      # remote history is not automated

These are illustrative verdicts for supported inputs, not commands to run or proof that every harness intercepts them. Ignored, regenerable targets may be ALLOW; a Git snapshot that cannot be made fails closed. When recovery or safe rewriting is possible, the agent can keep working. Otherwise, the guard asks the human or blocks the operation.

Quick start with your coding agent

Use the scoped npm package after it is available in the registry, or keep a stable Git checkout. Never substitute the unrelated unscoped agent-guard package.

Pinned npm installation into a stable, user-owned prefix:

npm install --prefix /absolute/path/to/agent-guard-install @mokuyoaxis/agent-guard@0.2.2

0.2.2 remains the stable recommendation. The published prerelease @mokuyoaxis/agent-guard@0.2.3-rc1 is also available through npm @rc, but does not contain guard-lab. This checkout is the 0.2.3-rc2 source candidate; check Releases and npm before assuming rc2 is published. Prereleases do not replace npm latest.

The package root is then /absolute/path/to/agent-guard-install/node_modules/@mokuyoaxis/agent-guard. Alternatively, clone the source (skip this if you already have a checkout):

git clone https://github.com/mokuyoaxis/agent-guard.git
cd agent-guard

Python 3.9+ and Git are required for the Core; native interception depends on the host's hook support. Then give your coding agent the following setup prompt (replace the path with your checkout):

Set up agent-guard for this workspace. Use either an existing Git checkout or
the exact scoped npm package @mokuyoaxis/agent-guard@0.2.2; never install the
unscoped package named agent-guard. Before installing, ask me to choose and
approve a stable user-owned prefix. Treat the checkout or installed package
root as /absolute/path/to/agent-guard below.
First identify the current harness and its actual hook/skill capabilities;
read this README and the matching adapter README. Check Python and Git.
Install the relevant Skills, then configure a native shell hook only if this
harness supports one. Preserve existing settings and show me the proposed
diff before editing user-wide configuration or installing dependencies.
For Claude Code use adapters/claude/README.md; for Kimi Code use
adapters/kimi-code/README.md; for DSH use adapters/dsh/README.md.
For another host, read adapters/INTEGRATION.md and do not invent a native
hook. If no blocking pre-tool hook is verified, use only Skill/CLI and say
plainly that automatic interception is not enabled.
Verify a harmless command and pass a BLOCK-shaped command only as data to
check.py; never execute a destructive test command. For Claude/Kimi, run the
local doctor but do not treat its PASS as proof of host interception. Report
the host version, tool coverage, what was installed, what the host actually
intercepted, and any unverified paths.

For manual setup and evidence limits, see the adapter matrix and the adapter README for your host.

Design principles

PrincipleGuarantee
Stay in scopeThe guard blocks deletion of the workspace root, .git, and outside paths when the operation reaches it
Make it recoverableSupported deletions relocate to .agent-trash/ with a manifest; destructive Git overwrites snapshot first
Constrain authorizationAuthorization is session-scoped; a veto downgrades one-way, and only a human restores it
Leave a durable trailEnforced verdicts, compensation intents, outcomes, and restores use append-only JSONL; mutation fails closed if its intent cannot be stored

One rule runs through all four: uncertainty increases restriction.

How it decides

Each inspected operation is classified by its effect and then mapped to the least restrictive decision that preserves the relevant safety or recovery guarantee. The stable interface is a Decision Protocol, not a binary allow/block check:

Effect → Classifier → Policy → Decision   ∈ { ALLOW, SANITIZE, RELOCATE,
                                            SNAPSHOT, ASK, BLOCK }
                                + ReasonCode   (stable, machine-readable)
                                + Explanation  (human-facing)
                                + RecoveryPlan (txids, strategy)
TierDecisionsWhat the agent experiences
SAFEALLOW · SANITIZE · RELOCATE · SNAPSHOTRuns silently; compensation is applied first where needed; recoverable mutations are restorable via txid. SANITIZE returns a plan for the payload owner to rewrite (not a command rewrite)
AMBIGUOUSASKSingle-execution authorization (ASK_ONCE) — e.g. compound shapes the guard cannot safely automate
FORBIDDENBLOCKRefused with reason and remediation; never askable

Precedence when several decisions meet in one operation, weakest to strongest:

ALLOW < SANITIZE < RELOCATE < SNAPSHOT < ASK < BLOCK

SANITIZE ranks below ASK deliberately: it is automatic (SAFE tier), while ASK forfeits automation. A payload carrying both a sanitizable secret and a shape that cannot be rewritten must ASK — you cannot silently proceed when part of the emission is uninspectable.

True effect uncertainty ($VAR targets, bash -c, find -delete, stdin-fed lists) stays on the BLOCK path: allowing it would forfeit the core guarantee. Adapters map decisions onto their harness natively — DSH PreToolDecision, Claude Code PreToolUse ask, or a deny carrying the explanation where no ask exists.

What Agent Guard includes

delete-guard

Answers "if this destroys something, can we come back?" It runs before a delete or destructive Git action when invoked through a supported adapter or CLI, and compensates first when recovery is possible.

exfil-guard

Answers "if this leaves the machine, was it supposed to?" Its cooperative CLI checks text before an emission when the payload owner invokes it, returning a redaction or escalation decision for supported patterns.

recovery-audit

The incident-response companion for cases where prevention never ran or did not cover the path. It establishes source precedence, distinguishes recovered bytes from reconstructed behavior and known gaps, audits replay tooling, and keeps landing, commit, push, and release as separate authorization gates.

delete-guard and exfil-guard are the two preventive guard branches; recovery-audit handles evidence-led recovery after the fact.

recovery-audit

Sometimes prevention never ran: a harness had no adapter, a subagent bypassed the expected path, or an over-broad command removed the workspace before anyone could intervene. The working tree may be gone while the coding agent's session cache still preserves successful patches, file snapshots, tool results, diffs, and command context.

recovery-audit turns those remnants, Git remotes/reflogs/stashes, editor or tool caches, build artifacts, and project plans into an evidence-led recovery:

  • every unit is labelled recovered, reconstructed, or missing;
  • recorded tool effects are replayed in chronology and checked for divergence;
  • repeated replay must produce a byte-identical tree;
  • landing, commit, push, and release remain separate authorization gates.

It is not filesystem undelete and cannot recreate bytes no surviving source captured. Its promise is a fast, auditable path to the strongest project state the evidence actually supports, with gaps reported instead of hidden.

exfil-guard

exfil-guard checks text before an agent writes, sends, commits, or pushes it when the payload owner calls its CLI. It also offers an explicit, read-only safe view of selected JSON/dotenv configuration files. The text scanner is designed to catch two accidental disclosure classes: known credentials and host-identifying absolute paths. Depending on the channel, it can allow the payload, return a redaction plan, ask for a human decision, or block the emission.

DSH also offers a default-off text-read redaction prototype for complete native reads in a pinned composition. It reuses Core and regenerates both rendered text and presentation metadata; a zero-model native probe covers the next request and durable JSONL log. A separate official Flash direct-read trial observed supported synthetic-secret redaction while useful config stayed readable.

It is a prevention and redaction guard, not a compensation engine: after an emission there is nothing to recover. It is also not a security sandbox and does not attempt to stop adversarial exfiltration by an agent with equal OS privileges.

Decisions exfil-guard can return

The full Decision Protocol applies, but only four classes are reachable for a text payload (RELOCATE/SNAPSHOT belong to delete-guard — the guard cannot rewrite what it did not write):

DecisionMeaningExample
ALLOWnothing matched, a documented placeholder, or a workspace-relative pathecho "hello" | check_span.py
SANITIZEa redaction plan is returned; the payload owner rewrites and emitsa real key on file-write / llm-request
ASKthe channel cannot be rewritten and cannot be taken backa host path on shell-stdout
BLOCKrefuse: immutable/remote history, an un-scannable payload, invalid configa credential in git-push-payload

What is detected

T1 vendor credential patterns (secret/*, deterministic, near-zero false positives). Rule ids: secret/openai-key, secret/github-token, secret/aws-access-key-id, secret/gitlab-token, secret/slack-token, secret/stripe-key (live keys only — sk_test_ is exempt), secret/jwt (structural: the header must base64-decode to JSON containing alg), and secret/private-key-block (whole -----BEGIN ... PRIVATE KEY----- block, redacted in one piece). See skills/exfil-guard/references/rules.md for the frozen table.

Value-free secret references (secret/source-reference). The guard classifies an environment variable's name (*KEY*, *TOKEN*, *SECRET*, *PASSWORD*, *CRED*, *AUTH*) and a secret-store file name (.env, *.pem, id_rsa*, .netrc, kubeconfig, ...), and detects whole-environment expansions (printenv, env | ..., cat /proc/self/environ). This scanner does not resolve the variable value; that does not certify unrelated Guard output or existing audit records as secret-free.

Host-identifying paths (path/*). path/workspace-relative is ALLOW (the workspace is exempt); path/system (/usr, /etc, C:\Windows) is ALLOW; path/host-absolute (under HOME/TEMP, a CI root, or a workspace ancestor) is SANITIZE; path/generic-absolute (no host correlation) is ASK; path/device (UNC, \\?\, pipes) is SANITIZE.

Channels determine the disposition

A channel is defined by two facts: can it be rewritten, and does the emission persist? rewritable is what makes SANITIZE meaningful; persistence is what justifies BLOCK.

ChannelRewritablePersistenceDefault
llm-requestyesremoteSANITIZE
file-writeyesworkspaceSANITIZE
forge-comment / issue-body / pr-descriptionyespublicSANITIZE
git-commit-messageyes (rewrite argv)remote historyBLOCK
git-push-payloadnoremoteBLOCK
shell-stdoutnolocal transcriptASK
shell-file-redirectyeslocalASK
archive-uploadyesremoteASK
process-argvyeslocalASK

An unknown channel name is a configuration defect, not "no risk": check_span.py returns BLOCK_OUTPUT_UNSCANNABLE, never an implicit ALLOW.

Usage

check_span.py reads the payload on stdin and is a pure function — it never writes, never rewrites, and never prints the match. sanitize.py applies the plan the guard returned.

# a credential on a rewritable channel -> SANITIZE, exit 0
echo 'config: sk-proj-AbCdEf…' | python3 skills/exfil-guard/scripts/check_span.py --channel file-write

# a credential bound for remote history -> BLOCK, exit 2
echo 'token=ghp_abcdefghijklmnopqrstuvwxyz…' | python3 skills/exfil-guard/scripts/check_span.py --channel git-push-payload

# apply the redaction plan (format preserved: sk-<REDACTED>)
echo 'config: sk-proj-AbCdEf…' | python3 skills/exfil-guard/scripts/sanitize.py --channel file-write

Exit code contract: 0 = ALLOW/SANITIZED · 2 = BLOCK · 3 = ASK · 1 = ERROR. Use --json for the machine-readable verdict (offsets, rule ids and placeholders only — never the matched bytes), and --path to enable the repo-local exemption file for the file being written.

Read a config without printing its values

Introduced in the 0.2.0 source; this CLI is not in the earlier 0.2.0-rc2 preview tag.

python3 skills/exfil-guard/scripts/view.py --workspace /path/to/workspace .env
python3 skills/exfil-guard/scripts/view.py --workspace /path/to/workspace config.json

The path must be relative to that workspace. The JSON output preserves field names and structure, plus scalar types and set/empty states, but never scalar values. Known secret-shaped field names are hidden; unknown secrets in field names remain a limitation. Only UTF-8 JSON and a strict, single-line dotenv subset are supported (256 KiB maximum, 16 levels, 2048 nodes). The CLI refuses symlinks, hardlinks, special files, unsafe paths, malformed input, and platforms without safe descriptor-relative reads. Exit 0 means a view was produced, 2 means refused, and 1 means an internal error. This view is for diagnosis only: do not write it back over the original config. It does not intercept ordinary file reads made by a harness.

Relationship to delete-guard

They are two halves of the same promise, on opposite sides of the action:

delete-guardexfil-guard
Question"can we come back?""was this supposed to leave?"
Guardsbefore a deletebefore an emission
Responsecompensate, then proceedredact, then emit
Failure costrecoverable via txidirreversible
Entry pointcheck.py -- <command>check_span.py (stdin)

They share the vocabulary (core/policy.py), the aggregation (worst()), the exemption discipline, and the audit log. worst() is shared by both guards, which is why SANITIZE had to be ranked once, not twice.

Coverage and limitations

This is stated plainly, because a reliability tool that overstates its reach is a false security claim:

  • Not a sandbox. It does not prevent adversarial exfiltration. An agent that obfuscates a secret to evade the scanner is out of scope; this catches accidents.
  • Channels with no hook are unreachable by construction. A hosted model call with no proxy, the model's own tool calls, content produced inside a program, and the human clipboard get no verdict at all — no coverage is claimed there. See references/channels.md and docs/secret-guard-analysis.md §2.4 for the reachability table this claim is traceable to.
  • Not a file scanner. It is not a gitleaks replacement; it scans what the guard can see on the way out.
  • No history rewriting. Detecting a secret already in Git history is a report at most. Rewriting history is a human action with its own risks.
  • No T3 entropy detector in this release. It is the single largest false-positive source, and the named scenarios do not require it.

Manual, harness-neutral usage

The Core has zero third-party dependencies. Requirements: Python 3.9+, POSIX shell, and Git.

# delete something - it is quarantined, not destroyed:
python3 skills/delete-guard/scripts/safe_delete.py build/ --reason "stale"

# inspect and undo:
python3 skills/delete-guard/scripts/status.py
python3 skills/delete-guard/scripts/restore.py list
python3 skills/delete-guard/scripts/restore.py <txid>

# quarantine maintenance (dry plan by default):
python3 skills/delete-guard/scripts/gc.py

A harness adapter can invoke the guard before a supported shell command and map its exit status to the host's own tool decision:

python3 skills/delete-guard/scripts/check.py --enforce -- "$COMMAND"
exit 0 → host may run the original command
exit 2 → deny
exit 3 → ask the human if supported; otherwise deny
exit 1 → guard error; fail closed

What gets protected

These examples assume the command reaches the guard and its targets meet the stated conditions; see the capability matrix for what each host has actually demonstrated.

rm -rf build/            → RELOCATE  (tree quarantined, command proceeds)
rm -rf .                 → BLOCK     (workspace root)
rm -rf $DIR/             → BLOCK     (unresolvable target: fail closed)
rm *.log                 → BLOCK     (opaque glob; safe_delete expands it)
cd X && rm -rf build     → ASK_ONCE  (COMPOUND_CWD_DELETE)
touch f && rm f          → ASK_ONCE  (COMPOUND_CREATE_DELETE)
git clean -fd            → RELOCATE  (enumerate via -n, relocate, proceed)
git reset --hard         → SNAPSHOT  (when a Git snapshot can be made)
git push --force         → BLOCK     (remote history is never automated)
node_modules/ (ignored)  → ALLOW     (provably regenerable)
quarantine full          → BLOCK     (never fall back to permanent delete)

Integration and validation matrix

Node.js 20 smoke Codex Skill/CLI tested DSH 0.1.5-rc.1 host BLOCK tested ZCode win32 CLI evaluated Claude Code hook tested with scripted model Kimi Code 2.1.1 K3 Bash BLOCK

"The Core works", "a Skill-guided agent used it", and "the harness intercepts every matching tool call" are separate claims. This table keeps those evidence levels explicit:

Harness / tested versionIntegration pathEvidence and limit
Claude Code 2.1.270 / 2.1.273Native PreToolUse for BashReal CLI + scripted model: sampled allow/ask/deny and Python-startup failure; other tools unverified.
DSH 0.1.5-rc.1Native pre-execute adapterReal-host packaged-plugin probe: execution-level BLOCK and loaded-adapter Core failure. A real-model Lab baseline stayed quiet with guard off, so no L2 mitigation claim exists; model-proposed destructive bash enforcement remains unverified. The separate opt-in read path is listed below.
DSH CLI rc.1 / tools and FS rc.2 / Node 22Experimental text-read redaction, default offZero-model native probe: next request and durable JSONL. Official Flash direct-read off/on: supported synthetic secrets redacted in tool content/meta/session, useful config preserved; no injection L2 claim.
DSH 0.2.0-rc.2 / Node 22Default deletion and optional text-read adapterFresh contract review: native deletion blocking, read redaction on reviewed local/sandbox FS, next synthetic request and durable JSONL. Native v4 Lab support remains bounded; no new real-model L2 result.
Codex CLI 0.154.0 (tested session)Skill + production CLIOlder-source cooperative acceptance; no native hook claim.
ZCode (version unrecorded; win32)Historical Skill/CLI + hook trialOlder hook observation includes a persistent-permission bypass; current version unverified.
Kimi Code 0.42.0 / 2.1.1Native PreToolUse for BashBounded tests: authenticated 2.1.1 root-Bash PASS on an OAuth official model and maintainer-confirmed official K3 relay; older 0.42.0 root/child BLOCK, ASK denial and Python-failure refusal. Hook absence/timeout remains fail-open.
Other / unlisted hostsSelf-adaptation guideNo native claim without a blocking pre-tool event and independent non-execution check.

For agent-led setup: identify the actual host version and tool names, follow the matching guide above, preserve existing settings, then report separate configuration, local-probe, and real-host evidence. For an unlisted host, follow the self-adaptation checklist; a prompt or adapter exit code alone is not proof of interception. Ask before changing user-wide settings, security policy, or dependencies.

tests/test_conformance.py checks the shared Core and Claude adapter; DSH has a smoke test and Kimi has targeted adapter tests. The Kimi observations above cover only the tested calls, not every shell construct or a general concurrent-agent safety guarantee. Run python3 doctor.py kimi --probe or python3 doctor.py claude --probe to check a selected configuration file and the local shell bridge without a model call. Add --json for machine-readable configuration, local_probe, and host_interception statuses. The last status is always UNVERIFIED: this doctor cannot prove that a live session loaded or enforced the hook. Add --check-drift to query the local host version and compute privacy-safe configuration/runtime fingerprints. It reports CURRENT, STALE, DRIFTED, BROKEN, or UNVERIFIED; even CURRENT is static preflight evidence, not interception proof. Optional create-new baselines contain only the parsed version and SHA-256 fingerprints. See the host drift guide for the status and baseline contract. An explicit --live-sentinel can then spend one configured model call in a private blank fixture. It reports PASS, FAIL, or INCONCLUSIVE with NONE/NOTICE, CRITICAL, or WARNING; marker absence alone never passes. It retains a hashed/redacted result but not raw model output. This is a reliability probe for a trusted provider, not a sandbox for a malicious model. The optional Kimi doctor needs Python 3.11+ for TOML parsing; the Core and hook adapter continue to support Python 3.9+. See the Kimi and Claude adapter guides. See harness capabilities and evidence levels for the per-host scope and execution-level acceptance criteria.

Repository layout

agent-guard/
├── skills/delete-guard/   # agent-facing skill: SKILL.md + CLI scripts
├── skills/exfil-guard/    # egress skill: check_span.py · sanitize.py
├── skills/recovery-audit/ # evidence-led repository audit and recovery
├── core/                  # classifier · policy · recovery · audit · redaction
├── doctor.py              # local configuration, probes, and drift preflight
├── live_sentinel.py       # opt-in real-host sentinel and alarm evidence
├── guard_lab.py           # user-controlled offline synthetic honeytoken lab
├── adapters/INTEGRATION.md # checklist for an unlisted host
├── adapters/claude/       # Claude Code PreToolUse hook adapter
├── adapters/kimi-code/   # Kimi Code PreToolUse hook adapter
├── adapters/dsh/          # DeepSeek Harness integration bridge
├── adapters/codex/harness/# CLI acceptance driver; not a native hook
├── tests/                 # unittest suites incl. cross-harness conformance
└── docs/                  # architecture · threat-model · friction log

Skills guide agent behavior; constraints live in Core. Future git-guard, database-guard, cloud-guard skills plug into the same compensation engine without restructuring.

Documentation

ReadFor
CONTRIBUTING.mdhow to propose an Issue and submit a focused Pull Request
docs/architecture.mdpillars ↔ components, data flow, design decisions
docs/host-drift.mdzero-token host/version drift states and privacy-minimal baselines
docs/guard-lab.mdsynthetic Lab workflow, evidence semantics, and unsupported channels
docs/release-notes-0.2.3-rc2.mdoffline guard-lab source candidate
docs/release-notes-0.2.3-rc1.mdWindows Core fail-closed candidate and prerelease channel
docs/release-notes-0.2.2.md0.2.2 host drift, live sentinel and DSH package acceptance
docs/release-notes-0.2.1.md0.2.1 adapter/doctor changes and bounded evidence
docs/release-notes-0.2.0.md0.2.0 changes, evidence levels, and known limits
adapters/INTEGRATION.mdself-adaptation checklist for an unlisted host
docs/threat-model.mdhonest limits: what this is and is not
docs/friction.mdwhat real agents and host processes taught us (F1–F20)
docs/development-note-unguarded-deletion.mdde-identified incident exploration and the recovery-aware direction
docs/test-report-codex-gpt-5.6-sol.mdv0.1.1 Codex evaluation (medium + high)
docs/test-report-dsh-0.1.5-rc.1.mdcurrent DSH packaged-plugin and execution-level bash acceptance
docs/test-report-dsh-guard-lab.mdbounded DSH real-model Lab baseline; no L2 claim
docs/test-report-dsh-v0.1.1.mdv0.1.1 DSH live test (DeepSeek V4 Pro high, minimal mode)
skills/recovery-audit/SKILL.mdevidence hierarchy, deterministic replay, recovery and landing gates
skills/delete-guard/references/policy.mdfull rule table and decision codes
skills/exfil-guard/references/rules.mdexfil rule table, reason codes, exemption format, audit shape
skills/exfil-guard/references/channels.mdegress-channel taxonomy and the unreachable channels

Status & roadmap

The source version is v0.2.3-rc2; the latest stable release remains v0.2.2. The published v0.2.0 baseline includes recoverable destructive actions, cmd/PowerShell dialect parsing, the exfil-guard text CLI, a read-only config view, and recovery-audit. Later source releases add Claude/Kimi POSIX hook bridges, bounded host evidence, a local doctor, host-drift checks, an opt-in live sentinel and current DSH package acceptance. The rc2 source additionally contains the offline guard-lab MVP described below. See the 0.2.3-rc2 notes for its source scope, 0.2.3-rc1 notes for the published candidate, 0.2.2 notes for the stable scope, and GitHub Releases for publication status; a source version is not itself a published Release.

The release identity is harness-neutral. Existing DSH and Claude adapters, Codex/ZCode acceptance evidence, and the bounded Kimi exercise are entries in a growing compatibility matrix, not separate definitions of the product. The Kimi result includes sampled hook-compensation and execution-level BLOCK paths, not proof that its host enforces every BLOCK or that arbitrary agent actions are protected. A tested direct-Python hook startup failure was fail-open. The optional shell bridge blocked a sampled Python startup failure in a real Kimi host session, but only when the bridge itself ran; config validation alone cannot prove that the hook is live. The candidate adds a focused real-Windows Core gate for Issue #7, but native cmd/PowerShell execution through every harness and general concurrent-subagent safety remain explicit gaps. Later Guard branches (git-guard, database-guard, cloud-guard) reuse the same protocol and compensation engine.

0.2.3-rc2 source candidate: guard-lab

The current source adds an opt-in, offline honeytoken MVP using disposable projects and synthetic, non-secret markers. Its clean, mock-positive, mock-injection, and snapshot-positive controls call no model and make no external request. A bounded loopback observer records the fake Lab stub, bait URL, or a validated synthetic snapshot; an explicit post-run scan can detect a canary in selected output without retaining the raw value.

The user keeps the true controller and evidence outside the fixture shown to the tested agent. Broken controls or changed evidence produce INCONCLUSIVE, never an optimistic pass. This first candidate cannot observe ordinary file reads, prove remote transfer, universally block LLM calls, or resist a same-UID adversary. Start with the zero-token controls and read the guard-lab guide before involving a real harness. The built-in cases report CALIBRATION_ONLY. A separate manual injection-probe treats bait contact as EXPOSURE_OBSERVED, and compare accepts a mitigation result only after an effective unguarded baseline and a matched guarded trial with the same protocol, harness/version/model, group, and task hash. Manual real-model runs also require a user-verified completed host result, so an exit-0 startup/model failure cannot count as a quiet win. A quiet baseline stays INCONCLUSIVE; no result is a general model or vendor safety rating.

Contributing

Bug reports, design proposals, compatibility evidence, documentation fixes, and focused code changes are welcome. Please start with an Issue for non-trivial or security-boundary changes, and submit implementations as a focused Pull Request. Read the contribution guide before sharing logs or test evidence; credentials, private configuration, and unredacted incident data must not be posted publicly.

Community

LINUX DO community link

This project-made banner links to our LINUX DO project post. It does not imply official endorsement by the community.

License

MIT — see LICENSE.

Related plugins