Skip to main content
G

dsh-key-rotation

goodandready/dsh-key-rotation

Keeps a pool of API keys per provider with auto-created clone routes, switches to the next key on quota or rate-limit errors, probes cooled keys for recovery, and exposes the pools, cooldown and switch codes in a Settings section.

Install

dsh plugin --profile web add github:goodandready/dsh-key-rotation

README

📦 @goodandready/dsh-key-rotation

Enterprise-Grade Transparent API Key Rotation, Rate-Limit Pre-emption & Failover Cascade for DeepSeek Harness

npm version license DSH Plugin Node version

GoodAndReady Showcase

🇬🇧 English • 🇷🇺 Русский • 🇨🇳 中文说明

⭐ If you like this plugin, please star it on GitHub — it shows me that the plugin is useful to you and motivates me to keep developing it.

🐛 If you find a bug or would like to request a feature, open a GitHub issue in any language — I will review your proposal and implement useful suggestions in a future plugin version.

⚡ Overview & The Problem

🚀 What's New in v0.8.11 (One-click Updater & Quality Gate)

  • Plugin updater in Settings: see current/latest version and update from the card without leaving DSH (#307).
  • Safer best-effort side effects: intentional non-critical failures log at debug instead of silent empty catch (#315).
  • Theme-only client colors and production-path test cleanup (#311, #314).
  • Leaner publication set: agent-only files no longer ship in git/npm (#308).

🚀 What's New in v0.8.10 (Stream Concurrency Hardening & Auto-Pruning)

  • Zero Concurrency Leaks: Guaranteed release of stream concurrency slots via deterministic try ... finally block, preventing key starvation during clean finishes or client stream aborts.
  • Robust Probe Retry: Added transient network socket error retry (PROBE_RETRY_DELAY_MS) in SandboxRunner.probeModels before marking keys as broken.
  • Memory & State Pruning: Automated pruning of removed/stale keys from internal pool maps during periodic sweep cycles.

🚀 What's New in v0.8.9 (Routing Evolution & UI)

  • Proactive Rate-Limit Guard: Automatic key pausing based on x-ratelimit-remaining-* and Retry-After headers before hitting 429 errors.
  • Self-Healing / Auto-Unbreak: Periodic background probe via free /models endpoint to automatically revive broken keys without token burn.
  • Latency-Aware Routing: Selectable strategies: round-robin, least-loaded (concurrency), and lowest-latency (p95 latency).
  • UI Evolution in dsh-clinebot Style: Sequential batch "Test All Keys" runner with live progress, Live Event Stream drawer, and Quota Reset countdown badge.
  • Native Dual Language Support: Native English (en) and Chinese (zh) UI dictionaries.

🛠️ What's New in v0.8.0 (Stability)

  • 🔌 Circuit breaker: after N consecutive provider failures the circuit opens and requests fail fast (CIRCUIT_OPEN) until a cool-down; half-open probes recover automatically.
  • 🕒 Monotonic clock: cooldown/breaker durations use process monotonic time so NTP steps cannot invert remaining times.
  • 📮 Non-blocking webhooks: alerts go through a bounded queue with backoff — stream rotation never waits on webhook HTTP.
  • 🧱 Atomic I/O helpers: crash-safe writes; corrupt JSON never overwrites previous in-memory state.
  • 🧹 Clone-route GC: orphaned auto-created clone routes are dropped from the runtime set.
  • 🧭 Error taxonomy: explicit switch/surface/soft classification for 408/425/429/5xx, sockets and gRPC codes.
  • 📡 Status extras: per-provider circuit plus meta.expectedClones / meta.notifyQueue.
  • 🧪 Smoke harness: scripted 429 → next-key → success path in test/smoke-rotation-080.test.mjs.

🛠️ What's New in v0.7.33 (Stability & Bugfix Release)

  • 🔍 Resolved Key Probing BaseURL: Fixed resolveBaseUrl to map key credential refs to owning provider pools, restoring live probeModels testing.
  • 🛡️ Guarded Cascade Recursion: Prevented call stack overflow in cross-provider failover when circular cascade chains occur.
  • 🕒 Accurate Midnight PST Resets: Corrected UTC-8 timezone calculation offset sign for calendar quota reset windows.
  • 🧹 Lifecycle Timer Cleanup: Wrapped canaryTimer and selfHealTimer in Cordis effect scopes, eliminating background orphaned intervals on hot reload.
  • ⚡ Stale Lock Recovery in Load Balancer: Added expired lock detection to pickLeastLoaded for uninterrupted least-connections routing.
  • 📊 Load Distribution & Modals (Changed in v0.8.5): Interactive segmented load distribution charts per pool, modal action confirmation, and 429/5xx backoff jitter (#283, #284).
  • 🎨 Native Design System (Changed in v0.8.3): Unified with dsh-clinebot baseline: modular section cards, live pool telemetry stat boxes, pill badges, and complete semantic theme token styling (#281).
  • 🌐 Localization (Changed in v0.8.2): Source strings are English-only. Russian/Chinese UI comes from the DSH core locale service and translation plugins (props.t). Active locale: host snapshot → first navigator.languages entry → en (#277).

🚀 What's New in v0.7.31

  • ⚡ O(1) TokenBucket Accumulator: Upgraded rate limiting math to O(1) time and zero-allocation memory with adaptive header synchronization.
  • 🛡️ Soft vs Hard Backoff: Differentiates transient infrastructure drops (502/503/timeouts: 10s flat cooldown) from hard quota errors (progressive doubling).
  • ⏳ Penalty Decay: Stable keys that operate cleanly automatically decay their failure penalty multiplier every hour.
  • 🎲 Cooldown Jitter: Adds ±12.5% random dispersion to recovery timers, eliminating thundering herd stampedes.
  • 🎯 Addressable Canary Probing: Support for probing target pool models with lightweight single-token verification pings.
  • 📊 TTFT Percentiles (p50 / p95 / p99): Sub-second high-resolution latency percentile tracking across all key pools.
  • 🔔 Webhook Alert Digest: Aggregates multiple rapid switch/cooldown events into consolidated incident digests for Telegram, Discord, and Slack.
  • 🧹 30-Day Usage Compaction: Automatic bounded memory management with 30-day rolling window data pruning.
  • ✨ Optimistic UI & Filter Pills: Instant zero-latency UI updates on reset, plus All, Ready, In Cooldown, and With Errors quick filter chips.

High-throughput autonomous agent workflows, parallel subagent swarms, and multi-turn tool loops inevitably hit upstream API rate limits (HTTP 429, RPM/TPM exhaustion, daily quotas, or sudden provider outages). In standard DeepSeek Harness deployments, a single exhausted API key breaks the entire agent execution chain, requiring manual intervention and destroying the session's replay state.

dsh-key-rotation provides a seamless, enterprise-ready transparent API key pooling, pre-emptive rate-limiting, and cross-provider failover engine built natively on the Cordis microkernel architecture.

Unlike naive routing proxies that alter provider identifiers, dsh-key-rotation hooks into ctx.credentials.resolve and intercepts llm/stream at runtime:

  • The provider identity never changes: Agent replay states, multi-call turns, and tool schemas remain 100% consistent.
  • Pre-emptive Token Bucket: Throttled keys are skipped before issuing network calls, eliminating retry latency.
  • Least-Connections Concurrency Control: Balances in-flight streams across keys to prevent burst saturation.
  • Autonomous Self-Healing & Cascades: Lifts expired quarantines on idle keys and smoothly escalates to fallback providers if an entire pool is exhausted.

🏗️ Architecture & Request Lifecycle

graph LR
    subgraph ClientLayer ["Client & Agent Turn"]
        UserMsg["User / Subagent Message"] --> Adapter["pi-ai Model Adapter"]
    end

    subgraph RotationEngine ["dsh-key-rotation Core Engine"]
        Adapter --> StreamHook["llm/stream Interceptor"]
        StreamHook --> BucketCheck{"Token Bucket\nRPM / TPM Check"}
        BucketCheck -->|Under Limit| ConcurrencyCheck{"Concurrency Tracker\nLeast-Connections"}
        BucketCheck -->|Exceeded| NextKey1["Pick Next Healthy Key"]
        ConcurrencyCheck -->|Slot Available| KeyResolver["ctx.credentials.resolve"]
        ConcurrencyCheck -->|Saturated| NextKey1
        
        KeyResolver --> ActiveKey["Active Key (In Use)"]
        
        ActiveKey -.->|HTTP 429 / Quota / Error| Failover["Instant Failover Handler"]
        Failover --> BackoffCalc["Exponential Backoff & Quarantine"]
        Failover --> NextKey2["Retry Next Key (Zero Token Loss)"]
        Failover -.->|All Pool Keys Exhausted| CascadeEngine["Cross-Provider Cascade"]
        
        BackoffCalc --> QuotaWindow["Calendar Reset / Midnight Window"]
        BackoffCalc --> SelfHeal["Self-Heal Idle Sweep"]
        SelfHeal -->|Cooldown Expired| PoolReady["Restored to Ready Pool"]
    end

    subgraph UpstreamLayer ["Model Provider Endpoints"]
        ActiveKey --> UpstreamAPI["Primary Provider API"]
        CascadeEngine --> FallbackAPI["Backup Provider API"]
    end

    style ClientLayer fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
    style RotationEngine fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
    style UpstreamLayer fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4

✨ Full Feature Breakdown

🔄 1. Transparent Rotation & Failover

  • Unchanged Provider Identity: Rotates only the underlying resolved API credential ref, never the provider ID. Prevents INVALID_REPLAY_STATE crashes in pi-ai multi-turn sessions.
  • Zero-Token-Loss Stream Retries: If an API key encounters an error before the first content chunk is emitted, the request is transparently re-dispatched to the next healthy key in the pool.
  • Comprehensive Switch Codes: Automatically fails over on QUOTA, RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT, EMPTY_RESPONSE, UNKNOWN_MODEL, AUTH, and INVALID error codes.
  • Intelligent Message Pattern Matching: Fallback regex classifier (SWITCHABLE_MESSAGE_PATTERN) identifies text-based quota/rate-limit errors thrown as generic exceptions by upstream SDKs.
  • Non-Streaming Safety Net: Synchronous calls (e.g., embeddings, batch evaluations) are protected via the agent/request-error lifecycle hook.

⏱️ 2. Rate-Limit Pre-emption & Concurrency Control

  • Token Bucket / Leaky Bucket (lib/bucket.js): Sliding-window tracking of Requests Per Minute (rpmLimit) and Tokens Per Minute (tpmLimit). Quarantines saturated keys before dispatching network requests, preventing 429 roundtrips.
  • Least-Connections Balancer (lib/concurrency.js): Tracks active in-flight streams per key (inFlight). Distributes concurrent requests evenly across available credentials and enforces maxConcurrency limits.
  • Stale Lock Auto-Release: Deadlocks from disconnected clients or aborted network sockets are automatically purged after 5 minutes.

🛡️ 3. Autonomous Healing & Cascade Escalation

  • Cross-Provider Failover Cascade (lib/cascade.js): If all keys for a selected provider are in cooldown, requests automatically cascade to an alternative fallback provider pool (e.g., primary provider → fallback proxy / secondary provider).
  • Sandbox Key Probes (lib/sandbox.js): On-demand /models probes validate a key before returning it to rotation; idle cooldowns are lifted by the self-heal sweep.
  • Calendar & Rolling Quota Reset Windows (lib/quota-window.js): Supports scheduled quota reset alignments (midnight_utc, midnight_pst, and rolling_24h) so daily free/tier quotas unfreeze exactly when upstream resets them.
  • Adaptive Exponential Backoff (lib/pool.js): Successive failures on a key double its quarantine duration (base → ×2 → ×4 → cap ×8). Successful requests gradually restore healthy status.

🎯 4. Model-Aware Routing

  • Model Sub-Pools (lib/pool.js): Configure dedicated key pools for specific model tiers (e.g. reasoning/heavy models vs fast/cheap utility models).
  • Per-Model Per-Key Token Quotas (lib/model-quota.js): Give one credential its own local token budget for one model. A key that spends its budget is skipped for that model only — an exhausted claude-sonnet budget never disables the same key for claude-opus.
  • Tag-Based Routing: Assign operational tags (production, background, eval) to match key usage with workload priorities.
Per-Model Per-Key Token Quotas

tokenLimit applies to one credential in one configured model pool. The same credential can hold an independent budget in every model pool it belongs to:

dsh-key-rotation:
  quotaResetWindow:
    type: midnight_utc
    hour: 0

  providers:
    - provider: anthropic

      keys:
        - CLAUDE_KEY_A
        - CLAUDE_KEY_B

      models:
        claude-sonnet:
          keys:
            - CLAUDE_KEY_A
            - CLAUDE_KEY_B
          quotas:
            CLAUDE_KEY_A:
              tokenLimit: 1000000
            CLAUDE_KEY_B:
              tokenLimit: 1000000

        claude-opus:
          keys:
            - CLAUDE_KEY_A
            - CLAUDE_KEY_B
          quotas:
            CLAUDE_KEY_A:
              tokenLimit: 200000
            CLAUDE_KEY_B:
              tokenLimit: 200000

Behaviour:

claude-sonnet request
  → CLAUDE_KEY_A still has Sonnet budget
  → dispatched on CLAUDE_KEY_A
  → actual usage is charged to CLAUDE_KEY_A / Sonnet
  → CLAUDE_KEY_A / Sonnet reaches its limit
  → later Sonnet requests skip CLAUDE_KEY_A and use CLAUDE_KEY_B
  → claude-opus requests may still use CLAUDE_KEY_A

Rules and limits worth knowing:

  • tokenLimit is the number of tokens one credential may spend on one model pool within the current quota window. A missing, null, non-numeric or non-positive value means no local model token limit — omitting quotas entirely leaves behaviour byte-identical to previous releases.
  • quotaResetWindow controls reset timing for these budgets, reusing the existing midnight_utc / midnight_pst / rolling_24h settings. Resets are lazy: the counter returns to zero when the window elapses, with no per-key timer.
  • Usage is tracked from the usage returned by successful LLM responses. Only a completed request with usable usage is charged. Failed, aborted or refused requests are never billed.
  • Providers that do not report usage cannot be tracked precisely. Such requests are not guessed at and not deducted, so a model budget can only be exhausted by responses that actually reported token usage.
  • The final request that fits is allowed to overshoot the configured limit slightly, and concurrent in-flight requests can also overshoot. This is expected behaviour in this version — there is no token reservation.
  • Local model quotas fail closed: when no credential has budget left, the request is not sent upstream and the original credential is not used as a fallback. The pool enters the existing exhaustion / cascade flow instead.
  • Exhaustion is a budget state, not a credential failure. It never sets a cooldown, a failure count or a broken flag, and upstream QUOTA errors stay isolated to the model pool that served the request.
  • Editing a limit takes effect immediately: lowering it below the tokens already spent marks the credential exhausted at once, raising it restores headroom without clearing usage, and deleting quotas.<REF> restores Unlimited.
  • Only credential refs and counters are stored in the state file. The limit itself stays in Settings Config, and no API key value is ever written to disk or returned by the status API.

📊 5. Observability, Telemetry & Webhooks

  • Interactive Multi-Platform Webhooks (lib/webhook.js): Dispatches rich notifications with HMAC-signed action buttons for Telegram (Inline Keyboards), Discord (Action Rows), and Slack (Block Kit). Administrators can click buttons to reset cooldowns or pause providers directly from their mobile chat.
  • Usage & Cost Reporting (lib/usage-report.js): Per-key daily request counters and estimated cost breakdown with one-click CSV/JSON export (GET /dsh-key-rotation/usage-report).
  • Latency SLO & Histogram (lib/histogram.js): Tracks Time-To-First-Token (TTFT) and stream durations with health score degradation scoring (0..100).

🔁 8. One-click Plugin Updater

The settings card includes an Updater section:

  1. Check for updates — GET /api/dsh-key-rotation/update returns currentVersion, latestVersion, updateAvailable, canAutoUpdate (version metadata only; no secrets).
  2. Update now — POST with the same path installs the exact latest npm version through the standard dsh plugin add flow. The request is accepted only from loopback with a matching same-origin Origin/Host (and the dedicated header). Cross-origin or missing-origin POST is rejected with 403.
  3. After a successful install the UI tells you to restart DSH so the new host code loads.

No --force, no raw shell, no install from worktree/DEV paths. Update runs only after an explicit click.

🖥️ Rich Web GUI & Dashboard

Access full visual management under Settings → Key Rotation or via the Header quick-widget.

Interface FeatureDescription
Header Status WidgetCompact live badge in DSH header: 🟢 All Healthy | 🟡 Cooldown Active | 🔴 Pool Exhausted with quick popover actions.
1-Click Health Matrix"Health Matrix" dashboard running parallel sandbox probes across all providers, keys, and models with TTFT latency and status badges.
Instant Key ProvisioningAdd keys with auto-generated names (<PROVIDER>_API_KEY, _2, _3) and automatic key-tail disambiguation.
Live Status BadgesVisual states: In Use, Ready, Cooling Down (with live countdown timer), and Not Found.
Drag & Priority OrderingReorder keys with ↑ and ↓ buttons to fine-tune selection precedence.
Switch Code TogglesInteractive checkboxes for switchable error conditions.
Secret Leak DetectorReal-time input sanitizer (lib/keycheck.js) catching accidental pastes of private keys, SSH keys, or misplaced tokens.
Batch .env ImportParse standard .env key-value pairs directly into corresponding provider pools.
5-Second Undo BarNon-destructive undo bar for accidental key or pool removals.
Model Sub-Pools & Token Quota EditingMaintain per-model key lists under each provider, with per-key token limits, live used/limit, percentage, remaining tokens and reset countdown.
Usage Analytics ChartInteractive breakdown of lifetime requests and daily trends per key.

🔒 Security & Safe Storage

  • Zero Plaintext Secrets in Plugin Config: Configuration files store only environment variable reference names (e.g. MY_PROVIDER_API_KEY).
  • Secure Vault Storage: Actual secret values reside securely in $DSH_HOME/.credentials.yaml managed by the DSH Credentials service.
  • 5-Character Masking (keyTail): Full secret values are never sent to the client browser; only the trailing 5 characters are exposed for visual identification. Short keys (<= 5 characters) return a fixed masked placeholder (***) to prevent credential disclosure.
  • Fail-Closed Loopback & Same-Origin Fencing: Administrative endpoints strictly enforce loopback checks (isTrustedBridgeRequest): the socket peer and the Host header must both be loopback (127.0.0.1, ::1, localhost), and sec-fetch-site: cross-site is always refused. An attached Origin must be an http(s) loopback origin matching Host exactly. Origin is required on POST/PUT/PATCH/DELETE but optional on GET/HEAD/OPTIONS, because browsers omit it on same-origin reads — requiring it there would reject the Settings card's own requests.
  • Fail-Closed Resolver on Pool Exhaustion: When all credentials in a managed pool are exhausted, paused, expired, or blocked by RPM/TPM limits, the resolver fails closed with LOCAL_POOL_EXHAUSTED error rather than falling back to leaking unmanaged credentials.
  • SSRF Protection & Rebinding Guard: Remote provider pool import strictly enforces HTTPS-only URLs, validates all resolved IP addresses including IPv4-mapped and IPv4-compatible IPv6 addresses (::ffff:127.0.0.1, ::ffff:7f00:1, 64:ff9b::/96), and enforces connect-time DNS validation via undici agent dispatchers to prevent TOCTOU DNS rebinding.
  • Cross-Pool Revocation & Inheritance: Model pools automatically inherit pause, revoke, and expiry states from their base provider; runtime 401 permanent authentication failures revoke the credential across all shared pools immediately.

📦 Installation

# Install via DSH Plugin Manager (Web Profile):
dsh plugin --profile web add @goodandready/dsh-key-rotation

# Or directly from GitHub:
dsh plugin --profile web add github:GooDAnDReaDY/dsh-key-rotation

[!IMPORTANT] Restart DeepSeek Harness web service after installation and refresh your browser tab:

systemctl --user restart dsh-web

⚙️ Configuration Reference (settings.yaml)

dsh-key-rotation:
  switchCodes:
    - QUOTA
    - RATE_LIMIT
    - SERVER
    - TIMEOUT
    - TRANSPORT
    - EMPTY_RESPONSE
    - UNKNOWN_MODEL
    - AUTH
  cooldownMs: 60000
  # v0.8.0 circuit breaker
  circuitBreakerEnabled: true
  circuitBreakerThreshold: 5
  circuitBreakerOpenMs: 30000
  circuitBreakerHalfOpenProbes: 1
  concurrencyLimit: 5
  quotaResetWindow:
    type: midnight_utc
    hour: 0
  cascade:
    - provider: backup-provider-id
      model: your-backup-model-id
  webhookUrl: "https://api.telegram.org/bot<TOKEN>/sendMessage?chat_id=<CHAT_ID>"
  providers:
    - provider: your-primary-provider
      rpmLimit: 60
      tpmLimit: 100000
      keys:
        - PRIMARY_API_KEY
        - PRIMARY_API_KEY_2
        - PRIMARY_API_KEY_BACKUP
      models:
        reasoning-model-id:
          keys:
            - PRIMARY_API_KEY
            - PRIMARY_API_KEY_2
          quotas:
            PRIMARY_API_KEY:
              tokenLimit: 500000
            PRIMARY_API_KEY_2:
              tokenLimit: 1000000
    - provider: secondary-provider
      keys:
        - SECONDARY_API_KEY
        - SECONDARY_API_KEY_2

Parameter Reference

ParameterTypeDefaultDescription
switchCodesstring[][QUOTA, RATE_LIMIT, ...]List of error codes that immediately trigger failover.
cooldownMsnumber60000 (1 min)Base penalty duration (in ms) for quarantined keys.
circuitBreakerEnabledbooleantrueEnable per-provider circuit breaker (v0.8.0).
circuitBreakerThresholdnumber5Consecutive failures before opening the circuit.
circuitBreakerOpenMsnumber30000How long the circuit stays open (ms).
circuitBreakerHalfOpenProbesnumber1Probe requests allowed in half-open state.
verboseLoggingbooleanfalsePer-request rotation logs (noisy; off by default).
concurrencyLimitnumber0 (disabled)Max concurrent in-flight streams per key (0 = unlimited).
quotaResetWindowobjectnullCalendar reset alignment for both provider-reported quota cooldowns and local per-model token budgets (midnight_utc, midnight_pst, rolling_24h).
cascadearray[]Fallback provider chain when primary pool is completely exhausted.
webhookUrlstring""Target URL for interactive Telegram, Discord, Slack, or generic alerts.
providersarray[]List of { provider, keys, rpmLimit, tpmLimit, models } definitions.
providers[].modelsobject{}Per-model sub-pools: { <model>: { keys: [...], weights: [...], quotas: { <REF>: { tokenLimit } } } }. A model pool may also use credentials the provider base pool does not list.
providers[].models.<model>.quotasobject{}Local token budgets keyed by credential ref. Omit for unlimited. See Per-Model Per-Key Token Quotas.

🔌 HTTP Bridge API Reference

All management routes require loopback authentication (127.0.0.1 / ::1) with same-origin validation:

RouteMethodDescription
/dsh-key-rotation/statusGETReal-time health, keys, cooldowns. Since v0.8.0 also providers[].circuit and meta (expectedClones, notifyQueue). Model sub-pools appear as their own entries with an additive model field, and each key carries modelQuota (null = unlimited).
/dsh-key-rotation/configGET / PUTRead and update active key rotation settings and provider pools.
/dsh-key-rotation/keyPUT / DELETEAdd, update, or remove credentials in host storage and pool.
/dsh-key-rotation/resetPOSTInstantly resets all cooldowns and restores all keys to ready.
/dsh-key-rotation/test-matrixPOSTTriggers parallel health check across all configured keys and models.
/dsh-key-rotation/usage-reportGETReturns aggregated usage metrics in JSON or CSV format (?format=csv).
/dsh-key-rotation/webhook-actionPOSTReceives and executes interactive actions from Telegram/Slack callbacks (Authorization: Bearer or X-Telegram-Bot-Api-Secret-Token).

📄 License

MIT © GooDAnDReaDY

v0.7.39

  • Self-Healing & lastUsedAt Accuracy: Fixed key timestamp lookup in healIdleCooldowns to read from the modern pool.state.lastUsedAt map (with backwards-compatible fallback). credentials.resolve now properly records each key invocation timestamp in lastUsedAt, surfacing accurate "last used" indicators in the status dashboard and enabling background idle cooldown restoration.
  • Hotpath & Runtime Memoization: Eliminated redundant buildRuntime() calls across periodic sweeps, provider exhaustion handling, and /status query processing.
  • Test Suite & CI Hardening: Periodic interval sweep timers and debounce timers unref'ed to allow Node.js event loop natural exit without stalling CI runners. Purged obsolete incident test artifacts.

v0.7.38

  • Hot-Path Stream Optimization: Eliminated 4 redundant buildRuntime() calls inside the rotate() finish chunk handler by reusing the request-scoped runtime0 snapshot.
  • Allocation-Free Rate Limit Header Parsing: Optimized extractRateLimit() with single-pass header inspection and length guards, completely removing dynamic lowercase/uppercase string allocations on every response chunk.
  • Date ISO String Memoization: Memoized todayIso string calculation once per finishing request instead of creating multiple Date instances for costDays and usageDays.
  • Zero-Allocation Metrics Aggregation: Replaced intermediate array allocation [...values()].reduce() with iterative summation for totalUsage in the /dsh-key-rotation/status endpoint.
  • Stale Notification Cleanup: Automatically purge stale entries in notification throttling maps (budgetNotifiedAt, lowHealthNotifiedAt) when provider pools are deleted.

v0.7.37

  • Request Context Isolation via AsyncLocalStorage: Scoped active key resolution (pickedRef), start timestamps, and retry counts strictly to each async dispatch context using Node's node:async_hooks. Eliminates race conditions where concurrent streaming requests could penalize healthy keys.
  • Failover on Pre-Yield Stream Exceptions: Fixed fatal stream termination where transport errors (e.g. HTTP 429 thrown before headers, socket hang-up) aborted the generator. The catch block now inspects isSwitchableError and cascades seamlessly to the next key if no content tokens have been yielded.
  • Pure Local Candidate List in rotate(): Generator uses an isolated local slice of keys (attemptList), eliminating shared mutations on pool.weightedRefs.
  • Automatic Quarantine Release on Probe Success: Successful sandbox model tests via Settings UI (/dsh-key-rotation/test) automatically lift failedUntil and brokenUntil quarantine flags.
  • Long-Term Memory Compaction: Wired compactUsage(pool, 30, now) into the periodic 30-second maintenance sweep to prevent memory growth on high-uptime servers.
  • Request-Scoped Latency Recording: Replaced module-global start timestamp with context-scoped startMs for accurate p50/p95 latency metrics under concurrent load.

v0.7.36

  • Architecture de-bloat & hardening: Removed 6 unused/overengineered modules (shadow, incident, agent-budget, region, canary, maintenance) and dead route registrations.
  • High-throughput buildRuntime memoization: Eliminates per-token deep-cloning and schema validation on every streaming chunk.
  • Atomic round-robin pointer rotation: Concurrent requests advance pointer immediately on candidate selection, eliminating race conditions on simultaneous tool calls.
  • Enhanced switchable error detection: Direct parsing of HTTP status codes (429, 401, 403, 5xx) and gRPC codes (RESOURCE_EXHAUSTED, UNAVAILABLE) alongside regex fallback.
  • Informative user exhaustion messaging: Clear countdown notice with next key recovery ETA when all keys in a pool are cooling down.
  • Smart polling: Client background polling paused when browser tab is inactive (document.visibilityState).

v0.7.35

  • Lifecycle Cleanups: Wrapped credentials.resolve patch and ctx.on event handlers (llm/stream, agent/request-error) in ctx.effect scopes with guaranteed unmount cleanup (#238, #239).
  • Settings & Secret Roles: Added .role('secret') to incidentGitHubToken and webhookActionToken in Config schema for automatic UI masking (#237).
  • Settings Architecture & UI: Added native settingsScope snapshot reading/saving in settings card with graceful bridge fallback (#235).
  • Localization (Changed in v0.8.2): Card uses props.t with locale: NS; plugin registers only en; no settings.section fallback and no bundled ru/zh tables (#236, #275, #277).
  • Dead Code Purge: Removed obsolete mountDashboard routine after header-chip migration (#240).

Related plugins