跳过主要内容
Y

dsh-web-search-litellm

yunxiyang/dsh-web-search-litellm

ctx.web 搜索提供方:通过 LiteLLM 代理走 OpenAI Responses API 调用 DeepSeek 原生服务端 web_search,复用 LITELLM_API_KEY,无需新密钥。

安装

dsh plugin --profile web add github:yunxiyang/dsh-web-search-litellm

README

dsh-web-search-litellm

DSH web_search provider over the LiteLLM proxy using the OpenAI Responses protocol. The request carries the server-side web_search tool, executed natively by the DeepSeek Responses API; the grounded answer and the real URLs the model opened are returned to the harness web seam (ctx.web).

  • No Anthropic protocol — speaks POST {baseURL}/responses, not /messages.
  • No new keys — reuses the LITELLM_API_KEY credential your chat Models page already stores.
  • No third-party search service — search runs on DeepSeek's official server side, billed through your existing LiteLLM route.
  • Fully configurable in the Settings UI (web-search-litellm section).

简介 / 快速上手(中文)

这是 DeepSeek Harness ctx.web 能力的联网搜索提供方web_search 请求走 OpenAI Responses 协议发往你的 LiteLLM 代理,由 DeepSeek 官方 Responses API 在服务端原生执行搜索,返回带真实来源 URL 的答案。

  • 不需要 Anthropic 协议,也不需要新的 API Key——直接复用聊天模型页已配置的 LITELLM_API_KEY
  • 不接任何第三方搜索服务;搜索在 DeepSeek 官方服务端完成,走你现有的 LiteLLM 计费路由。
  • 安装:dsh plugin --profile <name> add dsh-web-search-litellm,然后在 profile 的 cordis.patch.yml 里把 websearchProvider 设为 litellm-responses(详见下方英文说明)。
  • 常见症状:web_searchAuthentication Fails, Your api key is invalid,且你的 DEEPSEEK_API_KEY 其实是 LiteLLM 代理 key——装这个插件并把 baseURL 指向代理即可。

何时使用 / When to use

Pick this provider when any of these is your situation:

  • web_search fails with Authentication Fails, Your api key: ****XXXX is invalid — usually because DEEPSEEK_API_KEY holds a LiteLLM proxy key, not a DeepSeek platform key.
  • All company traffic must go through LiteLLM (direct api.deepseek.com is blocked or forbidden).
  • You prefer the OpenAI Responses protocol over the Anthropic /messages format.
  • You want no free-tier / third-party search service (Tavily, Brave, Exa, …) — search stays on DeepSeek's official server side.
  • You use openai/deepseek-v4-flash or openai/deepseek-v4-pro through a LiteLLM proxy as your main model.

Install

dsh plugin --profile <name> add dsh-web-search-litellm
# or from a local checkout:
dsh plugin --profile <name> add ./dsh-web-search-litellm

Then route the seam (profile cordis.patch.yml):

- id: web
  config:
    searchProvider: litellm-responses

# optional: disable the shipped Anthropic-format DeepSeek provider
- id: web-search-deepseek
  disabled: true

Restart the profile (desktop: Settings → Desktop settings → Restart, or quit and reopen).

Configuration

Settings section web-search-litellm (harness Settings UI) or the bundle patch config:

Configuration — derive, don't hardcode

Every endpoint/model field is optional. When unset, the provider derives its values from dsh's active model configuration (the same provider the chat uses), so it works on any machine without baking in a proxy URL or model:

  • baseURL ← the active provider's baseURL (the chat's gateway).
  • apiKeyEnv ← the active provider's apiKeyEnv.
  • model ← the active model's id.
  • candidateModels ← the active provider's full models[] list, so discovery can race every model on that gateway and latch onto the first that actually runs web_search.

Only set a field here to override the derived value (e.g. to force a specific search model).

keydefaultmeaning
baseURLderived$LITELLM_SEARCH_BASE_URLhttp://127.0.0.1:4000/v1LiteLLM proxy root; /responses is appended
modelderived (active model)starting model id; the first pick
candidateModelsderived (active provider models[])fallback pool raced in parallel when the active model fails to actually run web_search; the fastest searcher wins and is cached
apiKeyEnvderivedLITELLM_API_KEYcredential reference resolved at each search
apiKeyoptional literal key (secret role)
maxTokens4096max_output_tokens for one search request
timeoutMs60000connect deadline + idle deadline for the response stream; resets whenever data arrives, so slow-but-active searches are never cut off (WEB_TIMEOUT only on real stalls)

How it works

  1. The model calls web_search with a query string.
  2. This provider POSTs to {baseURL}/responses with tools: [{"type": "web_search"}], stream: true.
  3. The LiteLLM proxy forwards the call; DeepSeek executes the search server-side and feeds results to the model.
  4. The provider parses the SSE stream: the final output_text becomes the result content, and every web_search_call item whose action is open_page contributes its URL to sources.

Session compatibility (why this plugin writes no custom session events)

This plugin appends no session events of its own. The harness reads session logs fail-closed: any event type outside the build's KNOWN_SESSION_EVENT_TYPES catalog aborts loading unless the event envelope carries ignorable: true. A third-party type can never be in that catalog, and the public session.append API offers no way to set ignorable, so a custom log-only event here would make older harness builds refuse to open any session this plugin ran in. Searches are still fully visible in the session through the standard web_search tool call/result events.

Known upstream limits (not configuration issues)

  • DeepSeek's Responses API documents include as not supported, so structured result items are consumed server-side; sources therefore carry url only (no title/snippet).
  • Each search costs one DeepSeek model turn (official mechanism).

License

MIT

相关插件