dsh-web-search-ddg
aooyoo
Zero-token DuckDuckGo search provider for the DeepSeek Harness (DSH) web seam — local headless browser, no API key, no m…
PROJECT TOPICS
PROJECT README
GitHub: bcdahb0-jpg/dsh-tool-vision
External vision model for DeepSeek Harness.
DeepSeek's own models are text-only, and the harness derives every model
request strictly from the session log (llm/stream requests must equal the
durable derivation — the agent-loop invariant). This plugin bridges the gap
in two ways:
inspect_image tool — sends an image (local file, or http(s) URL) to
any OpenAI-compatible /chat/completions endpoint that supports
image_url content parts, and returns the vision model's textual answer
into the agent loop.
Image bridge (v0.2.1) — pasted images are turned into inspect_image
hints before they enter the durable log, on the agent/pre-step
waterfall (the one seam where the harness lets a plugin replace the
messages of a proposed step). Images already logged by an older version
are repaired lazily with a surface replace on the session's first
pre-step. Only models listed in multimodalModels receive image blocks
directly; a model's declared inputModalities are never consulted,
because profiles routinely declare input: [text, image] on text-only
models just to pass the harness's prompt-admission check.
Images render in chat (v0.3.8): when the session model runs on the
bridge route (default provider tool-vision), the durable log KEEPS the
original image blocks, so the chat renders the pasted image instead of a
[User sent an image ... exported to: <path>] hint text. The bridging
logic stays fully behind the scenes: the bridge adapter rewrites image
blocks into inspect_image hints at stream time, so the text-only
upstream still receives exactly the same hint as before.
inspect_image.tool-vision namespace (API endpoint, write-only key, model, bridge
options) in settings.yaml; changes hot-apply without a restart. The API
key lives in settings.yaml, not the profile patch. Mount by package name
(name: 'dsh-tool-vision') so the web client bundle is discovered.Install from GitHub (recommended):
dsh plugin --profile <profile> add github:bcdahb0-jpg/dsh-tool-vision
Then mount it in a profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml):
- insert:
- id: tool-vision
name: 'dsh-tool-vision'
config:
baseURL: 'https://api.openai.com/v1'
apiKeyEnv: 'VISION_API_KEY'
model: 'gpt-4o-mini'
Or load it from a local path without installing the package:
- id: tool-vision
name: './plugins/dsh-tool-vision/index.js'
| Field | Default | Meaning |
|---|---|---|
baseURL |
https://api.openai.com/v1 |
OpenAI-compatible API base URL. |
apiKey |
'' |
API key (takes precedence over env). |
apiKeyEnv |
VISION_API_KEY |
Env var holding the key. |
model |
gpt-4o-mini |
Vision model id. |
maxTokens |
1024 |
Max output tokens. |
timeoutMs |
60000 |
Per-request timeout. |
maxImageBytes |
10MB |
Largest accepted local image. |
description |
default | Tool description shown to the model. |
bridgeTextOnly |
true |
Bridge pasted images to text hints on models that cannot see images. |
bridgeExportDir |
temp | Export dir for bridged images (os.tmpdir()/dsh-vision-bridge). |
multimodalModels |
[] |
Model ids that receive image blocks directly (e.g. mimo-v2.5). |
bridgeModel |
true |
Register a "bridge model entry": the model picker gains provider tool-vision with names suffixed (tool-vision 桥接); selecting one passes the harness prompt admission for pasted images — no manual input: [text, image] declaration in settings.yaml required (the admission check runs before any plugin hook, and dsh-llm-deepseek hardcodes DeepSeek models as text-only; this plugin now owns that gate). |
bridgeRoute |
tool-vision |
Provider route id of the bridge model entry (shown in the picker). |
bridgeProvider |
deepseek-official |
Text provider the bridge delegates to: text turns forward unchanged; image blocks are rewritten to inspect_image hints at request time. |
bridgeModelIds |
deepseek-v4-flash, deepseek-v4-pro |
Model ids mirrored onto the bridge route (empty = mirror all). |
When enabled, the plugin registers an image-admission route in the model picker
(default provider tool-vision, names like DeepSeek V4 Flash(tool-vision 桥接)).
Selecting it:
input: [text, image] under llm-pi-ai in settings.yaml —
now the plugin owns that gate end to end);agent/pre-step) turns images into inspect_image hints,
or keeps them as image blocks via the multimodalModels whitelist and
rewrites them at request time;bridgeProvider adapter (by
default the official DeepSeek route).Decoupling: the bridge model entry, the image bridge, the inspect_image
tool, and the settings namespace are all registered/unregistered by this plugin.
Removing the plugin removes the picker entry, the bridge, the tool, and the
settings together — nothing survives in settings.yaml; swap vision solutions by
swapping the plugin.
llm-pi-ai:
providers:
your-provider:
models:
- id: deepseek-v4-flash
input: [text, image]
- id: tool-vision
name: 'dsh-tool-vision'
config:
multimodalModels: ['mimo-v2.5', 'grok-4.5']
Then pasting an image:
inspect_image hint
([User sent an image, exported to: <path>. Inspect it with the inspect_image tool...]), and the agent inspects it through the configured
vision endpoint;Why not
llm/stream? The harness freezes every request and the agent-loop invariant fails any request whose messages diverge from the session-log derivation (log-reconstruction desync), and this cordis waterfall'snext()cannot replace request arguments. Theagent/pre-stepwaterfall is the supported seam: its decision messages become the durable log, so the invariant stays satisfied.
Key resolution order: config.apiKey → process.env[apiKeyEnv] →
process.env.OPENAI_API_KEY.
inspect_image| Arg | Required | Meaning |
|---|---|---|
path |
✅ | Image path (absolute, or relative to the current workspace) or http(s) URL. |
question |
– | Optional specific question about the image. |
detail |
– | auto / low / high resolution hint. |
Example endpoints (baseURL):
https://api.openai.com/v1 — gpt-4o, gpt-4o-minihttps://dashscope.aliyuncs.com/compatible-mode/v1 — qwen-vl-plus, qwen-vl-maxhttps://open.bigmodel.cn/api/paas/v4 — glm-4v-flash (free tier), glm-4v-plushttps://api.moonshot.cn/v1 — moonshot-v1-8k-vision-previewhttp://localhost:11434/v1 — llama3.2-vision (no key)inspect_image. On any
other text-only model the transcript stores the hint text itself.MIT
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。