dsh-plugin-verified-search
f0909172434
Verified current-source search workflow for DeepSeek Harness
akqwpeter-prog/dsh-media-skills
Free vision & image generation for DeepSeek Harness — paste an image into any chat, even text-only sessions. GLM-4V-Flash / Qwen3-VL / Gemini failover chain, ModLens-style structured evidence, Kolors generation. 免费读图·生图 · 三引擎容错 · 无 Key 入库
PROJECT TOPICS
PROJECT README
DeepSeek Harness is brilliant at reasoning — but a text-only model can't see the image you just dragged into the chat. This bundle fixes that with two free skills, a free vision model route, and a vision engine failover chain:
vision-review — analyze images and screenshots, catch UI visual bugs, detect watermarks, turn images into text.media-tools — generate illustrations, avatars, backgrounds and banners with a free, watermark-free model.No hardcoded keys, no paid API, no file saving, no session switching.
Why · Quick start · See it in action · Usage · Keys & privacy · FAQ · Examples
English · 简体中文 · 繁體中文 · 日本語 · 한국어 · Español · Deutsch · Português · Русский
Most DSH vision plugins only read images — and many push you through a shared third-party endpoint. dsh-media-skills takes a different stance:
| This bundle | Typical vision-only plugin | |
|---|---|---|
| Read images for free | ✅ Zhipu GLM-4V-Flash | ✅ |
| Generate images for free | ✅ SiliconFlow Kolors | ❌ usually absent |
| Auto model route in the picker | ✅ installed automatically | sometimes |
| Keys committed to the repo | ❌ never — keys stay local | ⚠️ often required |
| Docs in multiple languages | ✅ 9 languages | ❌ usually English only |
| Privacy | ✅ you choose the provider; images only go to your provider | shared free endpoints can see your images |
Why bring your own free key instead of a built-in anonymous endpoint? Privacy and reliability. Your images go only to the provider you choose, under your account and your rate limits — no shared third-party service in the middle.
| Capability | What it does | Model | Cost |
|---|---|---|---|
| 📎 Paste-image reading | In a text-only session, the input bar gains an “Add image” button (paperclip); pasted images are auto-described by the vision model and handed to the current model as text. (Harness-core feature: requires the core api-proxy admission patch; this bundle supplies the vision route + skill it depends on) | Zhipu GLM-4V-Flash | Free |
| 🧠 Vision model route | 「智谱 GLM-4V-Flash(视觉)」 appears in the model selector automatically — pick it for a new conversation and talk about images directly | Zhipu GLM-4V-Flash | Free |
👁️ vision-review |
Analyze / recognize / describe images & screenshots; catch UI visual bugs (overlap, overflow, misalignment); detect watermarks/logos; turn images into text. Optional --structured mode returns ModLens-style evidence JSON (summary, full OCR, reading-order layout, entities/relations, uncertainty). Engine failover chain: GLM-4V-Flash → SiliconFlow Qwen3-VL / Google Gemini (auto-join with free keys) → any OpenAI-compatible endpoint |
GLM-4V-Flash + Qwen3-VL + Gemini | Free |
🎨 media-tools |
Generate images, illustrations, avatars, backgrounds, banners | SiliconFlow Kolors | Free, no watermark |
dsh plugin --profile <name> add github:akqwpeter-prog/dsh-media-skills
Get two free keys (~2 minutes, no payment):
glm-4v-flash is free)Add them in the Web GUI (Settings → Models → the zhipu-vision provider's API Key field), or use the credentials file:
# ~/.dsh/.credentials.yaml (chmod 600)
GLM_API_KEY: <your key>
Restart dsh web, then hard-refresh (Cmd+Shift+R).
Verify: the model selector shows 智谱 GLM-4V-Flash(视觉). If your Harness build supports paste-image reading, the input bar also has a 📎 Add image button — paste an image in any session and it arrives as a text description.
Full walkthrough and troubleshooting: docs/SETUP_VISION_EN.md.
Paste an image in a text-only session → the free vision model describes it → your model answers. The same bundle also generates new images on demand.
How it works in one picture:
Three ways to read images:
| Way | How | When |
|---|---|---|
| A. Paste directly (recommended) | In any session, click the 📎 button / drag / paste an image and send | Everyday image questions — no file saving, no model switching |
| B. Vision model session | New conversation, pick 智谱 GLM-4V-Flash(视觉), paste images and chat | Multi-turn image conversations, native read_image |
| C. Files + skill | Put the image in the workspace and say “read this image with vision-review” | Batch review, scripted workflows |
Descriptions follow your message language (Chinese message → Chinese description; English message → English description; no text → Chinese).
Also just say:
vision-reviewmedia-toolsKeys are never stored in this repo. Skill scripts read, in order: environment variables → ~/.dsh/secrets/media-tools.env → ~/.codex/secrets/media-tools.env (legacy fallback). The vision model route reads GLM_API_KEY from DSH's credential store.
Where to get the keys (all free): Zhipu — open.bigmodel.cn → API Keys (glm-4v-flash). SiliconFlow — siliconflow.cn → API Keys (Kolors). Google (optional, joins the vision failover chain automatically) — aistudio.google.com → Get API key.
# ~/.dsh/secrets/media-tools.env (chmod 600, one KEY=value per line)
GLM_API_KEY=...
SILICONFLOW_API_KEY=...
GEMINI_API_KEY=... # optional
Your images are sent only to the provider you configure — never to this repo, never to a shared anonymous endpoint.
Privacy note on Gemini: Google's free-tier key comes with data-use terms — requests may be used to improve Google products. For sensitive images (IDs, internal docs, customer data), prefer the direct domestic engines (Zhipu / SiliconFlow).
Does paste-image reading require a DeepSeek Harness core patch?
The auto-describe pipeline lives in the Harness core (api-proxy image-admission logic; see docs/HARNESS_PATCH_EN.md). This bundle ships the model route + skills: the vision model works on any DSH build, but paste-image reading requires a Harness build with that core support — see FAQ Q1 in docs/SETUP_VISION_EN.md.
Why not just use a built-in free endpoint with no key at all? We prefer to let you own the route: your images go to the provider you pick, under your rate limits, with no shared middleman. The keys are free and take about two minutes to create.
Is media-tools really free?
Yes — SiliconFlow Kolors is free and watermark-free. If a model is temporarily disabled, the skill lists available models and you can switch.
Sample material to try instantly — 6 AI-generated images with their prompts, plus a purpose-built vision test card (title, buttons, bar-chart values) for checking reading accuracy:
dsh-media-skills/
├── package.json # dsh.bundle manifest
├── cordis.patch.yml # plugin layer
├── index.js # registers skills + seeds the zhipu-vision model route
├── skills/
│ ├── vision-review/ # image reading
│ └── media-tools/ # image generation
├── examples/ # sample images + vision test card
├── docs/
│ ├── screenshots/ # demo mockup & how-it-works diagram
│ ├── SETUP_VISION_EN.md # detailed setup guide (English)
│ ├── SETUP_VISION.md # 详细配置指南(中文)
│ ├── HARNESS_PATCH_EN.md# core patch notes (English)
│ ├── HARNESS_PATCH.md # 本体补丁说明(中文)
│ ├── COMPARE_MODLENS.md # 与 ModLens 的对比/共存(中文)
│ └── lang/ # READMEs in 9 languages
├── scripts/make-banner.py # regenerates docs/social-preview.png
└── docs/social-preview.png
Both this bundle and ModLens give text-only models vision. Installed together they do not conflict: ModLens intercepts pastes first (path → modlens_read_image tool), and this bundle's api-proxy fallback handles anything it doesn't take over. See docs/COMPARE_MODLENS.md (中文) for the full comparison, the paste routing order, and how to point ModLens at the same free Zhipu endpoint.
DeepSeek Harness developer preview is still in its testing phase for Harness developers; core plugins and base APIs will keep iterating. We look forward to exploring the upper limits of intelligence together with developers worldwide, on top of open-source, open, reusable, and composable infrastructure.
This repo is tagged
dsh-pluginand listed in the awesome-dsh-plugin curated list. PRs, issues and translations are welcome.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。