dsh-plugin-verified-search
f0909172434
Verified current-source search workflow for DeepSeek Harness
PROJECT TOPICS
PROJECT README
Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.
Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.
glance -- describe, ask about, or OCR an imageground -- locate a specific element (returns bounding box)detect -- find all instances of an element kindcrop -- cut a region from an imagedsh plugin --profile your-profile add dsh-plugin-vision-toolkit
Set environment variables:
export VISION_API_KEY=sk-xxx # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://... # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2 # Vision model name
# Describe an image
glance screenshot.png
# Ask a question
glance screenshot.png -q "What error is shown?"
# OCR
glance screenshot.png --ocr
# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870
# Find all buttons
detect screenshot.png "buttons"
# Crop a region
crop screenshot.png 450,820,620,870 button.png
The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:
The agent never sees raw pixels -- it gets text descriptions it can reason about.
MIT -- YYTbit
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。