返回目录
文件与数据 插件

dsh-tool-eyes

go-farther-and-farther/dsh-tool-eyes

DeepSeek Harness (DSH) 本地视觉眼睛插件:screen 工具(截图/图片交给本地视觉模型描述)+ ocr 工具(Windows 内置 OCR 逐字提取文字)。零云端、OCR 零 GPU、图片不出本机。

Stars
0
Forks
0
Issues
0
更新
1 天前

PROJECT TOPICS

项目标签

PROJECT README

README

dsh-tool-eyes

Local vision "eyes" for DeepSeek Harness (DSH).

Give your text-only agent eyes with two model-facing tools:

  • screen — capture the screen (or describe an existing image file) through a local OpenAI-compatible vision endpoint (llama.cpp with --mmproj, LM Studio, Ollama, ...) and return the vision model's text description.
  • ocr — extract ALL text verbatim with the Windows built-in OCR engine: zero model, zero GPU, zero cloud, milliseconds.
screen  =  screenshot / image  ->  local VLM  ->  text description (understanding)
ocr     =  screenshot / image  ->  Windows OCR -> verbatim text          (extraction)

Why

DeepSeek's chat-completions line is text-only. Instead of switching your whole conversation to a vision model, keep the text brain and add eyes as tools:

  • Private by default — point screen at a local endpoint and images never leave your machine.
  • Cheap — a 0.8B–4B local VLM is plenty for describing screens; ocr costs nothing at all.
  • Honest by default — the screen prompt tells the VLM to describe only what is visible and never guess app/game/character names unless confirmed by on-screen text (this measurably cuts small-model name hallucination).

Requirements

  • Windows 10/11 (the ocr tool uses WinRT OCR; screen uses .NET for capture)
  • Node.js >= 22.19, DeepSeek Harness >= 0.1.0-rc.6
  • screen additionally needs any OpenAI-compatible VLM endpoint, e.g.:
    • llama.cpp: llama-server -m model.gguf --mmproj mmproj.gguf --port 1235
    • LM Studio (loaded vision model), Ollama, or any OpenAI-compatible gateway

Install

This package is published on GitHub only (not on npm).

dsh plugin --profile web add https://github.com/go-farther-and-farther/dsh-tool-eyes

Then restart dsh web. The screen and ocr tools appear in the agent's toolkit automatically.

Manual install (offline / from source)

Copy this package into the profile's node_modules, then register it in $DSH_HOME/profiles/<profile>/cordis.patch.yml:

- insert:
    - id: tool-eyes
      name: 'dsh-tool-eyes'
      config:
        baseUrl: http://127.0.0.1:1235/v1
        model: ''
        timeoutMs: 180000

Configuration

Plugin config (all optional):

key default meaning
baseUrl http://127.0.0.1:1235/v1 OpenAI-compatible endpoint for screen
model '' model id to send; empty lets the server decide (llama.cpp serves one model)
timeoutMs 180000 hard cap for one capture call
captureScript bundled capture.ps1 override path to an alternate capture script

Override in your profile's cordis.patch.yml (id-targeted):

- id: tool-eyes
  name: 'dsh-tool-eyes'
  config:
    baseUrl: http://127.0.0.1:1235/v1
    model: qwen3.5-4b
    timeoutMs: 120000

Usage

In a conversation, the agent can now:

  • screen — "what is on my screen?", "describe this image file", with an optional prompt to focus on a region or detail.
  • ocr — "read all the text on screen", "transcribe this error dialog".

Both accept an optional image path; without it they capture the screen. The bundled PowerShell scripts can also be run standalone:

powershell -NoProfile -ExecutionPolicy Bypass -File lib\capture.ps1 -Prompt "..." -BaseUrl http://127.0.0.1:1235/v1
powershell -NoProfile -ExecutionPolicy Bypass -File lib\ocr.ps1 -Image C:\path\x.png

Privacy

  • ocr is fully local (WinRT OCR, no network).
  • screen sends the captured image to the configured baseUrl. Point it at a local endpoint (llama.cpp / LM Studio / Ollama) to keep images on your machine.

Related

  • dsh-vision-proxy — automatic transcription of attached images in the chat input (Chatbox-style), so you don't need to give file paths. Pairs well with this plugin.

Development

npm test    # node --test tests/

License

MIT

CLASSIFICATION EVIDENCE

分类依据

项目类型插件
功能分类文件与数据
规则置信度

系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。