dsh-thinking-language-zh
ayanJava111
deepseek harness思考过程中文插件
PROJECT TOPICS
PROJECT README
中文(README.zh.md,推荐) | English
Semantic long-term memory for DeepSeek Harness.
A dsh-plugin (Cordis plugin) that gives the model a persistent, embedding-based
memory across sessions — unlike the harness's built-in session_query (literal
FTS5), this store retrieves by meaning.
Zero-config out of the box: install, restart, and start a new session. The local embedding model downloads itself on first use (~100 MB); the four tools, per-question recall, and 5-turn auto-summarization all work with defaults. You only configure when you want something different (see Configuration).
$DSH_HOME/memories/memories.jsonl.local (default): ONNX inference via @huggingface/transformers with
Xenova/bge-small-zh-v1.5 (offline, ~100 MB model, cached in
~/.cache/huggingface).api: any OpenAI-compatible /embeddings endpoint (e.g. SiliconFlow,
Zhipu, DashScope).memory_write — persist a fact / decision / preference / note (content
hash dedup, repeats update in place).memory_search — semantic top-k recall with kind/tag/workspace filters.memory_forget — delete by id.memory_stats — store summary.memory_write on its own when the user states a durable preference, an
established fact, or an explicit decision (no need to say "remember").auto). One in-flight summary per session; silent on
failure; only active when llm and agentDefaultModel services exist.The package ships an in-package cordis.patch.yml declared via
dsh.bundle.patch, so the plugin command mounts it automatically with no
manual profile edits. DSH delegates plugin installation to pnpm, so pnpm
must be available on PATH regardless of how DSH itself is launched.
Use the command matching your DSH launcher:
# DSH run through npx (no global `dsh` command)
npx @deepseek-ai/dsh plugin --profile web add dsh-plugin-semantic-memory
# DSH run from a deepseek-harness source checkout (run from that repo root)
pnpm dsh plugin --profile web add dsh-plugin-semantic-memory
# DSH installed with a global `dsh` command
dsh plugin --profile web add dsh-plugin-semantic-memory
For a local checkout, build it first, then use the same launcher prefix:
cd C:/path/to/dsh-semantic-memory
npm install
npm run build
npx @deepseek-ai/dsh plugin --profile web add file:C:/path/to/dsh-semantic-memory
Then restart the Web profile with the same launcher
(npx @deepseek-ai/dsh web, pnpm dsh web, or dsh web) and start a new
session. All knobs have schema defaults; the in-package cordis.patch.yml is
the deployment config source — in-package config overrides outer layers
(settings.yaml and user patch rows only fill keys the package does not declare,
they do not override it).
Applying local configuration changes: the Web profile uses a copied snapshot for a
file:dependency rather than reading the checkout live. After changingcordis.patch.ymlor rebuilding the plugin, delete<DSH_HOME>\profiles\web\node_modules\dsh-plugin-semantic-memory, run<your DSH launcher> plugin --profile web install, and restart the Web profile.
Manual equivalent (for older installs): add the dependency to the profile's
package.json, insert a mount row — new entries must be inserted (a bare
- id: row only overrides an existing bundle id and is silently ignored):
- insert:
- id: semantic-memory
name: 'dsh-plugin-semantic-memory'
Leave mode/provider unset unless you need an explicit switch: selection is
automatic (see below).
The embedding provider is chosen by mode (explicit deployment switch), falling
back to the automatic selection:
| Configuration | Provider |
|---|---|
mode: 'cloud' |
API (OpenAI-compatible /embeddings endpoint); requires apiKey |
mode: 'local' |
local (ONNX via @huggingface/transformers, offline), even with an apiKey set |
no mode, apiKey present (non-empty) |
API |
no mode, no apiKey |
local |
provider: 'local' (explicit) |
local, even with an apiKey set |
provider: 'api' (explicit) |
API; requires apiKey |
Switching deployment mode means editing mode in the in-package
cordis.patch.yml and restarting the Web profile with the same DSH launcher —
in-package config overrides outer layers (settings.yaml
or user profile patch rows only fill keys the package does not declare; they do
not override it). A restart is needed after patch-file changes; the settings
document (~/.dsh/settings.yaml, semantic-memory: section) hot-reloads for
the keys it is allowed to supply. The first local embed downloads the model
(~100 MB, cached in ~/.cache/huggingface; use remoteHost for a mirror).
Open a new session (existing sessions keep their original tool set) and ask
the model: "Do you have memory_ tools?"* — it should list memory_write,
memory_search, memory_forget, and memory_stats. The system prompt also
carries a ## Long-term memory section once memories exist.
memory_write without being asked.memory_write.memory_search digs deeper
(supports kind, tags, workspace, limit, min_score).memory_forget <id> deletes; memory_stats summarizes the store.| Trigger | Behavior |
|---|---|
| Every user message | Asynchronous embedding + search; the freshest per-session hits are injected into the next prompt assembly (## Long-term memory (recalled for your current question)) |
| Every N user messages (default 5) | The harness LLM distills only the messages since the last summary (per-session seq cursor — no re-digesting, nothing skipped) into memory entries, written with the auto tag; the cadence can be set with the DSH_SEMANTIC_MEMORY_SUMMARIZE_EVERY environment variable (0 disables, overrides the config document) |
| Prompt assembly, no fresh recall | Strongest memories (importance × recency × access) injected as fallback |
$DSH_HOME/memories/memories.jsonl (one JSON line per entry, vectors
included; edit/backup freely).cordis.patch.yml (deployment source of truth — package
config overrides outer layers); ~/.dsh/settings.yaml under semantic-memory:
only fills keys the package does not declare (hot-reloaded).remoteHost: https://hf-mirror.com in restricted networks.api provider errors — confirm mode/apiKey are set and apiBase
points at an OpenAI-compatible endpoint (a /v1 base gets /embeddings
appended).llm and agentDefaultModel
services (present in the standard web profile) and autoSummarizeEvery > 0.| Key | Default | Meaning |
|---|---|---|
mode |
(unset) | Deployment switch: local forces the local model, cloud forces the API (requires apiKey). Unset keeps the automatic selection. |
provider |
auto |
auto selects by apiKey (non-empty → api, else local); explicit local/api overrides. An explicit mode overrides both. |
localModel |
Xenova/bge-small-zh-v1.5 |
Local transformer model id. |
remoteHost |
https://huggingface.co |
Model download host; set https://hf-mirror.com in restricted networks. |
apiBase |
https://api.siliconflow.cn/v1 |
API base URL (an /embeddings route is appended). |
apiKey |
'' |
API key. When non-empty and provider is not explicitly local, the API provider is used. |
apiModel |
BAAI/bge-m3 |
API embedding model name. |
memoryPath |
$DSH_HOME/memories/memories.jsonl |
Store file path. |
promptTopK |
3 |
Memories injected per system-prompt assembly (0 disables). |
maxSearchResults |
10 |
Default memory_search hit cap. |
minScore |
0.35 |
Default minimum relevance for search hits. |
halfLifeMs |
30 days | Memory strength half-life. |
autoSummarizeEvery |
5 |
Auto-summarize every N user messages (0 disables; needs llm + agentDefaultModel). The DSH_SEMANTIC_MEMORY_SUMMARIZE_EVERY env var overrides this (0..100). |
summarizeWindow |
12 |
Most recent messages included in one auto-summary. |
summarizeMaxTokens |
800 |
Token budget for the summary call. |
summarizeTemperature |
0.2 |
Sampling temperature for the summary call. |
interface MemoryEntry {
id: string // sha1(kind + content), 16 hex chars — upsert key
kind: 'fact' | 'decision' | 'preference' | 'note'
content: string // one-sentence, self-contained text
tags: string[]
workspace?: string // caller session cwd at write time
source?: { sessionId: string; seq: number }
importance: number // 1..5
embedding: number[] // normalized vector
createdAt: number
updatedAt: number
accessCount: number
lastAccessAt: number
}
Effective strength = importance / 5 × 0.5^(age / halfLife);
search rank = cosine(query, entry) × strength.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。