Skillsdsh-plugin
Stars
0
Forks
0
Open issues
0
Last push
Aug 22, 2026
Latest release
—
24h Growth
+0
7d Growth
+0
Installation
No reliable install command was detected. Check the project README for installation instructions.
Overview
Vision translation plugin for DeepSeek Harness: converts images into structured <vision-context> primitives via an auxiliary VLM bridge, letting text-only agents consume visual information.
- Converts images into structured `<vision-context>` text primitives for text-only agents.
- Uses an auxiliary VLM to inspect the image while the Core generates norm-1000 xyxy visual primitives.
- Computes spatial relations (left_of/right_of/above/below/inside/overlaps) from geometry rather than language.
- Extracts OCR visible text into the output context.
- Fail-closed behavior: returns unavailable after three invalid VLM outputs and never fabricates context.
- Any OpenAI-compatible /chat/completions endpoint works; defaults are OpenRouter-compatible.
- Provides native Hermes Bridge, generic MCP Bridge, and CLI cross-language entry point.
- Core runtime uses only Python standard library; optional Pillow for rotation/resize.