Bounded GGUF parsing
Architecture, context, tensor count, parameter hints and quantization metadata are read locally without loading the model for inference.
02 · Flagship case study · Source available · 1.0 · Local AI systems
A local-first workstation manager for the parts of local AI that become difficult to trust once a setup grows: model files, multimodal projectors, launch flags, runtime ownership, RAM/VRAM pressure and benchmark history.
The problem
My local AI setup worked, but it was increasingly difficult to reason about. Models lived across drives. Commands depended on exact flags. Vision required the right projector. Ports could already be occupied. A benchmark could look fast even when the machine was already under load.
VektorDeck turns that into one persistent local control surface. It discovers assets, inspects GGUF metadata, validates saved profiles, launches runtimes safely, watches the machine, records evidence and explains why a recommendation exists.
The project deliberately avoids becoming another chat frontend. Its job is to make the workstation itself understandable and trustworthy.
Architecture
The browser never launches arbitrary commands or scans drives directly. System-specific behavior stays behind a loopback FastAPI service and explicit runtime adapters.
VektorDeck manages llama.cpp, AUTOMATIC1111 and Hermes locally. It does not require a cloud account, and model/runtime control stays on the workstation.
Model intelligence + Smart Launch
GGUF inspection and launch recommendations are deliberately linked but kept conceptually separate. Metadata describes what is installed; Smart Launch explains what the current machine should try.
Architecture, context, tensor count, parameter hints and quantization metadata are read locally without loading the model for inference.
Vision projectors are suggested from folder/name evidence, shown with reasons, and never silently applied.
Saved configurations are checked against indexed assets, local-only host rules and GGUF-advertised context limits before launch.
The current 32 GB-class RAM / 16 GB VRAM workstation maps to an 8K starting point for this 11.15 GB multimodal model, with every reduction rule visible.
Real workstation evidence
The application preserves old results but separates legacy history from controlled evidence. If the machine is already busy, Protocol v2 refuses the run instead of creating a fake clean benchmark.
Safety rules
Do not replace it. Identify ownership where possible and fail closed when the process is not VektorDeck-managed.
REFUSEDo not launch. Model and projector paths must resolve to assets the local inventory knows about.
BLOCKRefuse the controlled run and surface likely contaminating processes instead of manufacturing clean evidence.
DEFERKeep the recommendation experimental. Promotion requires repeated Protocol-v2 HEALTHY/TIGHT evidence for the same profile.
NO PROMOTIONFailure → redesign
Inference-time telemetry could show pressure, but it could not prove the machine was quiet before the run began.
Protocol v2 records a pre-inference baseline, source, contaminators and inference-time evidence. Busy machines are allowed to say “not now.”
The file on disk was correct while the current update process was still running stale logic.
The updater hashes itself before and after pull and automatically restarts once when its own script changes.
PIDs are reusable, and blindly adopting a process would create a dangerous ownership assumption.
Runtime recovery verifies the saved profile, PID, executable path, port, model path and live llama-compatible API before reclaiming ownership.
Engineering surface
Local APIs, runtime adapters, process ownership, telemetry, GGUF parsing, validation and orchestration.
Stateful control surfaces for runtimes, profiles, telemetry, evidence, Model Intelligence and release readiness.
Persistent model inventory, launch profiles, benchmark history, runtime leases and Protocol-v2 provenance.
Ports, process trees, Windows paths, PowerShell lifecycle issues, loopback networking and fail-closed ownership checks.
Explicit HEALTHY/TIGHT/PRESSURED/UNVERIFIED states, contamination checks and transparent recommendation reasons.
GitHub Actions, local regression/build gates, repair scripts, diagnostics and live post-relaunch smoke tests.
Built a local-first Windows AI workstation manager using Python/FastAPI, React/TypeScript and SQLite, with safe runtime ownership, GGUF metadata inspection, hardware telemetry, evidence-aware benchmarking, transparent launch recommendations, CI and live smoke validation.
VektorDeck 1.0
VektorDeck 1.0 deliberately stops at a complete Windows local-workstation product. Installer packaging, ComfyUI, remote control and plugins are future extensions rather than unfinished core requirements.