GPU control plane
Sparkplane provisions vLLM on Kubernetes across your GPU cluster, hands you an OpenAI-compatible endpoint, and keeps every token on hardware you control.
Serving stack
vLLM
Orchestration
Kubernetes
API surface
OpenAI
Data path
Private
01 — Flight sequence
Llama, Qwen, Mistral, Gemma and DeepSeek weights, sized against your available VRAM.
A vLLM pod and service land on a GPU node with tensor parallelism set for you.
An OpenAI-compatible endpoint plus scoped keys, ready the moment the pod reports healthy.
A built-in context window with system prompt, temperature and token controls.
02 — In the catalog
Cluster credentials are encrypted at rest and never leave the backend. Your models run on private infrastructure reached over an internal network path — the platform itself is public, the GPUs are not.