Run every model on your own GPUs — at production economics.
Soika Stack is the inference management layer for enterprises and sovereign clouds. Deploy open and proprietary models across mixed GPU estates, route traffic intelligently, meter every token, and prove exactly what ran where.
- llama-3.3-70bH200 · node-04
- qwen-2.5-32bH100 · node-11
- soika-embed-v2L40S · node-19
TTFT
184ms
Tok/s
4.2k
Cache hit
71%
Throughput per GPU versus an unmanaged baseline deployment.
Lower cost per million tokens through batching and cache reuse.
Serving availability with multi-node failover and drain-safe upgrades.
Bytes of inference data that leave your network boundary.
What Soika Stack does
LLM inference management for private GPU fleets.
Unified model registry
Version, sign and promote open-weight, fine-tuned and third-party models through dev, staging and production gates.
Intelligent routing
Route each request by cost, latency, sensitivity or tenant policy — with automatic fallback across model families.
GPU scheduling & autoscaling
Bin-pack workloads across H100, H200, GB200 and L40S nodes, with pre-emption, priority classes and scale-to-zero.
Serving optimisation
Continuous batching, paged KV-cache, speculative decoding and quantisation pipelines tuned per model.
Full-fidelity observability
Token-level telemetry, per-tenant cost attribution, latency histograms and GPU utilisation in one console.
Air-gapped ready
Install from a signed offline bundle. No telemetry callbacks, no external licence checks, no internet dependency.
One gateway for every model your organisation uses
Applications talk to a single OpenAI-compatible endpoint. Behind it, Soika Stack decides which engine, which node and which precision serves the request — and rewrites that decision as your estate changes.
- OpenAI-compatible REST and gRPC APIs, plus native streaming
- vLLM, TensorRT-LLM and SGLang runtimes managed side by side
- Blue/green and canary rollouts with automatic rollback on eval regression
- Per-tenant rate limits, quotas and burst budgets
- llama-3.3-70bH200 · node-04
- qwen-2.5-32bH100 · node-11
- soika-embed-v2L40S · node-19
TTFT
184ms
Tok/s
4.2k
Cache hit
71%
Know the cost of every token before finance asks
Inference spend is the line item nobody can explain. Soika Stack attributes every token to a tenant, application, agent and business unit, then shows where the money actually goes.
- Chargeback and showback reporting per department
- Cache-hit analytics and prompt-cost regression alerts
- Capacity forecasting against committed GPU inventory
- Budget guardrails that throttle instead of surprising you
- Customer operations$18,400
412M tokens
- Risk & compliance$11,950
268M tokens
- Engineering$6,320
147M tokens
- HR shared services$2,480
58M tokens
Governed by design, not by policy document
Every request is authenticated, classified, logged and retained according to your rules. Regulators get evidence, not assurances.
- SSO, SCIM and fine-grained RBAC down to the model version
- Prompt and response retention policies with configurable redaction
- Immutable audit log of who invoked which model, when and why
- Data-residency enforcement at the routing layer
- 14:02:11invoke · llama-3.3-70bsvc/claims-agent
- 14:02:11redact · 3 spanspolicy/pii
- 14:02:12retrieve · policy_v7svc/claims-agent
- 14:02:19approve · payout £4,120user/a.moreau
- 14:02:19seal · block 84,213audit/ledger
Specifications
- Deployment
- Kubernetes, bare metal, air-gapped, sovereign cloud
- Accelerators
- NVIDIA H100 · H200 · GB200 · L40S · RTX 6000 Ada
- Runtimes
- vLLM · TensorRT-LLM · SGLang · custom containers
- Model formats
- Safetensors, GGUF, AWQ, FP8, INT4
- Interfaces
- OpenAI-compatible REST, gRPC, WebSocket streaming
- Identity
- OIDC / SAML SSO, SCIM provisioning, mTLS service auth
Where it earns its place
National AI cloud
Serve multiple ministries from one GPU estate with hard tenant isolation and per-agency reporting.
Bank-wide model service
Give every internal team a governed endpoint instead of shadow API keys and unreviewed vendors.
GPU service provider
Turn a data-centre investment into a metered, sellable inference product with billing-grade telemetry.
The other two products
Every Soika product is bought, deployed and run on its own. Soika Stack is the exception in one direction only: it can serve the models behind either platform.
See Soika Stack on your own data
We run a scoped discovery workshop, deploy into your environment and prove the outcome against your real workload — with your local Soika partner alongside.