Skip to content
Inference control plane

Run every model on your own GPUs — at production economics.

Soika Stack is the inference management layer for enterprises and sovereign clouds. Deploy open and proprietary models across mixed GPU estates, route traffic intelligently, meter every token, and prove exactly what ran where.

3.4×
Throughput per GPU versus an unmanaged baseline deployment.
68%
Lower cost per million tokens through batching and cache reuse.
99.95%
Serving availability with multi-node failover and drain-safe upgrades.
0
Bytes of inference data that leave your network boundary.
Capabilities

What Soika Stack does

The control plane that turns raw GPUs into a governed, metered, production-grade model service — on your hardware, in your jurisdiction.

Unified model registry

Version, sign and promote open-weight, fine-tuned and third-party models through dev, staging and production gates.

Intelligent routing

Route each request by cost, latency, sensitivity or tenant policy — with automatic fallback across model families.

GPU scheduling & autoscaling

Bin-pack workloads across B300, B200, H200, H100 and L40S nodes, with pre-emption, priority classes and scale-to-zero.

Serving optimisation

Continuous batching, paged KV-cache, speculative decoding and quantisation pipelines tuned per model.

Full-fidelity observability

Token-level telemetry, per-tenant cost attribution, latency histograms and GPU utilisation in one console.

Air-gapped ready

Install from a signed offline bundle. No telemetry callbacks, no external licence checks, no internet dependency.

Serving

One gateway for every model your organisation uses

Applications talk to a single OpenAI-compatible endpoint. Behind it, Soika Stack decides which engine, which node and which precision serves the request — and rewrites that decision as your estate changes.

  • OpenAI-compatible REST and gRPC APIs, plus native streaming
  • vLLM, TensorRT-LLM and SGLang runtimes managed side by side
  • Blue/green and canary rollouts with automatic rollback on eval regression
  • Per-tenant rate limits, quotas and burst budgets

Economics

Know the cost of every token before finance asks

Inference spend is the line item nobody can explain. Soika Stack attributes every token to a tenant, application, agent and business unit, then shows where the money actually goes.

  • Chargeback and showback reporting per department
  • Cache-hit analytics and prompt-cost regression alerts
  • Capacity forecasting against committed GPU inventory
  • Budget guardrails that throttle instead of surprising you

Control

Governed by design, not by policy document

Every request is authenticated, classified, logged and retained according to your rules. Regulators get evidence, not assurances.

  • SSO, SCIM and fine-grained RBAC down to the model version
  • Prompt and response retention policies with configurable redaction
  • Immutable audit log of who invoked which model, when and why
  • Data-residency enforcement at the routing layer
Where it fits

Common ways teams deploy Soika Stack

National AI cloud

Serve multiple ministries from one GPU estate with hard tenant isolation and per-agency reporting.

Bank-wide model service

Give every internal team a governed endpoint instead of shadow API keys and unreviewed vendors.

GPU service provider

Turn a data-centre investment into a metered, sellable inference product with billing-grade telemetry.

Specification

Technical footprint

Deployment
Kubernetes, bare metal, air-gapped, sovereign cloud
Accelerators
NVIDIA B300 · B200 · H200 · H100 · L40S · RTX PRO · RTX 6000 Ada
Runtimes
vLLM · TensorRT-LLM · SGLang · custom containers
Model formats
Safetensors, GGUF, AWQ, FP8, INT4
Interfaces
OpenAI-compatible REST, gRPC, WebSocket streaming
Identity
OIDC / SAML SSO, SCIM provisioning, mTLS service auth
Get started

See Soika Stack against your own workload

Bring a real process or dataset. We will show you the deployment, not a slide deck.