Skip to content
Inference control planeShared foundation · Serves models to any application you run

Run every model on your own GPUs — at production economics.

Soika Stack is the inference management layer for enterprises and sovereign clouds. Deploy open and proprietary models across mixed GPU estates, route traffic intelligently, meter every token, and prove exactly what ran where.

Inference gatewayLive
  • llama-3.3-70bH200 · node-04
  • qwen-2.5-32bH100 · node-11
  • soika-embed-v2L40S · node-19

TTFT

184ms

Tok/s

4.2k

Cache hit

71%

3.4×

Throughput per GPU versus an unmanaged baseline deployment.

68%

Lower cost per million tokens through batching and cache reuse.

99.95%

Serving availability with multi-node failover and drain-safe upgrades.

0

Bytes of inference data that leave your network boundary.

Capabilities

What Soika Stack does

LLM inference management for private GPU fleets.

Unified model registry

Version, sign and promote open-weight, fine-tuned and third-party models through dev, staging and production gates.

Intelligent routing

Route each request by cost, latency, sensitivity or tenant policy — with automatic fallback across model families.

GPU scheduling & autoscaling

Bin-pack workloads across H100, H200, GB200 and L40S nodes, with pre-emption, priority classes and scale-to-zero.

Serving optimisation

Continuous batching, paged KV-cache, speculative decoding and quantisation pipelines tuned per model.

Full-fidelity observability

Token-level telemetry, per-tenant cost attribution, latency histograms and GPU utilisation in one console.

Air-gapped ready

Install from a signed offline bundle. No telemetry callbacks, no external licence checks, no internet dependency.

Serving

One gateway for every model your organisation uses

Applications talk to a single OpenAI-compatible endpoint. Behind it, Soika Stack decides which engine, which node and which precision serves the request — and rewrites that decision as your estate changes.

  • OpenAI-compatible REST and gRPC APIs, plus native streaming
  • vLLM, TensorRT-LLM and SGLang runtimes managed side by side
  • Blue/green and canary rollouts with automatic rollback on eval regression
  • Per-tenant rate limits, quotas and burst budgets
Inference gatewayLive
  • llama-3.3-70bH200 · node-04
  • qwen-2.5-32bH100 · node-11
  • soika-embed-v2L40S · node-19

TTFT

184ms

Tok/s

4.2k

Cache hit

71%

Economics

Know the cost of every token before finance asks

Inference spend is the line item nobody can explain. Soika Stack attributes every token to a tenant, application, agent and business unit, then shows where the money actually goes.

  • Chargeback and showback reporting per department
  • Cache-hit analytics and prompt-cost regression alerts
  • Capacity forecasting against committed GPU inventory
  • Budget guardrails that throttle instead of surprising you
Cost attributionThis month
  • Customer operations$18,400

    412M tokens

  • Risk & compliance$11,950

    268M tokens

  • Engineering$6,320

    147M tokens

  • HR shared services$2,480

    58M tokens

Budget guardrail82% of quarter · throttle at 95%
Control

Governed by design, not by policy document

Every request is authenticated, classified, logged and retained according to your rules. Regulators get evidence, not assurances.

  • SSO, SCIM and fine-grained RBAC down to the model version
  • Prompt and response retention policies with configurable redaction
  • Immutable audit log of who invoked which model, when and why
  • Data-residency enforcement at the routing layer
Audit ledgerImmutable
  • 14:02:11invoke · llama-3.3-70b
  • 14:02:11redact · 3 spans
  • 14:02:12retrieve · policy_v7
  • 14:02:19approve · payout £4,120
  • 14:02:19seal · block 84,213
Retention7 years · residency eu-west
Technical profile

Specifications

Deployment
Kubernetes, bare metal, air-gapped, sovereign cloud
Accelerators
NVIDIA H100 · H200 · GB200 · L40S · RTX 6000 Ada
Runtimes
vLLM · TensorRT-LLM · SGLang · custom containers
Model formats
Safetensors, GGUF, AWQ, FP8, INT4
Interfaces
OpenAI-compatible REST, gRPC, WebSocket streaming
Identity
OIDC / SAML SSO, SCIM provisioning, mTLS service auth
In production

Where it earns its place

National AI cloud

Serve multiple ministries from one GPU estate with hard tenant isolation and per-agency reporting.

Bank-wide model service

Give every internal team a governed endpoint instead of shadow API keys and unreviewed vendors.

GPU service provider

Turn a data-centre investment into a metered, sellable inference product with billing-grade telemetry.

Get started

See Soika Stack on your own data

We run a scoped discovery workshop, deploy into your environment and prove the outcome against your real workload — with your local Soika partner alongside.