Skip to content
AI hardware

Systems that arrive already running Soika

Every Soika workstation, laptop and cluster ships with the Soika Enterprise licence, Soika Stack for GPU and model management, and Mockingjay Network for no-code agent design — commissioned, tuned and supported. No hardware setup headache before the first agent runs.

Soika Stack

GPU clustering, LLM deployment and inference management, pre-configured and licensed.

Mockingjay Network

No-code AI agent design and agent-fleet management with the Enterprise licence.

Ubuntu LTS

Hardened Linux base image with the full CUDA and container toolchain already in place.

AI workstations

Desktop systems for the heaviest workloads

Eight configurations from a two-GPU Blackwell system to an H200 platform with 141 GB of HBM3e per accelerator — all on the same Xeon platform, all with the same licence and the same three-year service.

SM RTX PRO 4000

2× NVIDIA RTX PRO 4000 Blackwell

The smallest system that still runs the full Soika software stack — for teams putting their first no-code agents into production.

GPU memory
24 GB
Bandwidth
672 GB/s
CUDA cores
8,960
Interface
PCIe Gen 5
Datasheet

SM RTX PRO 4500

2× NVIDIA RTX PRO 4500 Blackwell

Twice the CPU platform of the 4000 for teams running retrieval, tooling and inference side by side on one machine.

GPU memory
32 GB
Bandwidth
896 GB/s
CPU cores
120
Storage
16 TB
Datasheet

SM RTX PRO 5000

2× NVIDIA RTX PRO 5000 Blackwell

The volume choice for departmental deployments — enough memory to serve a capable model and its retrieval index without quantising down.

GPU memory
48 GB
Bandwidth
1,344 GB/s
CUDA cores
14,080
Interface
PCIe Gen 5
Datasheet

SM RTX PRO 6000

2× NVIDIA RTX PRO 6000 Blackwell

Maximum single-node capability short of a data-centre platform, for teams running frontier-class open models on their own floor.

GPU memory
up to 96 GB
Bandwidth
1,792 GB/s
CUDA cores
24,064
Interface
512-bit
Datasheet

SM 5000

3× NVIDIA RTX 5000 Ada Generation

A proven three-GPU configuration for mixed inference, rendering and simulation work under one Soika Enterprise licence.

GPU memory
32 GB
Bandwidth
576 GB/s
Tensor
1,044.4 TFLOPS
CUDA cores
12,800
Datasheet

SM 5880

3–4× NVIDIA RTX 5880 Ada Generation

More memory per GPU and a fourth slot — for teams that outgrew 32 GB cards but do not need a data-centre platform.

GPU memory
48 GB
Bandwidth
960 GB/s
Tensor
1,108.4 TFLOPS
GPUs
up to 4
Datasheet

SM 6000

3–4× NVIDIA RTX 6000 Ada Generation

Four RTX 6000 Ada GPUs, 18,176 CUDA cores each, with the Agent-to-Agent network option enabled.

GPU memory
48 GB
Bandwidth
960 GB/s
Tensor
1,457 TFLOPS
CUDA cores
18,176
Datasheet

SM H200

2–4× NVIDIA H200 SXM

141 GB of HBM3e per GPU and 4.8 TB/s of bandwidth — the platform for national-scale inference and fine-tuning.

GPU memory
141 GB HBM3e
Bandwidth
4.8 TB/s
FP8
4 petaFLOPS
CPU cores
120
Datasheet
GPU servers & clusters

Delivered and deployed as a service

From a single NVIDIA HGX B300 node to GB300 NVL72 racks and next-generation Vera Rubin systems: we source, deliver, install and commission the hardware inside your data centre or sovereign cloud, then operate it with your partner — so you consume AI infrastructure as a service, not as a project.

Blackwell Ultra · 8-GPU node

NVIDIA HGX B300

Eight Blackwell Ultra GPUs on a single HGX baseboard — the building block for dense inference and training servers, delivered in air- or liquid-cooled configurations from our OEM partners.

Blackwell Ultra · rack-scale

NVIDIA GB300 NVL72

72 Blackwell Ultra GPUs and 36 Grace CPUs in one liquid-cooled rack, NVLink-connected to behave as a single accelerator for frontier-scale inference and reasoning workloads.

Next generation

NVIDIA Vera Rubin

NVIDIA's next-generation Rubin GPU and Vera CPU platform. We plan capacity and reserve allocation ahead of availability, so your roadmap is not gated by the supply queue.

  • Delivery, installation & commissioning
  • Validated fabric, power & cooling
  • Operated as a service with your partner

Delivery & deployment as a service

We source, deliver, install and commission the systems, then run GPU management, clustering, inference-as-a-service and LLM deployment on top — handed over to your team or operated for you.

  • Sourcing, delivery & commissioning
  • Inference service enablement
  • Capacity and tenancy planning

AI Training Infrastructure

Turn-key fine-tuning and training environments: datasets, storage, data processing and cluster management.

  • Fine-tuning & model training
  • Dataset and storage architecture
  • Data processing pipelines

Storage & networking

High-throughput parallel storage and InfiniBand/RoCE fabrics engineered for sustained inference and retrieval.

  • Parallel filesystem design
  • InfiniBand / RoCE fabric
  • Power and thermal planning
Model catalogue

200+ optimised models, pre-loaded

Every system arrives with an adapted and specialised open-model library already tuned for the accelerator it ships with.

  • Qwen
  • Llama
  • DeepSeek
  • Mistral
  • MONAI
  • Meditron
  • BLOOM
  • Calme
  • + 200 more
Parts
3 years
Live service
3 years, engineer-led
Support tier
Standard 3-year; on-site service and SLA available through partners
Sizing

Tell us the workload, not the part number

Send us the models you want to run, the concurrency you expect and the space you have. We will come back with a configuration, a power and cooling budget, and the partner who will deliver it.