Skip to content
2–4× NVIDIA H200 SXM

Soika AI Workstation SM H200

When the model is the constraint, not the budget.

141 GB of HBM3e per GPU at 4.8 TB/s, four petaFLOPS of FP8 and up to seven MIG instances per card. This is the configuration behind sovereign inference services and in-house fine-tuning programmes.

WorkstationSOIKASMH200
PWR

2–4× NVIDIA H200 SXM

GPU memory
141 GB HBM3e
Bandwidth
4.8 TB/s
FP8
4 petaFLOPS
CPU cores
120

GPU memory

141 GB HBM3e

Bandwidth

4.8 TB/s

FP8

4 petaFLOPS

CPU cores

120

Key features

Data-centre class HBM3e for serious inference and training.

141 GB of HBM3e per GPU and 4.8 TB/s of bandwidth — the platform for national-scale inference and fine-tuning.

  • 141 GB of HBM3e GPU memory
  • 4.8 TB/s of memory bandwidth
  • 4 petaFLOPS of FP8 performance
  • Roughly 2× LLM inference performance over the previous generation
  • Up to 7 MIG instances at 18 GB each
  • Confidential computing supported
Specifications

The full configuration

Parts
3 years
Live service
3 years, engineer-led
Support tier
Standard 3-year; on-site service and SLA available through partners
Model
SOIKASMH200
Device type
Dual-socket workstation, 120 CPU cores
Graphics
2–4× NVIDIA H200 (SXM form factor)
Processor
2× Intel® Xeon® w9-3595X — 120 cores / 240 threads
Memory
8× 64GB DDR5-4800 2Rx4 ECC RDIMM (512GB)
Storage
4× 8TB NVMe
GPU memory
141 GB HBM3e
Memory bandwidth
4.8 TB/s
Multi-Instance GPU
Up to 7 MIGs @ 18 GB each
Decoders
7 NVDEC, 7 JPEG
Thermal design power
Up to 700 W, configurable
Confidential computing
Supported
Network
2× 10GbE RJ45 + 1× management LAN
Included software

Licensed, installed and tuned before it ships

Soika hardware exists to remove the setup problem. GPU clustering, model serving and agent design are configured at the factory, not on your time.

Soika Stack

GPU clustering, LLM deployment and inference management, pre-configured and licensed.

Soika Mockingjay

No-code AI agent design and agent-fleet management with the Enterprise licence.

Ubuntu LTS

Hardened Linux base image with the full CUDA and container toolchain already in place.

Pre-loaded model catalogue

  • Qwen
  • Llama
  • DeepSeek
  • Mistral
  • MONAI
  • Meditron
  • BLOOM
  • Calme
  • 200+ optimised models
Where it earns its place

What teams run on these systems

The same box serves very different work depending on who owns it.

Healthcare

Record summarisation, diagnostic support, research and treatment planning.

Government

Citizen-service agents, smart-city systems and in-country model hosting.

Finance

Reporting, risk assessment, trading research and investment analysis.

Legal

Contract drafting, case-law summarisation and legal co-pilots.

Software

Code generation, debugging, application design and internal developer tooling.

Research & academia

Advanced data analysis, reasoning tasks, language and AI research.

E-commerce

Personalised recommendation and supply-chain intelligence.

Edge & autonomous systems

Local inference for robotics, drones and IoT without cloud access.

Order

Configure a SM H200

Tell us the models you plan to run and the environment it will live in. We will confirm the configuration, power and cooling budget, and introduce the partner who will deliver and commission it.