Soika AI Workstation SM H200
When the model is the constraint, not the budget.
141 GB of HBM3e per GPU at 4.8 TB/s, four petaFLOPS of FP8 and up to seven MIG instances per card. This is the configuration behind sovereign inference services and in-house fine-tuning programmes.
2–4× NVIDIA H200 SXM
- GPU memory
- 141 GB HBM3e
- Bandwidth
- 4.8 TB/s
- FP8
- 4 petaFLOPS
- CPU cores
- 120
GPU memory
141 GB HBM3e
Bandwidth
4.8 TB/s
FP8
4 petaFLOPS
CPU cores
120
Data-centre class HBM3e for serious inference and training.
141 GB of HBM3e per GPU and 4.8 TB/s of bandwidth — the platform for national-scale inference and fine-tuning.
- 141 GB of HBM3e GPU memory
- 4.8 TB/s of memory bandwidth
- 4 petaFLOPS of FP8 performance
- Roughly 2× LLM inference performance over the previous generation
- Up to 7 MIG instances at 18 GB each
- Confidential computing supported
The full configuration
- Parts
- 3 years
- Live service
- 3 years, engineer-led
- Support tier
- Standard 3-year; on-site service and SLA available through partners
- Model
- SOIKASMH200
- Device type
- Dual-socket workstation, 120 CPU cores
- Graphics
- 2–4× NVIDIA H200 (SXM form factor)
- Processor
- 2× Intel® Xeon® w9-3595X — 120 cores / 240 threads
- Memory
- 8× 64GB DDR5-4800 2Rx4 ECC RDIMM (512GB)
- Storage
- 4× 8TB NVMe
- GPU memory
- 141 GB HBM3e
- Memory bandwidth
- 4.8 TB/s
- Multi-Instance GPU
- Up to 7 MIGs @ 18 GB each
- Decoders
- 7 NVDEC, 7 JPEG
- Thermal design power
- Up to 700 W, configurable
- Confidential computing
- Supported
- Network
- 2× 10GbE RJ45 + 1× management LAN
Licensed, installed and tuned before it ships
Soika hardware exists to remove the setup problem. GPU clustering, model serving and agent design are configured at the factory, not on your time.
Soika Stack
GPU clustering, LLM deployment and inference management, pre-configured and licensed.
Soika Mockingjay
No-code AI agent design and agent-fleet management with the Enterprise licence.
Ubuntu LTS
Hardened Linux base image with the full CUDA and container toolchain already in place.
Pre-loaded model catalogue
- Qwen
- Llama
- DeepSeek
- Mistral
- MONAI
- Meditron
- BLOOM
- Calme
- 200+ optimised models
What teams run on these systems
The same box serves very different work depending on who owns it.
Healthcare
Record summarisation, diagnostic support, research and treatment planning.
Government
Citizen-service agents, smart-city systems and in-country model hosting.
Finance
Reporting, risk assessment, trading research and investment analysis.
Legal
Contract drafting, case-law summarisation and legal co-pilots.
Software
Code generation, debugging, application design and internal developer tooling.
Research & academia
Advanced data analysis, reasoning tasks, language and AI research.
E-commerce
Personalised recommendation and supply-chain intelligence.
Edge & autonomous systems
Local inference for robotics, drones and IoT without cloud access.
Other workstations
Configure a SM H200
Tell us the models you plan to run and the environment it will live in. We will confirm the configuration, power and cooling budget, and introduce the partner who will deliver and commission it.