01
CPU & ECC memory
Choose a host platform that can feed the GPUs. We scope CPU core count, clock characteristics, memory channels, DIMM population, capacity, and NUMA placement against the application.
- AMD EPYC or Intel Xeon options on a qualified server platform
- ECC RDIMM capacity, speed, and channel population matched to the motherboard
- PCIe lane allocation for GPUs, NICs, NVMe, and management devices
02
NVMe storage & data movement
Separate boot, persistent datasets, checkpoints, scratch, and cache. Raw drive capacity is only one part of the design.
- Enterprise drive endurance, power-loss protection, and hot-swap support
- Mirrored boot options and workload-appropriate data protection
- Local versus shared storage, usable capacity, and recovery requirements
03
Network fabric
Define the client, management, storage, and GPU scale-out networks before choosing NICs and switches.
- 25/100/200/400 GbE and InfiniBand options, platform-dependent
- Port count, optics, transceivers, cable length, and switch compatibility
- RDMA, congestion management, redundancy, and GPU-to-NIC affinity
04
Power, cooling & rack integration
A rack-ready server needs a facility-ready plan. The quotation defines the electrical and thermal assumptions for the chosen configuration.
- Rack height, depth, mass, rail compatibility, and service clearance
- Sustained system load, input voltage, plugs, PDUs, and A/B feeds
- Supported air or liquid cooling and full-load PSU failover requirements
