Star
Code
GitHub

How to Build an 8x RTX 5090 GPU Server: 256GB VRAM

How to Build an 8x RTX 5090 GPU Server: 256GB VRAM

An 8x RTX 5090 build gives you 256GB of combined VRAM, 1,676 FP32 TFLOPS, and 14,336 GB/s of aggregate bandwidth in a single on-prem machine. Eight 32GB cards hold the largest open models and serve them to a whole team at once, with room to run several models in parallel. Built with RTX 4090s instead, the same housing delivers 192GB, since each 4090 carries 24GB.

Around the cards sits a server platform: two AMD EPYC 9004 Genoa CPUs on an ASRock Rack GENOA2D24G-2L+ board, 192GB of DDR5 ECC RDIMM, a 1TB NVMe boot drive, dual 1 GbE networking, and a BMC for lights-out management. The dual-socket board and its 24 DIMM slots exist to feed eight cards with enough PCIe lanes and memory bandwidth, which a single-socket platform cannot do.

The problem eight cards create, and how this solves it

Fitting eight full-speed GPUs in one machine is a real engineering problem, not just a bigger version of a four-card build. Eight cards need eight full PCIe Gen5 x16 links to the CPUs, and no consumer or single-socket board has the lanes. They also will not physically fit in a normal case, which is why open builds at this scale usually end up on improvised frames.

This build solves both with MCIO cabling and a purpose-built housing. Each GPU is fed by two MCIO cables that together carry a full x16 link, routed from the board through eight MCIO-to-PCIe adapter boards, so every card gets full bandwidth straight from the CPUs with no switch in the path. The cards mount to a CNC-milled aluminium housing designed around them, rather than a frame improvised to make them fit. That is the difference between a machine built to run continuously and one assembled to prove a point.

The problem eight cards create, and how this solves it

The full build

The build is fully specified, and every part is listed below with its exact type. Quantities are per machine.

Component

Part

Detail

GPUs

8x NVIDIA RTX 5090 (or 4090)

32GB each (24GB for 4090), 256GB total (192GB with 4090)

Motherboard

ASRock Rack GENOA2D24G-2L+

Dual SP5, PCIe over MCIO, BMC

CPUs

2x AMD EPYC 9004 (Genoa)

SP5 socket

Memory

DDR5 ECC RDIMM, 192GB

Configurable 1 to 24 DIMMs

Storage

1TB NVMe

Capacity configurable

GPU interconnect

16 MCIO cables, 8 MCIO-to-PCIe adapters

Two cables per GPU, x16 as 2x x8

GPU power

8x 12V-2x6, 600W rated

One per card

Power

4x 2000W CRPS, power distribution board

8000W total, 5,100W draw

Cooling

12 case fans, fan hub

Center fan wall

Chassis

CNC-milled anodised aluminium housing

Full set, milled from STEP files

The number to notice is the interconnect: sixteen MCIO cables into eight adapters. That is not extra complexity for its own sake. It is what delivers a genuine Gen5 x16 to each of the eight cards from the two CPUs, which is the whole reason to build on a dual-EPYC board rather than improvise lanes with switches or risers that drop to a narrower link.

Power and the circuit a server needs

Eight RTX 5090s plus dual EPYC CPUs draw about 5,100W under load, delivered by four 2000W CRPS supplies for 8000W total through a power distribution board. That draw is well past anything a standard outlet carries, so this build needs a dedicated high-current circuit, a 240V supply sorted before delivery, the same as a piece of server-room equipment. This is not a machine you plug into a wall socket.

The redundant power design is the payoff. Four CRPS modules feeding a distribution board means a single supply failing does not take the machine down, the same standard used in data-center servers. At 5,100W draw against 8000W of supply, the system runs around 64 percent of capacity, which keeps the supplies cool and leaves margin for the transient spikes eight cards produce under load.

Power and the circuit a server needs

Cooling eight cards

Eight cards in one housing is the hardest thermal problem in the lineup, and it is solved with a center fan wall rather than by hoping airflow finds its way through. Twelve case fans on a controlled hub push air across all eight cards in a defined path, with the fan wall sitting between the card banks so no card is starved of intake air. Airflow at this scale is engineered into the housing, not added afterward.

This is also why the housing is CNC-milled rather than printed. Eight cards, their adapters, the cabling, and a 12-fan wall need a rigid, precise structure to hold alignment and route air, which is what the milled aluminium provides. The result runs loud, like any dense server, which is part of why this build belongs on a server floor or in a dedicated room rather than an office.

Cooling eight cards

Building it, step by step

The mechanical build is documented as 39 steps from bare plates to a running machine, with a photo of every part, in the build repository. The housing is milled from the published STEP files, and a component checklist covers the electronics layout. Rather than reproduce all 39 steps, here are the stages that matter most and the order that keeps the build on track.

Step 1: Lay out and check every part

Check every part against the component checklist before anything goes together. At 39 steps and this many parts, a missing adapter or cable found mid-build is expensive, so confirm the count first.

Step 2: Build the dual-EPYC board

Seat both EPYC 9004 CPUs in their SP5 sockets, fit both coolers, and populate the DDR5 ECC RDIMM across the board's channels. With 24 DIMM slots and dual sockets, follow the board manual for the population order rather than filling slots by eye.

Step 3: Mount the board and the power system

Secure the board to the housing, install the four CRPS supplies and the power distribution board, then run the power-board-to-motherboard harness. Getting the power system in before the GPUs keeps the dense center of the machine reachable.

Step 4: Install the MCIO adapters and cabling

Fit the eight MCIO-to-PCIe adapter boards and run the sixteen MCIO cables, two per card, from the board to the adapters. This is the step that defines the build, so route the cables clear of the fan wall and label them, because tracing an MCIO pair later is difficult once the cards are in.

Step 5: Install the eight GPUs

Seat each of the eight cards into its adapter and secure it to the housing. Work from one side to the other so you are never reaching past an installed card, and check each sits square before moving on.

Step 6: Cable GPU power and wire the fans

Connect a 12V-2x6 to each of the eight cards, seating each until it clicks, then connect the twelve fans to the fan hub and verify the PWM headers. A loose 12V-2x6 is the connector behind most melting incidents, so confirm each with a gentle tug.

Step 7: Final pass and first boot

Confirm all eight power cables, all sixteen MCIO cables, every adapter, and the fan wall before closing the housing. Then connect the power and perform first boot into the BIOS step.

Building it, step by step

If sourcing and assembling a server-class machine is not where you want to spend your time, the Rig ships the housing, adapters, cooling, and power as one wired and burn-tested unit, and you supply the board, memory, and cards.

BIOS settings that matter

On the GENOA2D24G-2L+ the setting that matters most is PCIe link width, because the board feeds all eight GPUs over MCIO rather than physical slots. Set all eight MCIO pairs to Gen5 x16, enable Resizable BAR, and verify Above 4G Decoding, which is usually enabled by default on this platform. The exact link-width paths are:

Advanced -> Chipset Configuration -> PCIE link width

  -> set MCIO2/1, 4/3, 6/5, 8/7, 12/11, 14/13, 16/15, 18/17 to x16

Advanced -> PCI Subsystems Settings -> Enable Re-size BAR support

Setting all eight pairs to x16 is what turns the sixteen MCIO cables into eight full-bandwidth links. Beyond that, disable C-states and ASPM so power management does not throttle the cards, and run the memory at its rated speed. The board and BMC manuals in the repo carry the exact menu locations, and the multi-GPU BIOS setup guide has the cross-board sequence.

BIOS settings that matter

Verifying the build

A correct 8x RTX 5090 build shows all eight cards detected, each reporting its full 32GB, and each linked at a full PCIe Gen5 x16. Boot into Linux, install the NVIDIA driver and CUDA toolkit, then run nvidia-smi and nvtop. What to confirm:

  • All eight GPUs appear in nvidia-smi
  • Each card reports its full 32GB of VRAM
  • PCIe link width is x16 on every card, with none dropped to x8, x4, or x1
  • Temperatures and power draw are sane at both idle and under load

The failure to watch for is a card linking at a narrower width, which on an MCIO build usually points to a cable pair seated wrong or a BIOS link-width setting rather than a bad card. With eight cards, one dropped link costs an eighth of your throughput, so confirm all eight before you commit the machine to work.

Serving models

With 256GB of VRAM verified, the build serves the largest open models to a team. The common options are:

  • vLLM, for high-throughput serving
  • Ollama, for the simplest setup
  • llama.cpp, for flexibility across hardware
  • Grid, an open orchestrator that pools this machine with any others you own into one local network

On top of a Grid network, apps like OpenClaw and Hermes Agent, or your own, run against the pooled compute. Nothing you serve leaves the building, which for a business keeping IP and customer data in-house is the entire reason to run AI on-prem rather than in the cloud.

With 256GB of VRAM verified, the build serves the largest open models to a team.

Specs at a glance

Prices change periodically. Check the product page for current pricing.

Spec

8x RTX 5090 server

GPUs

8x RTX 5090, 256GB VRAM (or 8x 4090, 192GB)

VRAM per card

32GB (24GB for 4090)

Memory bandwidth

14,336 GB/s aggregate

Compute

1,676 FP32 TFLOPS

Platform

2x AMD EPYC 9004 Genoa, ASRock Rack GENOA2D24G-2L+

System RAM

192GB DDR5 ECC RDIMM

Storage

1TB NVMe

Interconnect

16 MCIO cables, 8 adapters, Gen5 x16 per card

Networking

2x 1 GbE, BMC

Power

4x 2000W CRPS, 8000W total, 5,100W draw

Cooling

12 case fans, center fan wall

Chassis

CNC-milled anodised aluminium, 15.5 x 15.5 x 24 in, 110 lb

Frequently asked questions

How much VRAM does an 8x RTX 5090 server have?

An 8x RTX 5090 server has 256GB of VRAM, 32GB per card. That holds the largest open models and serves several at once to a team. Built with RTX 4090s instead, the same housing gives 192GB. Aggregate memory bandwidth is 14,336 GB/s across the eight cards.

How do you get eight full PCIe x16 links to the GPUs?

Eight full Gen5 x16 links come from a dual-EPYC board and MCIO cabling. Each card is fed by two MCIO cables carrying a full x16, routed through eight MCIO-to-PCIe adapters straight from the two CPUs. This avoids the switches and risers that drop cards to a narrower link on lesser platforms.

What power does an 8x GPU server need?

An 8x RTX 5090 server draws about 5,100W under load from four 2000W CRPS supplies. That needs a dedicated 240V high-current circuit, sorted before delivery, like any server-room equipment. The redundant supplies mean a single unit failing does not take the machine down mid-run.

Can you build a GPU server with RTX 4090s instead of 5090s?

You can build this server with either card. The same housing and platform take eight RTX 4090s at 24GB each for 192GB total, or eight RTX 5090s at 32GB each for 256GB. The 5090s add VRAM and bandwidth, while 4090s can lower cost depending on current pricing.

Why is the housing CNC-milled instead of 3D-printed?

The housing is CNC-milled aluminium because eight cards, their adapters, the cabling, and a 12-fan wall need a rigid, precise structure to hold alignment and route airflow. Milled aluminium provides that strength and doubles as a heat path, where a printed frame would flex under the weight of eight cards.

Is an on-prem GPU server worth it versus the cloud?

For a business running AI continuously on private data, an on-prem server keeps IP and customer data in the building and replaces per-hour cloud billing with a fixed hardware cost. Heavy continuous use crosses the break-even point relatively quickly. Light or occasional use often still favours renting.

The bottom line

An 8x RTX 5090 server is the point where local AI becomes a piece of business infrastructure: 256GB of VRAM, dual EPYC CPUs, redundant power, and a milled aluminium housing built to run continuously on your own floor. What makes it work is the engineering around the cards, MCIO cabling for eight true x16 links, a fan wall sized for eight cards, and redundant power, rather than eight GPUs forced into a frame that was never meant to hold them.

References

  • NVIDIA, RTX 5090 specifications and Blackwell architecture overview, nvidia.com
  • Autonomous open hardware build repository, 8x RTX 5090 build and bill of materials, github.com/autonomous-ai/autonomous-computer

How to Build an 8x RTX 5090 GPU Server: 256GB VRAM