Custom GPU Inference Rig, Built in Quebec
What is a custom GPU inference rig in Canada? An x86 tower with a discrete GPU, assembled in Laval, Quebec, quoted from current GPU street. D-Central used to build GPU mining rigs; local inference uses the same PSU, PCIe, airflow, heat, and noise skills. It is not a named “Workstation 24” on a shelf. Spark is $9,449 CAD; Strix Halo 128 GB is $7,349 CAD; the tower is quoted because GPU prices move.
Credit NVIDIA or AMD for the card you pick. Credit llama.cpp, Ollama, and vLLM for the software. We assemble, burn-in, and hand over a runbook.
When a tower beats a 128 GB brick
- The model fits in 24–48 GB of GDDR and you care about memory bandwidth.
- You want to swap the GPU later.
- You already have a card, a 240 V circuit, or a room that can take 500–1,200 W and fan noise.
- You want x86 + CUDA without Arm/DGX OS.
When it loses: you need a single 128 GB pool at ~240 W on a desk. That is Spark or Strix Halo. Architecture: Spark vs custom GPU, VRAM vs unified.
How a quote works
- You name the model class (8–14B, 30–32B, 70B quantized), users, and whether the box may reach the internet.
- We pick a GPU from current street (used 3090-class is still the usual 24 GB conversation — 3090 guide).
- Chassis, PSU (transient response matters; mining taught that), CPU/RAM, exhaust.
- You pay before we buy the GPU. Made-to-order is non-refundable once parts are ordered, except DOA/defect.
- Open stack: Ollama, llama.cpp, or vLLM, plus a private UI if you want it.
We will not quote 70B full precision on a 24 GB card. We will not invent tokens/s. Dual-GPU is a separate quote (split VRAM, not a 48 GB unified pool).
Heat and the rest of the shop
A GPU tower makes heat. If you already duct miners, the same 120 mm language applies — see heating with inference and GPU mining rig to inference. Heater SKUs stay on those pages, not on a Law 25 article.
Related products, repair, and setup paths
- self-hosted AI for Bitcoiners hub
- plebs guide to self-hosted AI
- install Ollama in 10 minutes
- LM Studio vs Ollama vs llama.cpp
- connect local AI to Home Assistant and Obsidian
- self-hosted AI troubleshooting
- repurpose mining hardware into an AI hashcenter
- local AI model leaderboards
Last reviewed August 24, 2026.
