Description
D-Central used to build GPU mining rigs. Local inference uses the same shop skills: PSU headroom, PCIe, airflow, heat, noise, and a box that stays up. A custom GPU inference rig is a discrete-GPU x86 tower we assemble to your model size — not a vapour “Workstation 24” with a GPU we do not have on the shelf.
What an engagement is: you tell us the model class (for example 8–14B, 30–32B, 70B quantized), concurrent users, whether the machine may reach the internet, and whether you already own a GPU. We quote a chassis, PSU, CPU/RAM, and a specific GPU from current street, then build and burn-in in Laval. The inference stack is the open one: Ollama, llama.cpp, or vLLM, plus a private UI if you want it. Credit the model teams and those runtimes; our job is the hardware and the handover.
Pricing is quoted because GPU street prices move, especially in a memory crunch. A 24 GB class (used RTX 3090-class or current equivalent) is the usual starting conversation. Dual-GPU and 48 GB+ class are quoted separately. You pay before we buy the GPU. Made-to-order work is non-refundable once parts are ordered, except DOA or documented defect.
Honest limits: a discrete GPU is faster than Spark on models that fit in its VRAM, and slower or impossible on models that do not. System RAM is not a substitute for VRAM. We will not quote a 70B full-precision job on a 24 GB card. Airflow and power for a 350–600 W GPU are a mining problem we already know; we will say when a desk is the wrong place for the tower.
Prefer a sealed 128 GB appliance? DGX Spark — $9,449 CAD or Strix Halo 128 GB — $7,349 CAD. Architecture notes: Spark vs custom GPU build, used 3090 for LLMs, hashcenter retrofit.


Reviews
There are no reviews yet.