Skip to content
Small team, full backlog, zero orders dropped. Support replies are slower than we’d like. Read our status update → Zero orders dropped. Status → 📬 Check your spam folder — most of our replies land there. We do answer. Status update → 📬 Check your spam folder. Status →

Custom GPU Inference Rig

Priced by quote

Built-to-order discrete-GPU inference tower. Same workshop that used to assemble GPU mining rigs. Quoted from current GPU street. Quebec.

  • Condition: Used
  • Availability: Built to order — priced by quote
  • Lead time: Hand-built to order in Montreal, Quebec. We keep inventory lean and build to your spec; your lead time is confirmed with the quote.
  • Returns: Made-to-order work is non-refundable once sourcing or assembly starts, except dead-on-arrival or defective units.
  • Warranty: As stated on this listing at purchase: manufacturer's warranty on new hardware, D-Central's own warranty on refurbished.
  • Support: Ships from Canada with D-Central mining hardware and repair support

Shipping, returns, and warranty details

Request a quote

Tell us your use case — we reply with a spec and a firm quote. Or call 1-855-753-9997.

We accept Bitcoin and Lightning at checkout

We accept:
Shipping calculated at checkout
Ships from Canada
Bitcoin Accepted
Warranty terms shown before purchase
Ships from Canada
Secure Checkout
In-House Repair Experts

Save 3% when you pay in Bitcoin.

Shipping options and lead times are confirmed at checkout
SKU: DC-AI-GPURIG Category:

Questions about this product?

Talk to a mining expert → or call 1-855-753-9997

Description

D-Central used to build GPU mining rigs. Local inference uses the same shop skills: PSU headroom, PCIe, airflow, heat, noise, and a box that stays up. A custom GPU inference rig is a discrete-GPU x86 tower we assemble to your model size — not a vapour “Workstation 24” with a GPU we do not have on the shelf.

What an engagement is: you tell us the model class (for example 8–14B, 30–32B, 70B quantized), concurrent users, whether the machine may reach the internet, and whether you already own a GPU. We quote a chassis, PSU, CPU/RAM, and a specific GPU from current street, then build and burn-in in Laval. The inference stack is the open one: Ollama, llama.cpp, or vLLM, plus a private UI if you want it. Credit the model teams and those runtimes; our job is the hardware and the handover.

Pricing is quoted because GPU street prices move, especially in a memory crunch. A 24 GB class (used RTX 3090-class or current equivalent) is the usual starting conversation. Dual-GPU and 48 GB+ class are quoted separately. You pay before we buy the GPU. Made-to-order work is non-refundable once parts are ordered, except DOA or documented defect.

Honest limits: a discrete GPU is faster than Spark on models that fit in its VRAM, and slower or impossible on models that do not. System RAM is not a substitute for VRAM. We will not quote a 70B full-precision job on a 24 GB card. Airflow and power for a 350–600 W GPU are a mining problem we already know; we will say when a desk is the wrong place for the tower.

Prefer a sealed 128 GB appliance? DGX Spark — $9,449 CAD or Strix Halo 128 GB — $7,349 CAD. Architecture notes: Spark vs custom GPU build, used 3090 for LLMs, hashcenter retrofit.

Complete Your Setup