Skip to content

Custom GPU Inference Rig

Priced by quote

Built-to-order discrete-GPU inference tower. Same workshop that used to assemble GPU mining rigs. Quoted from current GPU street. Quebec.

  • Condition: Used
  • Availability: Built to order — priced by quote
  • Lead time: Hand-built to order in Montreal, Quebec. We keep inventory lean and build to your spec; your lead time is confirmed with the quote.
  • Returns: Made-to-order work is non-refundable once sourcing or assembly starts, except dead-on-arrival or defective units.
  • Warranty: As stated on this listing at purchase: manufacturer's warranty on new hardware, D-Central's own warranty on refurbished.
  • Support: Ships from Canada with D-Central mining hardware and repair support

Shipping, returns, and warranty details

Request a quote

Tell us your use case — we reply with a spec and a firm quote. Or call 1-855-753-9997.

We accept Bitcoin and Lightning at checkout

We accept:
Shipping calculated at checkout
Ships from Canada
Bitcoin Accepted
Warranty terms shown before purchase
Ships from Canada
Secure Checkout
In-House Repair Experts

Save 3% when you pay in Bitcoin.

Shipping options and lead times are confirmed at checkout
SKU: DC-AI-GPURIG Category:

Questions about this product?

Talk to a mining expert → or call 1-855-753-9997

Description

A D-Central custom GPU inference rig is a quoted, built-to-order x86 tower designed around your actual local-AI workload. There is no fictional base configuration and no fixed price hidden behind the quote button. You provide the model, context, concurrent-user and operating constraints; D-Central specifies the GPU, CPU, system memory, storage, motherboard, power supply, chassis and cooling, discloses the condition and source class of key parts, and builds the approved system in Laval, Quebec.

D-Central’s experience assembling and maintaining GPU mining systems transfers directly to power delivery, PCIe layout, sustained heat, airflow, cable management and serviceable components. The workload is different: inference needs the right accelerator software, memory capacity, latency and data controls. NVIDIA or AMD manufactures the GPU and maintains its platform software; component makers manufacture the remaining parts; model authors and runtime projects own their work. D-Central is the independent system builder and integrator. We do not claim manufacturer status, partnership, invented benchmark results or ownership of third-party software.

Pricing: written quote only

This product is priced by quote. GPU pricing, condition and supply move too quickly for an honest universal number. Your quote names every material component, whether a key part is new, refurbished or used, the included assembly and configuration work, applicable tax, shipping, payment schedule and the then-current sourcing estimate. No stock level, delivery date or specific GPU is promised by this page. D-Central does not purchase parts until the commercial terms and technical scope are accepted through the order process.

A 24 GB card can be a sensible starting discussion for models that fit, but it is not the default promise. A used GeForce RTX 3090, a current GeForce card, a professional NVIDIA card, an AMD Radeon or workstation accelerator each carries different memory, power, software, physical and warranty boundaries. Multi-GPU systems add PCIe-lane, chassis, cooling, NUMA and runtime complexity. The quote follows the workload instead of forcing every buyer into the same parts list.

What we need to design the rig

  • Exact model: repository or file name, architecture, parameter count, quantization and model license—not only “I need a 70B.”
  • Serving target: context window, typical prompt size, expected output length, simultaneous users, batching and acceptable response latency.
  • Software: Ollama, llama.cpp, vLLM, a specific CUDA or ROCm release, containers, Python packages, private UI and any application that must connect.
  • Data boundary: fully offline, LAN-only or Internet-connected; encryption, identity, backup and remote-access expectations.
  • Physical boundary: available voltage and circuit, room and ambient conditions, acoustic limit, tower dimensions, network speed, display/KVM needs and service access.
  • Lifecycle: whether future GPU replacement, a second GPU, additional storage, rack conversion or remote management is likely.

Memory fit before brand preference

Inference memory includes model weights, runtime overhead, KV cache, context, temporary buffers and—in multi-user serving—concurrent request state. Quantization can reduce weight memory, but it can also change accuracy, runtime compatibility and output characteristics. System RAM is useful for the operating system, data and CPU offload; it is not a drop-in substitute for fast GPU memory. A model that spills heavily to CPU may technically launch while missing the latency target.

That is why the quote documents a capacity assumption rather than a slogan. We will not promise a full-precision 70B workload on a 24 GB GPU, describe system RAM as VRAM, or use a single tokens-per-second result from another configuration. If a measured acceptance target is needed, the quote must define the exact model file, prompt/context method, runtime version, concurrency, power mode and test procedure. Until that test is performed on the completed build, performance remains an engineering estimate, not a published fact.

NVIDIA, AMD and runtime choices

NVIDIA is usually the direct path when a workload explicitly requires CUDA, TensorRT, a CUDA-only extension or a vendor container. vLLM’s official GPU installation documentation lists its current NVIDIA compute-capability and CUDA requirements and provides official CUDA container guidance. A particular NVIDIA model still has to meet memory, power, slot and driver requirements; the brand alone does not solve those constraints.

AMD can be appropriate when the target workload has a documented ROCm, HIP or Vulkan path and the chosen card appears in the applicable matrix. AMD’s ROCm compatibility matrix is the source of record for supported operating systems, accelerators and frameworks. A Radeon card working through Vulkan in llama.cpp does not imply that every ROCm framework or Python wheel supports it. The quote names the software path we will validate and avoids calling community support “official.”

Ollama documents NVIDIA, AMD and Vulkan GPU options. llama.cpp documents CUDA, HIP and Vulkan builds. vLLM documents NVIDIA CUDA, AMD ROCm and CPU installation paths, with native GPU serving on Linux rather than Windows. These projects and their maintainers deserve credit. D-Central’s contribution is selecting a supportable combination, installing the agreed versions and documenting the handoff.

Single GPU or multi-GPU?

A single GPU is simpler to cool, power, update and debug. If it meets the model and concurrency target, it is normally the cleanest service boundary. Two cards are not automatically twice as fast: the runtime must support the required parallelism, the motherboard and CPU must expose usable PCIe lanes, the chassis must physically separate and cool the cards, and the power supply and branch circuit must handle transients and sustained load. Consumer GPUs may not provide the high-speed peer interconnect or virtualization features assumed by enterprise software.

When the model cannot fit on one available card, the design can consider tensor or pipeline parallelism, multiple model replicas, CPU offload, a different quantization or a platform with a larger unified memory pool. Each option changes latency and operational complexity. The quote will state which strategy the parts list is intended to support; it will not imply that every runtime can pool the VRAM of arbitrary cards.

Power, thermals, noise and location

A high-end accelerator turns electricity into sustained heat. The GPU’s nominal board-power figure is only part of the system budget: CPU, memory, storage, fans, conversion losses and transient behavior matter. D-Central sizes the power supply and connectors for the named components and chooses a chassis with a viable airflow route. The quote also states the expected electrical input class and whether the build is suitable for a normal desk environment.

The customer must provide a compliant circuit, ventilation, clear intake/exhaust space and an ambient environment within component specifications. A quiet office target can constrain GPU power and performance. A rack or utility-room target can tolerate more airflow but needs different networking and physical access. We will say when the requested acoustic, thermal and performance goals conflict instead of representing a high-wattage tower as silent.

Storage, network and security

Model libraries can consume storage quickly, and temporary downloads may require space beyond final model size. The quote identifies the boot drive, model/data storage, filesystem assumptions and expansion path. RAID, snapshots and backup are distinct choices; RAID alone is not backup. Unless explicitly scoped, D-Central does not retain a copy of customer data or operate a remote backup service.

A LAN-only interface still needs authentication, firewall rules and an owner for updates. Internet-exposed inference requires a deliberately secured reverse proxy, certificates, identity, logging and maintenance plan; it is not included in a baseline local setup. For offline systems, dependencies, containers and model files need a controlled transfer path. The buyer remains responsible for model licenses, privacy requirements, user authorization and data retention. D-Central can implement a written boundary but cannot infer governance requirements that were not supplied.

Build, configuration and acceptance

After approved parts arrive, D-Central assembles the tower, checks cabling and airflow, updates the agreed baseline firmware where appropriate, installs the named operating system and drivers, and performs component and thermal checks suitable for the configuration. We then install the agreed inference runtime and confirm that its basic GPU path launches with a compatible test workload. The handoff records the material hardware, relevant versions and the actual checks completed.

“Burn-in” is not a promise of future uptime. It is a documented pre-delivery test intended to catch assembly, cooling or component faults under the defined conditions. If your purchase requires an application-level acceptance test, remote-user test, target throughput or a particular model output, that test and pass/fail method must be in the quote. No private production data is used unless both parties agree on handling.

Serviceability and upgrade boundaries

A standard tower can make GPU, storage, fan and power-supply replacement more practical than a tightly integrated appliance. Compatibility still has limits. A later GPU may need more slots, a different connector, higher PSU capacity, newer firmware, a larger chassis or a different software stack. “Upgradeable” means the design uses documented standard interfaces; it does not guarantee compatibility with an unannounced future card.

Opening, modifying, overclocking, reflashing or rewiring parts can affect defect review or manufacturer terms. Ask before changing the delivered build if return eligibility matters. D-Central support covers the configuration and work identified in the order. Component-maker processes, software communities and model publishers retain their own terms. This page does not create an unstated parts, labour or performance warranty.

When to choose another architecture

Choose a DGX Spark at $9,449 CAD when its 128 GB unified memory and integrated Arm64/CUDA platform are a closer fit than component expansion. Choose the AMD Strix Halo 128 GB desktop at $7,349 CAD when compact x86 and a large shared memory pool matter more than replaceable discrete GPUs. Choose a custom tower when a specific discrete accelerator, standard PC architecture, service access or future component change is central. Our architecture comparison and used RTX 3090 guide explain two common decision branches.

Exact compatibility and order boundary

  • Only the signed/accepted quote defines the build: component manufacturer and model, condition, quantities, storage, OS, software scope, validation procedure, price, tax, shipping and sourcing estimate.
  • No silent substitutions: any material GPU, motherboard, PSU, storage or chassis change requires documented approval.
  • No assumed software support: the requested model, quantization, runtime and OS must match official or clearly identified community support. CUDA-only needs NVIDIA; ROCm support must match AMD’s matrix; native vLLM GPU serving is Linux.
  • No undisclosed performance promise: model fit and throughput are not guaranteed unless the quote includes a reproducible acceptance test.
  • Buyer-provided environment: suitable electrical service, ventilation, network security, physical access, software/model rights and post-handoff backups.

Payment and returns for a custom build

The rig is quote-priced. The written quote states when payment is due and when sourcing or build work begins. No price or stock field on this page represents a prebuilt unit. Changes after parts are committed may be impossible or may require a new quote.

D-Central’s Return and Refund Policy states that custom builds, special orders and made-to-order work are non-refundable once sourcing, assembly, configuration or service work has started, except where they arrive defective or dead on arrival. The policy provides 30 days from delivery to report a defect or DOA for review; it is not a buyer’s-remorse window. Contact support with the order number, fault description and evidence, and wait for authorization before returning anything. The system must be complete and not damaged by misuse, installation error, modification or improper handling. Shipping and any repair, replacement, exchange or refund decision follow the approved process. The current policy and accepted quote control; this page adds no separate warranty claim.

Custom GPU inference rig FAQ

How much does a rig cost?

It is quote-only because the GPU, condition, memory, platform and service scope drive the price. The quote itemizes the actual build. D-Central does not advertise a teaser price for hardware that has not been specified.

Can I supply my own GPU?

Potentially. Provide the exact model, photos, condition and ownership history. The quote will state whether D-Central accepts it, what compatibility can be checked, who bears failure risk and whether separate diagnostics are required.

Do you use new or used parts?

Only as disclosed in the quote. A proposed used or refurbished GPU will be labelled as such; it will not be represented as new. Other component conditions are also stated.

Will two GPUs combine their memory?

Not automatically. The runtime and model must support an appropriate parallel or offload strategy. The quote identifies the intended method and hardware topology; it does not promise transparent memory pooling.

Can you guarantee tokens per second?

Only if a paid scope defines a reproducible acceptance test and the accepted quote makes that result a requirement. Otherwise we provide sourced component facts, a fit rationale and recorded build checks—not invented benchmarks.

Can the rig run fully offline?

Yes, if planned. Specify the model and dependency transfer method, update policy, local authentication and interface requirements. Third-party license or activation rules still apply.

What happens if the system arrives defective?

Report the defect or DOA within 30 days of delivery with the order number and evidence, then wait for return authorization. Made-to-order systems are not returnable for change of mind once sourcing, assembly, configuration or service has begun.

Sources and review: official NVIDIA/CUDA, AMD ROCm compatibility, Ollama GPU, llama.cpp build and vLLM GPU installation documentation; D-Central Return and Refund Policy. Reviewed 25 August 2026. The accepted quote, not this general page, defines the actual components and validation scope.

Complete Your Setup