AMD Strix Halo 128GB in Canada: Ryzen AI Max+ 395 Guide (2026)
AMD Strix Halo 128GB is a 16-core Zen 5 APU with Radeon 8060S (40 CU RDNA 3.5) and 128GB unified LPDDR5X-8000 on a 256-bit bus (256 GB/s). It runs x86 Windows 11 or Linux with ROCm 6.x/Vulkan, gives ~96GB usable VRAM for 70B-class models locally, and games at ~RTX 4060 Laptop level. In Canada it is $7,349 CAD in-stock from D-Central vs DGX Spark at $9,449 CAD when native CUDA is required.
Updated 24 August 2026 – If you want one local box that runs large open-weight models, Adobe and Steam, without learning a new architecture, this is the shortlist. AMD’s Ryzen AI Max+ 395 – codename Strix Halo – puts CPU, GPU and NPU on one die with all 128GB in one pool. No discrete GPU, no 24GB VRAM ceiling, no Arm translation. Below is the Canadian, in-stock, benchmark-honest version for buyers in Montreal and Toronto. D-Central stocks the AMD Strix Halo 128GB desktop at $7,349 CAD and the NVIDIA DGX Spark at $9,449 CAD when you need CUDA. Benchmarks are attributed, limits stated, lab numbers pending with methodology.
For context on where Halo sits in the Canadian local-AI stack, bookmark our local AI hardware guide and the 60-second chooser if you want a fast filter before you read 4,000 words. Developers mapping models to VRAM should also keep our model database and GPU/LLM compatibility table open in another tab.
What AMD Strix Halo Actually Is
Strix Halo is not a laptop chip in a desktop case. It is AMD’s first true halo APU for workstations: the Ryzen AI Max+ 395. Think of it as a 16-core desktop CPU and a mid-range discrete GPU fused into one SoC with a memory system wide enough to feed both. The whole point is unified memory. On a normal desktop you have 32GB or 64GB of system RAM and 8-24GB of VRAM bolted to the graphics card, and moving a 70B model between them is the bottleneck. On Halo there is one 128GB pool of LPDDR5X-8000 that the CPU and the Radeon 8060S both address. The BIOS lets you carve up to ~96GB to the GPU, leaving ~32GB for the OS, which is why a 128GB Halo can hold models that would OOM on a 24GB RTX 4090 without offloading.
AMD credits the platform as Zen 5 + RDNA 3.5 + XDNA2. Zen 5 gives native x86_64, RDNA 3.5 gives Vulkan and ROCm compute, XDNA2 gives ~50 TOPS for Windows Studio Effects and Copilot+ – not for LLM decode today, which runs on the GPU. The chassis we sell is a Framework Desktop-class / Minisforum-class 120W desk box, not the Micro Center bare-board SKU seen in US press. Ours ships as a complete system from Laval, QC, with Canadian warranty, no US brokerage or duties, and Windows 11 and Linux validated. If you are comparing to NVIDIA, read our DGX Spark vs Strix Halo side-by-side.
| Specification | AMD Strix Halo 128GB (Ryzen AI Max+ 395) |
|---|---|
| SoC | AMD Ryzen AI Max+ 395 (Strix Halo) – single-die APU |
| CPU | 16x Zen 5 cores / 32 threads, desktop-class IPC, x86_64 |
| GPU | Radeon 8060S – 40 Compute Units, RDNA 3.5, Vulkan + ROCm 6.x |
| NPU | XDNA2 – ~50 TOPS (Windows Studio Effects, future Copilot+ offload) |
| Memory | 128GB LPDDR5X-8000 unified, soldered (not upgradeable) |
| Bus / Bandwidth | 256-bit memory bus – 256 GB/s unified bandwidth |
| TDP | cTDP 45-120W configurable; ~120W full-system wall in our chassis |
| OS | Windows 11 Pro (x86) or Linux (Ubuntu/Fedora) – native drivers |
| Price (Canada) | $7,349 CAD – in-stock Canada, Montreal warranty & support |
Two notes because they are purchase blockers: 128GB is soldered and not upgradeable, and there is no CUDA on this GPU. You run ROCm 6.x, Vulkan, or DirectML. For many open-weight LLM workflows that is fine – Ollama and llama.cpp ship Vulkan builds for RDNA 3.5 – but CUDA-only kernels need porting or a different box. Credits: APU per AMD disclosure; memory/bandwidth per AMD sheets; ROCm/Vulkan per AMD docs and Ollama/llama.cpp notes.
Why x86 + Windows Matters — The One Thing DGX Spark Cannot Do
DGX Spark is superb if your stack is pure Linux + CUDA. It runs DGX OS (Ubuntu) on a 20-core Arm Grace CPU with Blackwell GPU. What it does not run is your Windows world. If your day job is Lightroom, Premiere, SolidWorks, Excel with COM add-ins, legacy x86 tools, or Steam with kernel anti-cheat, Spark will feel foreign – it is a Linux appliance. Strix Halo is a normal Windows 11 PC that happens to have 96GB of VRAM. That is the easiest buying filter we use:
- If you need Windows and x86 compatibility on the same desk that runs LLMs, Halo is the default. See also DGX Spark vs custom AI build for the tower-PC alternative.
- If you need native CUDA, Blackwell-optimised kernels, and 4TB of NVMe on the box, Spark is the default. Our DGX Spark Canada page covers that system at $9,449 CAD.
The table below is deliberately boring – it is every app you will open before you ever launch Ollama.
| Software / Task | Strix Halo 128GB – x86 Windows 11 / Linux | DGX Spark – Arm DGX OS (Ubuntu) |
|---|---|---|
| Adobe CC (Premiere, Photoshop, Lightroom), Office 365, VS Code, Chrome extensions | Native – install and sign in | Not supported natively; needs remote desktop or second PC |
| Steam, Battle.net, kernel anti-cheat titles | Native – Radeon 8060S, see gaming section | No Windows / no native anti-cheat support |
| Legacy x86 enterprise apps, fiscal software, browser plugins | Native x86_64 | x86 emulation only; often blocked by licensing |
| Ollama, llama.cpp, ComfyUI, Whisper, PyTorch | Native via Vulkan / ROCm / DirectML | Native via CUDA – faster per StorageReview July 2026 (see performance section) |
| CUDA-only libraries (TensorRT, some FlashAttention builds, custom CUDA kernels) | Requires ROCm port or Vulkan path – not drop-in | Native CUDA – this is the Spark advantage |
| Deployment target | Day workstation + night AI lab on one box | Dedicated AI appliance best paired with a separate PC/Mac |
For Quebec buyers: our Laval counter demos Halo in French or English, invoices are bilingual, and we advise on Law 25 for local inference (personal data stays on-premise). The box ships with a 1-year D-Central warranty handled in Montreal, not a cross-border RMA. The Strix Halo 128GB at $7,349 CAD is stocked for Quebec and Canada-wide courier; we do not drop-ship a US Micro Center SKU or bill in USD.
128GB Unified Memory: How It Works & What It Fits
Two numbers matter: how many gigabytes the weights need, and how much of the 128GB you can give the GPU. UEFI exposes a configurable UMA frame buffer. Our builds ship with 96GB for graphics and ~32GB for Windows/Linux, adjustable in BIOS. That 96GB is LPDDR5X-8000 on a 256-bit bus at 256 GB/s – less bandwidth than GDDR6 on an RTX 4090 (1 TB/s) but far more capacity, and capacity determines whether a model loads at all. Bandwidth governs tokens per second.
How many parameters fit? A useful rule of thumb for VRAM budgeting:
- Q4_0 / Q4_K_M quantisation ≈ 0.55-0.60 GB per billion parameters
- Q8_0 ≈ 1.05 GB per billion parameters
- FP16 ≈ 2.0 GB per billion parameters
- Add ~10-15% KV cache for long context (32k-128k), plus OS reserve
The mini-table below uses those rules and labels the fit against a 96GB carve-out. Keep our model database open for current weight sizes – new distils appear monthly.
| Model (example) | Precision | Approx. Weight Size | Fits in ~96GB Unified? |
|---|---|---|---|
| Llama 3.1 8B | Q4_K_M | ~4.9 GB | Yes – trivial, with full context headroom |
| Qwen3 32B / Gemma 3 27B | Q4_K_M | ~18-19 GB | Yes – plus long context |
| Llama 3.1 70B | Q4_K_M | ~40-42 GB | Yes – the Halo sweet spot |
| Llama 3.1 70B | Q8_0 | ~73 GB | Yes – within 96GB, tight with 128k context |
| Llama 3.1 70B | FP16 | ~140 GB | No – needs offload or smaller quant |
| DeepSeek R1 70B distil / Qwen3 72B | Q4_K_M | ~41-44 GB | Yes – similar to Llama 70B Q4 |
| 405B-class (Llama 3.1 405B) | Q4_K_M | ~230+ GB | No – exceeds unified pool; needs quant+offload or multi-box |
In practice, most teams standardise on a 70B Q4 for reasoning plus a 27-32B Q8 for speed, with a 405B quant on a NAS for occasional offload. Both the 70B Q4 and 32B Q8 fit together if you manage context, which is why we position this as the first 70B-native Windows PC under $8,000 CAD. For sizing help, see our GPU/LLM compatibility checker and open-weight comparison. The Strix Halo product page has BIOS allocation guidance.
Unified memory is not magic. Give 96GB to the GPU and Windows has 32GB left. Opening 40 Chrome tabs, Premiere, and a 70B Q8 at 64k context will pressure the system partition. Measure with the context you actually intend to use. Our lab protocol will publish VRAM allocations per model.
ROCm vs CUDA on Strix Halo: Honest Software Stack in 2026
ROCm in 2026 is no longer the experiment it was in 2023, but it is still not CUDA. Here is the plain version from a shop that sells both stacks:
- CUDA (DGX Spark) is the incumbent. If a paper ships code, it usually ships CUDA kernels. Everything just builds.
- ROCm 6.x (Strix Halo) covers PyTorch, ONNX, and most of the Hugging Face ecosystem on Linux, plus Vulkan and DirectML paths on Windows. Ollama and llama.cpp both publish Vulkan-accelerated builds that run well on the 8060S without touching ROCm.
- Vulkan / DirectML is often the easiest path on Halo under Windows: no ROCm install, just a Vulkan runtime, and you get GPU inference for Ollama/llama.cpp immediately. ROCm on Linux gives you broader PyTorch coverage when you need it.
Where that leaves you day-to-day is summarised below. Green = works today without porting. Yellow = works with caveats or a build flag. Red = needs a different kernel or a different box.
| Workload / Framework | On Strix Halo (ROCm 6.x / Vulkan / DirectML) | Notes |
|---|---|---|
| Ollama, llama.cpp, LM Studio | Green | Vulkan build recommended on Windows; models load into unified VRAM. Credit: Ollama + llama.cpp. |
| PyTorch + Transformers (inference) | Green | ROCm wheels on Linux; DirectML/PyTorch on Windows. Most open-weight checkpoints work. |
| Whisper / Faster-Whisper, Piper TTS, STT | Green | Vulkan/ROCm paths exist; real-time on 16 cores + 8060S. |
| ComfyUI / Stable Diffusion / SDXL | Yellow | ROCm + DirectML back-ends work; some custom nodes are CUDA-only and need alternatives. |
| vLLM, TensorRT-LLM, high-throughput serving | Yellow | vLLM has ROCm support but lags CUDA in feature velocity; see StorageReview July 2026 note in next section. |
| FlashAttention-2 (CUDA build), custom CUDA kernels, Triton-CUDA | Red | Needs ROCm port, HIP translation, or framework fallback; not drop-in. Choose DGX Spark if this is your daily stack. |
| XDNA2 NPU local acceleration | Yellow | ~50 TOPS for OS features and emerging Windows ML; not the LLM decode engine today. |
If you are unsure where your workflow lands, bring a requirements list. D-Central stocks both stacks so we can say no when Halo is not the right answer – we would rather sell you the DGX Spark at $9,449 CAD than have you fight a port you did not budget for. See our Spark vs Halo comparison and local hardware guide.
How Fast Is Strix Halo? Tokens Per Second Without Inventing Numbers
We do not invent tok/s. Any site publishing a clean “70B at 45 tok/s on Strix Halo” without a dated source, model, quant, prompt length, and sampler is guessing. Here is what we can responsibly say on 24 August 2026.
Third-party data point: StorageReview, July 2026, DGX Spark review found that on vLLM-optimised serving, Spark held a 2–4× throughput lead over Strix Halo-class Minisforum/Framework Desktop implementations on long-context inference, while noting Halo’s x86/Windows advantage for mixed use. That delta reflects CUDA kernel maturity and vLLM optimisation history, not just silicon. The same review noted llama.cpp/Vulkan on Halo closed the gap for single-user chat when the model fit in unified memory, but tok/s varied by quant and context.
Second data point: AMD positions the 8060S as RDNA 3.5 with 40 CU and 256 GB/s unified bandwidth, which governs token generation once resident. Discrete cards with 500-1,000 GB/s GDDR6 will generally generate faster if the model fits in VRAM; Halo wins when the model does not fit at all – avoiding PCIe spill matters more than bandwidth alone.
What D-Central is doing: our lab is running a reproducible harness for the Strix Halo 128GB at $7,349 CAD covering Llama 3.1 70B (Q4_K_M, Q6_K, Q8_0), Qwen3 32B, and Gemma 3 27B at 4k/32k/128k on Ollama and llama.cpp (Vulkan) under Windows 11 and Ubuntu 24.04, plus ROCm PyTorch. We will publish tok/s with prompt length, context, sampler, power limit, and UMA carve. Until then we are not quoting house numbers. If a decision depends on tok/s, ask for the run log or wait for the methodology post. See also the model database and compatibility notes.
Credits: AMD (8060S/RDNA 3.5 disclosure), Ollama and llama.cpp teams (Vulkan builds), StorageReview (July 2026 DGX Spark vs Halo-class comparison). No D-Central tok/s figures are stated in this section by design.
Gaming + AI: The Moat Spark Has No Answer For
No one buys Halo purely to game, but that it can is why it replaces a second PC. The 8060S at 40 CU RDNA 3.5 lands near an RTX 4060 Laptop GPU in rasterised titles at 1080p and playable 1440p, before ray tracing – well above any previous iGPU and above DGX Spark, which has no gaming story (Arm, no Windows, no GeForce, no anti-cheat). You buy Spark to run CUDA; you buy Halo to run everything else plus AI after hours.
What that means in practice on the Halo box we sell:
- 1080p high/ultra: eSports and most AAA titles at 60-144 fps without upscaling; 100+ fps with FSR in supported games.
- 1440p medium/high: 60 fps in many AAA titles, higher with FSR. Native 4K AAA at ultra is not the target – this is a 40 CU integrated part, not an RTX 4090.
- Creator apps by day, LLMs by night: Premiere, DaVinci, Lightroom and SolidWorks run natively on x86 Windows; after 18:00 the same box loads a 70B Q4 into ~96GB VRAM and serves a local assistant over the LAN.
- One-box apartment setup: No tower, no 600W PSU, no riser, no second GPU. See the thermals section for why that matters in February as much as August.
Honesty note: ray tracing still favours NVIDIA, and some CUDA creator plugins (DaVinci Neural Engine paths, Topaz, Blender CUDA) have slower ROCm/CPU fallbacks. If your income depends on one, validate on ROCm first or budget for Spark as a second box. For everyone else, a single Windows machine that games, edits, and hosts a 70B model offline is the moat Spark cannot match. The Strix Halo at $7,349 CAD is the simplest way to get it on a Canadian warranty – and if pure CUDA throughput is the mandate, the cross-sell is DGX Spark at $9,449 CAD via our side-by-side.
Thermostics, Noise & Power: 120W Desk Box in a Canadian Apartment
Halo idles under 15W, sips 35-55W in desktop use, and tops around 120W at the wall under sustained load (cTDP-limited) – roughly half of DGX Spark’s ~240W envelope and one-fifth of a 600W tower. In a 500 sq ft Montreal apartment in January that 120W is a quiet heater; in July you can still run it without tripping a 15A circuit. Spark is quiet for its class but audible at sustained load; a 450-600W tower is materially louder and hotter.
Noise: our 120W Halo build is effectively inaudible at idle, ~28-33 dBA at desk at sustained load – conversation level, not vacuum level. Spark’s blower is audible under serve load. Thermals: both Halo and Spark are designed for sustained inference, but only Halo is designed to sit on the same desk as a gaming monitor and run at 1440p without a second power brick.
Power costs – because Hydro bills are not abstract in this country:
| Assumption: 8 hrs/day @ 100% load, 30 days | kWh / month | QC – Hydro-Québec ~$0.111/kWh* | ON – ~$0.15/kWh* (TOU blended) | AB – ~$0.18/kWh* |
|---|---|---|---|---|
| Strix Halo 120W | 28.8 | ~$3.20 | ~$4.32 | ~$5.18 |
| DGX Spark ~240W | 57.6 | ~$6.39 | ~$8.64 | ~$10.37 |
| Gaming tower ~600W | 144.0 | ~$15.98 | ~$21.60 | ~$25.92 |
*Blended 2026 estimates incl. energy + delivery; tariff governs. Proportion: Halo costs half of Spark and one-fifth of a tower to leave on all day. In Quebec that 120W offsets furnace runtime for months, and local inference keeps data residency simple under Law 25 – personal information never leaves the premises. See the Strix Halo at $7,349 CAD for power/thermal specs and our local AI hardware guide for placement.
Petite note bilingue pour le Québec : service en français ou en anglais à Laval, facturation bilingue, et la page produit est disponible en français sur demande – on vous répond dans votre langue, pas en traduction automatique.
DGX Spark vs Strix Halo: $9,449 vs $7,349 — The $2,000 Question
These are the only two turnkey 128GB unified-memory desks worth stocking in Canada in 2026. Both hold ~96GB VRAM; both handle 70B Q4 natively. The $2,100 difference – DGX Spark $9,449 CAD vs Strix Halo $7,349 CAD – is not about capacity. It is about software stack, operating system, and power. The table below is the buying filter we use in-store. Full detail lives on DGX Spark vs Strix Halo and which local AI computer if you want the 60-second version.
| Dimension | AMD Strix Halo 128GB – $7,349 CAD | NVIDIA DGX Spark – $9,449 CAD |
|---|---|---|
| Architecture | 16-core Zen 5 x86 + Radeon 8060S 40 CU RDNA 3.5 (single die) | 20-core Arm Grace + Blackwell GPU (GB10) |
| Memory / Bandwidth | 128GB LPDDR5X-8000, 256-bit, 256 GB/s (up to ~96GB VRAM) | 128GB LPDDR5X, 256-bit, 273 GB/s + 4TB NVMe on-board |
| Software stack | ROCm 6.x / Vulkan / DirectML – Ollama, llama.cpp excellent | CUDA – widest kernel coverage, 2–4× vLLM lead per StorageReview July 2026 |
| Operating system | Windows 11 Pro or Linux – Adobe, Steam, Office native | DGX OS (Ubuntu) – Linux/CUDA appliance, no native Windows apps |
| Power / thermals | cTDP 45-120W; ~120W wall; quiet; apartment-friendly | ~240W system / 140W SoC; audible under serve load; also quiet for class |
| Gaming / creator | Yes – ~4060 Laptop class at 1080p/1440p, FSR, native x86 apps | No gaming/creator Windows story – pure AI box |
| Ideal buyer | One-box Windows user: dev, professor, studio, SOC wanting local 70B + daily PC | CUDA shop needing maximum tokens/sec and largest-batch serving |
Verdict: Choose Strix Halo if you want one x86 Windows/Linux box that handles documents, Adobe/Steam, and a 70B Q4 in ~96GB VRAM for $7,349 CAD. Choose DGX Spark if your team ships CUDA kernels, lives in vLLM/TensorRT, or needs Blackwell optimisation for maximum throughput at $9,449 CAD. Not sure? Full comparison → or 60-second chooser →
On pricing: Spark’s US MSRP is ~$3,999-$4,699 USD. Add exchange, brokerage, duties and early US demand, and landed cost in Canada is often within a few hundred dollars of our $9,449 CAD with a Canadian invoice and Montreal warranty. Same for Halo: Framework Desktop previewed is a US pre-order for Q4 2026, while our Strix Halo at $7,349 CAD is a Minisforum/Framework-class equivalent we source and support here today – a timeline note, not an anti-Framework one.
Where to Buy Strix Halo in Canada Without a US Import Surprise
What burns Canadians every launch: a US review says “now at Micro Center for $X USD” and the button is store-only in Ohio or a US pre-order. Through a forwarder you pay exchange plus $80-$150 brokerage plus HST on landed value plus a US-only warranty if it fails. Framework Desktop with Ryzen AI Max+ 395 is legitimate, but as of 24 August 2026 it remains a US pre-order for Q4 2026, not a stocked Canadian SKU. Micro Center G15/Z19-class Halo boxes are US exclusives with no Canadian channel.
D-Central’s position is boring on purpose: we stock a Framework/Minisforum-class Strix Halo 128GB at $7,349 CAD in CAD, HST at checkout, no duties or brokerage. Boxes ship from Laval, QC, with a 1-year D-Central warranty handled in Montreal, English or French, with in-store setup if you want to walk out with Ollama running. Prefer to validate? Book a demo – we will load your model, set the UMA carve, and show the same box running a 70B Q4 then launching Steam without rebooting.
What you get: 120W Strix Halo with Ryzen AI Max+ 395, 128GB LPDDR5X-8000 (soldered), Wi-Fi 6E/7 and 2.5GbE per bundle, HDMI/DP and USB4, Canadian PSU. We can pre-install Windows 11 Pro with Ollama + Vulkan and a starter library (Llama 3.1 8B + Qwen3 32B Q4), or ship bare with Linux for ROCm. Either way, no US pre-order wait and no self-import. If you need CUDA, the parallel SKU is DGX Spark at $9,449 CAD; indecision page is which local AI computer.
Sourcing note: Strix Halo APU is AMD silicon; our chassis is sourced via AMD’s Canadian distribution as a complete system in the Framework/Minisforum equivalence class – hence a 120W thermal target and BIOS with UMA tuning, not an unbranded board. Prices are current to 24 August 2026 and move with CAD/USD; product page governs. For broader context, see local AI hardware guide and open-weight comparison.
Frequently asked questions
Is Strix Halo upgradeable?
No. On Ryzen AI Max+ 395 the 128GB LPDDR5X-8000 is soldered to the board to achieve the 256-bit / 256 GB/s interface. You cannot add SO-DIMMs later, which is why the 128GB ceiling is fixed at purchase. Storage (NVMe) is upgradeable and the chassis exposes USB4 and standard display outputs, but system RAM is not. If you think you will need more than ~96GB VRAM within two years, consider whether a tower with discrete VRAM or a second appliance makes more sense, or speak to us about future 192GB Halo-class parts – none are stocked in Canada as of 24 August 2026. The box we sell at $7,349 CAD is sold strictly as 128GB unified for this reason.
Can it run CUDA?
No – CUDA is NVIDIA-only. Strix Halo’s Radeon 8060S runs ROCm 6.x on Linux, and Vulkan/DirectML on Windows. Ollama and llama.cpp both publish Vulkan paths that run well on RDNA 3.5, and PyTorch has ROCm wheels. If your workload is CUDA-native – TensorRT, certain FlashAttention-2 builds, or custom kernels a lab shipped only for NVIDIA – you need a ROCm port, a Vulkan fallback, or a different system. That different system is DGX Spark at $9,449 CAD, which is native CUDA/Blackwell. We stock both so you do not have to port on faith.
How much VRAM does 8060S get?
Up to ~96GB out of the 128GB unified pool. The BIOS exposes a UMA frame-buffer carve; our default is 96GB graphics / 32GB system, which you can adjust. That 96GB is not GDDR6 – it is LPDDR5X-8000 at 256 GB/s – but it is directly addressable by the GPU without PCIe copy, so a 70B Q4 (~41GB) or even a 70B Q8 (~73GB) sits comfortably. The 8060S itself is 40 CU RDNA 3.5 with hardware ray tracing and AV1 encode/decode. For model sizing, use our model database or the table in the memory section above.
Is Strix Halo available in Canada?
Yes – from D-Central in Montreal/Laval at $7,349 CAD, in-stock Canada with a Canadian invoice, HST, and no brokerage or duties. The SKUs you may have seen previewed in the US – Micro Center-exclusive boards and the Framework Desktop Halo – are either US-only retail or US pre-order for Q4 2026 as of 24 August 2026. Our chassis is a Framework Desktop-class / Minisforum-class build through AMD’s Canadian channel, fully supported here. If you want to compare availability and warranty against the CUDA alternative, see DGX Spark at $9,449 CAD and our DGX Spark Canada guide.
How is it different from DGX Spark?
Three differences above capacity: operating system, software stack, and power. Strix Halo is x86 Windows 11 or Linux, runs ROCm/Vulkan, and lives at ~120W – it is a workstation that does AI. DGX Spark is Arm DGX OS (Ubuntu), runs CUDA on a Grace Blackwell GB10, and lives at ~240W – it is an AI appliance that pairs with a separate PC. Both have 128GB unified memory and can host a 70B Q4 in ~96GB VRAM. Halo costs $7,349 CAD; Spark costs $9,449 CAD at D-Central. The long answer lives on DGX Spark vs Strix Halo and the quick filter on which computer to buy.
Does it game?
Yes – credibly. The 8060S at 40 CU RDNA 3.5 with 96GB available VRAM games like an RTX 4060 Laptop-class GPU at 1080p high/ultra (60-144 fps) and at 1440p medium/high (often 60 fps, higher with FSR). It handles eSports, AAA rasterised titles, and creator apps the same way a normal Windows PC does, because it is one. Spark, by design, does not game. If you want one box in a small apartment that trains, infers, edits, and games, this is the trade Halo makes deliberately. For hardware context beyond Halo, see the local hardware guide and Halo vs custom build.
What models fit?
With a 96GB VRAM carve: comfortably up to 70B Q8 (~73GB). A 70B Q4 (~41GB), a 32B Q4 (~18GB) and a 27B Q8 together fit with context headroom. FP16 70B (~140GB) does not fit without offload; 405B Q4 (~230GB+) does not fit at all. Quantisation is the lever: Q4_K_M saves ~45% vs Q8 with modest quality loss on large models, and our tables err conservative by adding KV-cache margin. Current fit guidance is maintained on GPU/LLM compatibility and scrolled daily on the model database. If you need a one-sentence rule: Strix Halo is a 70B-native box at $7,349 CAD.
Windows or Linux?
Both are supported and both run LLMs. Choose Windows 11 if you want Adobe, Office, Steam, and Ollama/llama.cpp on one machine – use the Vulkan builds for GPU inference and keep ROCm out of the picture. Choose Linux (Ubuntu 24.04 recommended) if you want ROCm 6.x breadth for PyTorch, vLLM, and ComfyUI – you get more frameworks native, but you lose native Windows apps. Either way, the 128GB unified pool and 256 GB/s layout are the same, and BIOS UMA tuning works on both. We can ship either OS pre-configured from Laval; state your preference at checkout on the Strix Halo 128GB page or bring your workload to the chooser if you are torn – French or English.
Check Strix Halo 128GB Availability – $7,349 CAD → · Need CUDA? DGX Spark $9,449 → · Compare Side-by-Side →
Related products, repair, and setup paths
- how D-Central diagnoses ASIC repairs
- ASIC troubleshooting library
- ASIC manuals and repair guides
- replacement hashboards
- ASIC control boards
- ASIC power supplies
- S19 family replacement hashboard
- C52 replacement control board
- APW12 S19 power supply
- immersion cooling hub
- home immersion cooling guide
- ASIC miners for immersion planning
- ASIC cooling parts
- airflow shroud before immersion
- compare miner specs in the database
- ASIC repair support
- compare ASIC miner specs
- ASIC miner database
- Antminer S19 specs and profitability
- buy a tested Antminer S19
- Antminer S19 maintenance guide
- Antminer S19 repair service
- Antminer S21 specs
- Bitmain Antminer S21
- Antminer S21 maintenance guide
- BM1370BC S21 Pro chip
- Antminer S9 specs
- Bitmain Antminer S9
- Antminer S9 maintenance guide
- S9 hashboard repair parts bundle
Last reviewed August 24, 2026.
