Enterprise model shortlist / reviewed August 24, 2026
Open weights are now near the frontier. Deployment is the hard part.
A current Canadian buyer’s guide to capability, licences, infrastructure class, origin, and where each leading open-weight model actually fits.
How close are open-weight models to the frontier? On August 24, 2026, Artificial Analysis Intelligence Index v4.1.1 scores the overall leader at 63 and Kimi K3, the leading open-weight model, at 60. Qwen3.8-2.4T scores 58; DeepSeek V4 Pro and GLM-5.2 each score 53. This is strong evidence of near-frontier aggregate capability—not proof that every open model matches every proprietary model on every legal, medical, coding, multilingual, safety, latency, or production task.
Read the date, index version, and licence before the score
Model leaderboards decay quickly. Artificial Analysis v4.1.1 combines nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. A composite score is useful for shortlisting; a production decision needs representative customer tasks, the exact downloadable checkpoint, serving configuration, and a licence the intended use can satisfy.
“Open weight” means the trained parameters can be obtained. It does not mean the model is open source, unrestricted, cheap to run, safe for confidential data, or legal to offer as a hosted service. MIT and Apache 2.0 checkpoints offer the cleanest commercial baseline in this list. Custom licences may add hosted-inference, revenue, branding, attribution, or acceptable-use conditions.
Current enterprise-relevant shortlist
| Model | AA v4.1.1 | Weights / licence | Infrastructure class | Best fit | Key caveat |
|---|---|---|---|---|---|
| Kimi K3 Moonshot AI, China |
60 | 2.8T total / 104B active Custom Kimi K3 licence |
Rack-scale; native checkpoint is roughly 1.6 TB | Frontier-class multimodal, long context, reasoning and coding | Hosted-service and revenue terms require review; major interconnect and runtime demands |
| Qwen3.8-2.4T-A95B Alibaba, China |
58 | 2.4T / 95B active Custom Qwen3.8-Max licence |
Rack-scale; 8×B300 class for a 4-bit recipe, more for higher precision | Text reasoning, coding, professional work and agents | Downloadable checkpoint is text-only; hosted “Max” features do not automatically transfer |
| DeepSeek V4 Pro 0813 DeepSeek, China |
53 | 1.6T / 49B active MIT |
Large enterprise node; official example uses 4×GB300 | Long-context coding, reasoning and general agent work | Very new; dedicated encoding/integration path needs production validation |
| GLM-5.2 Z.ai, China |
53 | 753B / 40B active MIT |
Large enterprise node; FP8 recipe uses 8×H200/H20 | Long-horizon tasks, coding and configurable reasoning | Long context and concurrency add memory far beyond weight storage |
| Qwen3.8-27B Alibaba, China |
52 | 27B dense Apache 2.0 |
Single high-memory GPU; roughly 52 GB BF16, around 25–32 GB packaged NVFP4 class | Best capability-per-box candidate; text, image, video, coding and agents | Extremely new; test reasoning verbosity, latency, and real workload reliability |
| MiniMax-M3 MiniMax, China |
45 | 428B / 23B active Custom community licence |
Enterprise node; reference BF16 deployment uses 8×H200/H20 | Multimodal, 1M context, adaptive and direct modes | Attribution plus prior authorization above a stated revenue threshold |
| NVIDIA Nemotron 3 Ultra 550B A55B NVIDIA, U.S. |
38 | 550B / 55B active OpenMDW 1.1 |
4× current Blackwell or 8×H100 class for listed NVFP4 support | Reasoning, agents, multilingual RAG and tool use | Custom terms; Canadian hosting does not remove U.S. model, GPU, firmware, or runtime origins |
| Mistral Medium 3.5 Mistral AI, France |
30 | 128B dense Modified MIT |
About 130 GB at FP8; 2 Blackwell or 4 Hopper-class GPUs in serving guidance | Multimodal, 24 languages, coding and a European upstream option | Licence rights change above a stated global revenue threshold |
| gpt-oss-120b OpenAI, U.S. |
24 | 117B / 5.1B active Apache 2.0 |
One 80 GB GPU class; 20B sibling around 16 GB | Mature ecosystem, structured output, function calling, simpler local route | Text-only, 2024 knowledge cutoff, Harmony response format; U.S.-origin technology |
| Cohere Command A+ Canadian origin |
23 | 218B / 25B active Apache 2.0 |
One B200 or two H100 class at W4A4 | Enterprise RAG, citations, agents, vision, 48 languages, private deployment | Credible Canadian-first choice, but not the aggregate intelligence leader |
The practical deployment tiers
Qwen3.8-27B
The standout local appliance candidate in the current data: a score of 52, Apache 2.0 terms, multimodal input, and a footprint compatible with one current high-memory GPU when using an appropriate low-precision checkpoint.
Command A+ or gpt-oss-120b
Command A+ is the Canadian-first RAG and multilingual path; gpt-oss-120b is the straightforward Apache-licensed ecosystem path. Both are easier to operationalize than a trillion-parameter frontier cluster.
GLM-5.2 or DeepSeek V4 Pro
MIT licensing and scores of 53 make both compelling where the workload justifies multiple very high-memory GPUs and the organization accepts their model origins and newness.
Kimi K3 or Qwen3.8-2.4T
Scores of 60 and 58 respectively put open weights near the overall frontier. Their scale, custom licences, runtime maturity, and interconnect demands make them infrastructure programs, not “download and run” products.
Yes, there is a good Canadian model
Saying “Canada has no good model” is no longer accurate. Cohere was founded in Toronto, and Command A+ is a commercially usable Apache-2.0 checkpoint designed for private enterprise deployment. It supports RAG, citations, agents, vision, and 48 languages including French, with a comparatively attractive serving footprint.
The honest qualifier is equally important: Command A+’s current Artificial Analysis v4.1.1 score is 23, not the open-weight frontier. That can still be the right choice for a Canadian organization’s retrieval-heavy, multilingual, private workflow. A leaderboard score should not erase origin, licence, operational fit, citations, or the customer’s acceptance tests.
Canada does not need to wait for a single domestic model to achieve AI sovereignty. We can deploy Canadian-developed models where they fit, and operate the world’s best commercially usable open weights on Canadian infrastructure where frontier capability is required.
Licence traps for a hosted inference provider
D-Central’s intended enterprise hosting use makes licence discipline non-negotiable. The following is a technical procurement summary, not legal advice:
- MIT and Apache 2.0: the cleanest starting point for commercial internal deployment and hosted-service evaluation, subject to their notices and the customer’s use.
- Kimi K3: custom terms address model-as-a-service revenue and large-product naming or UI obligations.
- Qwen3.8 flagship: custom terms address certain MaaS and AI-work-assistant revenue and branding thresholds.
- MiniMax-M3: custom terms include notice/attribution and prior authorization above a stated annual revenue level.
- Mistral Medium 3.5: modified terms change rights above a stated consolidated monthly revenue level.
- NVIDIA Nemotron: OpenMDW is a custom model-development-weights licence with acceptable-use obligations.
Always preserve and review the exact licence shipped with the exact checkpoint. A blog post cannot grant deployment rights.
Model origin and inference location are different controls
Hosting Qwen, DeepSeek, Kimi, Mistral, Cohere, Nemotron, or gpt-oss on a Canadian machine means the prompts and execution can remain inside the Canadian operating boundary. It does not change where the weights were developed or eliminate dependencies on foreign GPUs, drivers, firmware, runtimes, upstream security updates, or licence holders.
Some regulated or public-sector organizations extend supply-chain rules to model origin. Others focus on data flow, operator, governing law, auditability, and continuity. Record both instead of using “sovereign” as a shortcut.
How D-Central turns the shortlist into a decision
- Define the workload: examples, languages, modalities, tools, retrieval, output, latency, concurrency, and failure cost.
- Apply hard constraints: licence, model origin, hardware, network boundary, retention, and procurement rules.
- Benchmark two or three candidates: use the same authorized evaluation set and record the exact checkpoint and serving settings.
- Measure the system: end-to-end quality, retrieval, time to first answer, throughput, memory, power, heat, recovery, and operator effort.
- Keep a replacement: preserve the evaluation and a second compatible model so today’s winner does not become tomorrow’s lock-in.
Scope a Canadian model evaluation →
Frequently asked questions
What is the best open-weight model right now?
By Artificial Analysis Intelligence Index v4.1.1 on August 24, 2026, Kimi K3 leads open weights at 60. “Best” for a deployment can instead be Qwen3.8-27B for a single GPU, Command A+ for Canadian-origin RAG, or a permissively licensed large model for a dedicated cluster.
Are open-weight models as good as frontier proprietary models?
The leading open-weight aggregate score is within three points of the overall leader on the current index. That supports “near frontier,” not universal parity. Task-level quality, language, modalities, tools, safety, latency, and reliability still vary.
Are all models in this table open source?
No. They all offer downloadable weights, but several use custom licences. Use “open weight” as the umbrella term and reserve “open source” for software or checkpoints whose actual licence supports that statement.
Can a Canadian company commercially host these models?
It depends on the exact checkpoint, licence, customer use, revenue, branding, and service design. MIT and Apache 2.0 are the cleanest initial candidates. Custom-licence models need qualified review before a hosted offering.
Will the scores stay current?
No. Leaderboards, model versions, index composition, and estimates change. This comparison is dated August 24, 2026 and names v4.1.1. Follow the live leaderboard and rerun customer evaluations before procurement.
Primary model sources: Artificial Analysis live comparison; Kimi K3; Qwen3.8-2.4T; Qwen3.8-27B; DeepSeek V4 Pro; GLM-5.2; Cohere Command A+; gpt-oss-120b.
Live Canadian SKUs (24 August 2026 prices): NVIDIA DGX Spark — $9,449 CAD · AMD Strix Halo 128 GB — $7,349 CAD · custom GPU inference rig (quote). Retired vapour names (Pleb AI Box, Workstation 24/48, Hashcenter AI Node 80+) are not for sale.
Related products, repair, and setup paths
- self-hosted AI for Bitcoiners hub
- plebs guide to self-hosted AI
- install Ollama in 10 minutes
- LM Studio vs Ollama vs llama.cpp
- connect local AI to Home Assistant and Obsidian
- self-hosted AI troubleshooting
- repurpose mining hardware into an AI hashcenter
- local AI model leaderboards
Last reviewed August 24, 2026.