Skip to content

Enterprise model shortlist / reviewed August 24, 2026

Open weights are now near the frontier. Deployment is the hard part.

A current Canadian buyer’s guide to capability, licences, infrastructure class, origin, and where each leading open-weight model actually fits.

How close are open-weight models to the frontier? On August 24, 2026, Artificial Analysis Intelligence Index v4.1.1 scores the overall leader at 63 and Kimi K3, the leading open-weight model, at 60. Qwen3.8-2.4T scores 58; DeepSeek V4 Pro and GLM-5.2 each score 53. This is strong evidence of near-frontier aggregate capability—not proof that every open model matches every proprietary model on every legal, medical, coding, multilingual, safety, latency, or production task.

Read the date, index version, and licence before the score

Model leaderboards decay quickly. Artificial Analysis v4.1.1 combines nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. A composite score is useful for shortlisting; a production decision needs representative customer tasks, the exact downloadable checkpoint, serving configuration, and a licence the intended use can satisfy.

“Open weight” means the trained parameters can be obtained. It does not mean the model is open source, unrestricted, cheap to run, safe for confidential data, or legal to offer as a hosted service. MIT and Apache 2.0 checkpoints offer the cleanest commercial baseline in this list. Custom licences may add hosted-inference, revenue, branding, attribution, or acceptable-use conditions.

Current enterprise-relevant shortlist

Model AA v4.1.1 Weights / licence Infrastructure class Best fit Key caveat
Kimi K3
Moonshot AI, China
60 2.8T total / 104B active
Custom Kimi K3 licence
Rack-scale; native checkpoint is roughly 1.6 TB Frontier-class multimodal, long context, reasoning and coding Hosted-service and revenue terms require review; major interconnect and runtime demands
Qwen3.8-2.4T-A95B
Alibaba, China
58 2.4T / 95B active
Custom Qwen3.8-Max licence
Rack-scale; 8×B300 class for a 4-bit recipe, more for higher precision Text reasoning, coding, professional work and agents Downloadable checkpoint is text-only; hosted “Max” features do not automatically transfer
DeepSeek V4 Pro 0813
DeepSeek, China
53 1.6T / 49B active
MIT
Large enterprise node; official example uses 4×GB300 Long-context coding, reasoning and general agent work Very new; dedicated encoding/integration path needs production validation
GLM-5.2
Z.ai, China
53 753B / 40B active
MIT
Large enterprise node; FP8 recipe uses 8×H200/H20 Long-horizon tasks, coding and configurable reasoning Long context and concurrency add memory far beyond weight storage
Qwen3.8-27B
Alibaba, China
52 27B dense
Apache 2.0
Single high-memory GPU; roughly 52 GB BF16, around 25–32 GB packaged NVFP4 class Best capability-per-box candidate; text, image, video, coding and agents Extremely new; test reasoning verbosity, latency, and real workload reliability
MiniMax-M3
MiniMax, China
45 428B / 23B active
Custom community licence
Enterprise node; reference BF16 deployment uses 8×H200/H20 Multimodal, 1M context, adaptive and direct modes Attribution plus prior authorization above a stated revenue threshold
NVIDIA Nemotron 3 Ultra 550B A55B
NVIDIA, U.S.
38 550B / 55B active
OpenMDW 1.1
4× current Blackwell or 8×H100 class for listed NVFP4 support Reasoning, agents, multilingual RAG and tool use Custom terms; Canadian hosting does not remove U.S. model, GPU, firmware, or runtime origins
Mistral Medium 3.5
Mistral AI, France
30 128B dense
Modified MIT
About 130 GB at FP8; 2 Blackwell or 4 Hopper-class GPUs in serving guidance Multimodal, 24 languages, coding and a European upstream option Licence rights change above a stated global revenue threshold
gpt-oss-120b
OpenAI, U.S.
24 117B / 5.1B active
Apache 2.0
One 80 GB GPU class; 20B sibling around 16 GB Mature ecosystem, structured output, function calling, simpler local route Text-only, 2024 knowledge cutoff, Harmony response format; U.S.-origin technology
Cohere Command A+
Canadian origin
23 218B / 25B active
Apache 2.0
One B200 or two H100 class at W4A4 Enterprise RAG, citations, agents, vision, 48 languages, private deployment Credible Canadian-first choice, but not the aggregate intelligence leader

The practical deployment tiers

Single GPU

Qwen3.8-27B

The standout local appliance candidate in the current data: a score of 52, Apache 2.0 terms, multimodal input, and a footprint compatible with one current high-memory GPU when using an appropriate low-precision checkpoint.

Private enterprise

Command A+ or gpt-oss-120b

Command A+ is the Canadian-first RAG and multilingual path; gpt-oss-120b is the straightforward Apache-licensed ecosystem path. Both are easier to operationalize than a trillion-parameter frontier cluster.

Large dedicated node

GLM-5.2 or DeepSeek V4 Pro

MIT licensing and scores of 53 make both compelling where the workload justifies multiple very high-memory GPUs and the organization accepts their model origins and newness.

Frontier cluster

Kimi K3 or Qwen3.8-2.4T

Scores of 60 and 58 respectively put open weights near the overall frontier. Their scale, custom licences, runtime maturity, and interconnect demands make them infrastructure programs, not “download and run” products.

Yes, there is a good Canadian model

Saying “Canada has no good model” is no longer accurate. Cohere was founded in Toronto, and Command A+ is a commercially usable Apache-2.0 checkpoint designed for private enterprise deployment. It supports RAG, citations, agents, vision, and 48 languages including French, with a comparatively attractive serving footprint.

The honest qualifier is equally important: Command A+’s current Artificial Analysis v4.1.1 score is 23, not the open-weight frontier. That can still be the right choice for a Canadian organization’s retrieval-heavy, multilingual, private workflow. A leaderboard score should not erase origin, licence, operational fit, citations, or the customer’s acceptance tests.

Canada does not need to wait for a single domestic model to achieve AI sovereignty. We can deploy Canadian-developed models where they fit, and operate the world’s best commercially usable open weights on Canadian infrastructure where frontier capability is required.

Licence traps for a hosted inference provider

D-Central’s intended enterprise hosting use makes licence discipline non-negotiable. The following is a technical procurement summary, not legal advice:

  • MIT and Apache 2.0: the cleanest starting point for commercial internal deployment and hosted-service evaluation, subject to their notices and the customer’s use.
  • Kimi K3: custom terms address model-as-a-service revenue and large-product naming or UI obligations.
  • Qwen3.8 flagship: custom terms address certain MaaS and AI-work-assistant revenue and branding thresholds.
  • MiniMax-M3: custom terms include notice/attribution and prior authorization above a stated annual revenue level.
  • Mistral Medium 3.5: modified terms change rights above a stated consolidated monthly revenue level.
  • NVIDIA Nemotron: OpenMDW is a custom model-development-weights licence with acceptable-use obligations.

Always preserve and review the exact licence shipped with the exact checkpoint. A blog post cannot grant deployment rights.

Model origin and inference location are different controls

Hosting Qwen, DeepSeek, Kimi, Mistral, Cohere, Nemotron, or gpt-oss on a Canadian machine means the prompts and execution can remain inside the Canadian operating boundary. It does not change where the weights were developed or eliminate dependencies on foreign GPUs, drivers, firmware, runtimes, upstream security updates, or licence holders.

Some regulated or public-sector organizations extend supply-chain rules to model origin. Others focus on data flow, operator, governing law, auditability, and continuity. Record both instead of using “sovereign” as a shortcut.

How D-Central turns the shortlist into a decision

  1. Define the workload: examples, languages, modalities, tools, retrieval, output, latency, concurrency, and failure cost.
  2. Apply hard constraints: licence, model origin, hardware, network boundary, retention, and procurement rules.
  3. Benchmark two or three candidates: use the same authorized evaluation set and record the exact checkpoint and serving settings.
  4. Measure the system: end-to-end quality, retrieval, time to first answer, throughput, memory, power, heat, recovery, and operator effort.
  5. Keep a replacement: preserve the evaluation and a second compatible model so today’s winner does not become tomorrow’s lock-in.

Scope a Canadian model evaluation →

Frequently asked questions

What is the best open-weight model right now?

By Artificial Analysis Intelligence Index v4.1.1 on August 24, 2026, Kimi K3 leads open weights at 60. “Best” for a deployment can instead be Qwen3.8-27B for a single GPU, Command A+ for Canadian-origin RAG, or a permissively licensed large model for a dedicated cluster.

Are open-weight models as good as frontier proprietary models?

The leading open-weight aggregate score is within three points of the overall leader on the current index. That supports “near frontier,” not universal parity. Task-level quality, language, modalities, tools, safety, latency, and reliability still vary.

Are all models in this table open source?

No. They all offer downloadable weights, but several use custom licences. Use “open weight” as the umbrella term and reserve “open source” for software or checkpoints whose actual licence supports that statement.

Can a Canadian company commercially host these models?

It depends on the exact checkpoint, licence, customer use, revenue, branding, and service design. MIT and Apache 2.0 are the cleanest initial candidates. Custom-licence models need qualified review before a hosted offering.

Will the scores stay current?

No. Leaderboards, model versions, index composition, and estimates change. This comparison is dated August 24, 2026 and names v4.1.1. Follow the live leaderboard and rerun customer evaluations before procurement.

Primary model sources: Artificial Analysis live comparison; Kimi K3; Qwen3.8-2.4T; Qwen3.8-27B; DeepSeek V4 Pro; GLM-5.2; Cohere Command A+; gpt-oss-120b.

Live Canadian SKUs (24 August 2026 prices): NVIDIA DGX Spark — $9,449 CAD · AMD Strix Halo 128 GB — $7,349 CAD · custom GPU inference rig (quote). Retired vapour names (Pleb AI Box, Workstation 24/48, Hashcenter AI Node 80+) are not for sale.