Skip to content
Small team, full backlog, zero orders dropped. Support replies are slower than we’d like. Read our status update → Zero orders dropped. Status → 📬 Check your spam folder — most of our replies land there. We do answer. Status update → 📬 Check your spam folder. Status →

Enterprise design partners / Canadian-operated inference

Your private endpoint should not be a foreign off-switch.

D-Central is accepting scoped conversations with Canadian organizations that need open-weight inference operated in Canada without owning the entire GPU stack themselves.

  • Canadian location and operation
  • Dedicated options by scope
  • Open-weight portability
  • Written data boundary
Offer status: design-partner engagements. This is not a public, self-serve cloud or a claim of unlimited ready capacity. D-Central will validate hardware, facility fit, model licence, isolation, security, data handling, support, availability, and commercial terms before accepting an engagement. No capacity, uptime, certification, or service level exists unless it is written into the customer’s agreement.

What is Canadian-hosted AI inference? An AI model runs on computing infrastructure physically located and operationally managed in Canada, and your team accesses it through a private interface or API. A defensible service also documents administrators, subprocessors, retention, backups, governing terms, model licence, portability, and exit. D-Central is exploring dedicated and isolated deployments for enterprise customers that want this middle path between a U.S.-controlled API and a server in their own office.

Who this path is for

You need Canadian operation

Your procurement, customer, privacy, or resilience requirements call for an inference path operated here, with documented Canadian data flows and accountable human operators.

You need more than a workstation

The workload has sustained throughput, multiple teams, larger models, long context, or availability requirements that justify dedicated server or rack-scale infrastructure.

You do not want to operate GPUs

Your team wants a stable private endpoint and model portability without owning power, cooling, drivers, serving, monitoring, and hardware lifecycle work.

What can be scoped

A hosted engagement starts as architecture and feasibility. Depending on the workload and the written statement of work, the design can address:

  • Dedicated or logically isolated serving appropriate to the customer and model.
  • Private network access or authenticated API, including an OpenAI-compatible interface where the selected runtime supports it.
  • Open-weight model serving selected for quality, language, licence, hardware fit, and portability—not for a logo.
  • Private RAG with a defined location and lifecycle for source documents, embeddings, vector storage, retrieved passages, and citations.
  • Retention and observability controls that separate what must be measured from what must not be stored.
  • Model evaluation and update gates so a new checkpoint does not silently replace a known production behaviour.
  • Capacity and recovery planning tied to measured concurrency, context, latency, maintenance, and failure expectations.
  • An exit package covering customer data export, configuration, evaluation cases, and a route to another compatible runtime or an on-premises build.

The questions we put in writing

Area Decision before production What we refuse to imply
Capacity Model, precision, context, request profile, concurrency, latency, growth That a parameter count predicts throughput
Isolation Dedicated hardware, process boundaries, storage, network, administration That “private endpoint” explains the tenant boundary
Data Ingress, retention, logging, retrieval, backups, support access, deletion That “no training” means no storage
Licence Commercial use, hosted inference, attribution, revenue thresholds, acceptable use That downloadable weights are automatically open source
Availability Maintenance, redundancy, recovery, support window, remedies An SLA that was never engineered or contracted
Exit Portability, data/config export, replacement model, cutover and deletion evidence That a Canadian vendor should become a new lock-in

Model classes for a Canadian service

The cleanest hosted launch candidates are models with permissive MIT or Apache 2.0 terms. As of August 24, 2026:

  • Cohere Command A+ is Canadian-origin, Apache 2.0, multilingual and enterprise-oriented, with RAG, citations, agents, vision, and private deployment as core use cases.
  • Qwen3.8-27B is Apache 2.0 and unusually capable for a single high-memory GPU class system, scoring 52 on Artificial Analysis v4.1.1.
  • DeepSeek V4 Pro and GLM-5.2 use MIT terms and score 53, but require a much larger enterprise infrastructure class.
  • gpt-oss-120b carries Apache 2.0 and has a mature deployment ecosystem, but it is U.S.-origin technology even when its inference runs entirely in Canada.

Leading models such as Kimi K3, the flagship Qwen3.8-2.4T, MiniMax-M3, Mistral Medium 3.5, and NVIDIA Nemotron use custom licences with hosted-service, revenue, branding, or acceptable-use terms that require explicit review. “Open weight” is not a commercial blank cheque.

See the dated model and licence comparison →

Canadian-hosted, on-premises, or hybrid?

Path Best when Trade-off
Canadian-hosted You need Canadian operation and enterprise capacity without running the hardware You still depend on an operator; contract, transparency, and exit matter
On premises Maximum organizational control and predictable local workloads matter most You own patching, recovery, security, power, cooling, and capacity
Hybrid Different workloads justify local, Canadian-hosted, and frontier API routes Routing and governance become a real engineering system

Apply for a scoped Canadian inference engagement

Useful first details: your workload, data sensitivity, preferred model or capability, expected users and concurrency, context size, monthly or hourly volume, network boundary, availability target, and desired start date. If the requirement is better served on premises—or is not a fit for current D-Central capacity—we will say so.

Frequently asked questions

Is the service generally available today?

No public self-serve service or standard capacity pool is being advertised here. D-Central is accepting design-partner enquiries and will confirm feasibility, scope, capacity, security, support, and commercial terms before any engagement.

Will our prompts be used to train a model?

The intended private-service posture is no customer-data training, but the complete data-handling terms—including logging, retention, troubleshooting access, backups, and deletion—must be explicit in the final agreement. Do not rely on one slogan.

Can you expose an OpenAI-compatible API?

Many open serving runtimes support OpenAI-compatible endpoints, which can reduce application changes. Compatibility is not perfect across every model, tool call, structured-output feature, embedding route, or streaming behaviour, so we test the customer’s actual integration.

Can we choose the model?

Yes, within capability, hardware, security, and licence constraints. We recommend carrying an evaluation set and a replacement option so the service does not turn one provider dependency into a model dependency.

Does running in Canada eliminate every foreign dependency?

No. GPUs, firmware, drivers, model origins, licences, and software supply chains can remain foreign. Canadian operation materially changes the data and control boundary; honest sovereignty work also documents the dependencies that remain.

Service principle: Canadian hosting should be a path out of concentration risk, not a new form of lock-in. Related: AI vendor exit plan; Canadian AI cloud residency index; CLOUD Act and Canadian AI.