Skip to content
Small team, full backlog, zero orders dropped. Support replies are slower than we’d like. Read our status update → Zero orders dropped. Status → 📬 Check your spam folder — most of our replies land there. We do answer. Status update → 📬 Check your spam folder. Status →

D-Central / private infrastructure / Canada

Put the model inside your boundary.

D-Central designs and deploys open-weight AI on hardware your organization controls—from a private workstation to a shared GPU server or a deliberately air-gapped system.

  • Local inference
  • Private RAG
  • LAN-only or air-gapped
  • Operator runbook

What is on-premises AI? The model executes on computing equipment inside a network and control boundary your organization administers. Prompts, retrieved documents, embeddings, outputs, and logs can remain local when the full architecture—not only the model process—is configured that way. D-Central scopes the workload, selects and benchmarks an open-weight model, supplies or integrates the hardware, deploys the runtime and interface, and hands over an operating runbook.

The box is only one part of private AI

Installing Ollama is easy. Building a dependable organizational service means deciding who can reach it, which documents it can retrieve, where embeddings and logs live, how updates cross the boundary, what happens when the GPU fails, and how the team can change models later. A private model with cloud telemetry, an unmanaged vector database, or shared administrator credentials is not a private system.

Three deployment shapes

One to five users

Private workstation

A high-memory GPU workstation for a professional, founder, researcher, or small team. Best when the workload is interactive, concurrency is low, and the simplest physical control boundary is the right one.

Team service

Shared inference server

A central GPU node with authenticated chat or an internal API, shared model serving, private document retrieval, usage controls, and an operating plan. Best when a department needs one governed service.

Highest isolation

LAN-only or air-gapped

A deliberately isolated system with controlled model and document import, local registries, restricted administration, and a defined update process. Best for high-sensitivity workloads that can accept operational friction.

What an engagement can include

Layer Typical deliverable Decision we document
Workload Use-case and data-flow inventory What should be local, hybrid, or not automated
Model Task-specific evaluation and licence review inputs Quality, language, context, tool use, origin, and commercial terms
Hardware GPU, memory, storage, network, power, heat, and expansion design Interactive latency, concurrency, availability, and growth
Serving Ollama, llama.cpp, or vLLM deployment; private UI or API Ease, throughput, hardware support, and portability
Knowledge Private RAG with local embeddings and vector storage Source permissions, citations, refresh, deletion, and evaluation
Security Authentication, roles, secrets, logging, telemetry audit, and patch plan Who can access what, and what leaves the boundary
Operations Backups, recovery, monitoring, model registry, and handover runbook How your team keeps working without depending on us

How we select the model

We do not sell a permanent allegiance to one model family. We start with a small set that satisfies the workload and licence constraints, then test it against examples your organization recognizes. An Artificial Analysis score can create a shortlist; it cannot tell a Quebec engineering firm whether the model reliably extracts the right fields from its bilingual maintenance reports.

As of August 24, 2026, Qwen3.8-27B is a particularly strong single-GPU candidate: it carries Apache 2.0 terms and scores 52 on Artificial Analysis v4.1.1. Cohere Command A+ is a Canadian-origin Apache-2.0 option designed for enterprise RAG, citations, vision, and multilingual use. Larger MIT-licensed models such as DeepSeek V4 Pro or GLM-5.2 move the deployment into a rack-scale class. We verify current model cards, licences, serving support, and actual task results before a production recommendation.

Read the current enterprise model shortlist →

A typical delivery path

1

Boundary and workload workshop.
Identify sensitive data, current vendors, users, volume, latency, languages, integrations, and recovery expectations.
2

Benchmark and architecture.
Run representative tasks, choose a model/runtime pair, size the hardware, and map every data flow.
3

Pilot.
Deploy one bounded workflow with acceptance tests. Measure answer quality, latency, throughput, retrieval quality, failure modes, and staff fit.
4

Production build.
Install the approved stack, authentication, storage, network rules, monitoring, backups, and interfaces.
5

Handover.
Deliver the runbook, administrative access, model and data inventory, recovery procedure, and known limits.

Privacy and compliance: architecture helps, governance finishes the job

On-premises inference can remove entire third-party data transfers and make retention and access easier to prove. It does not automatically establish compliance with Quebec Law 25, PIPEDA, professional duties, health-sector requirements, or a customer’s contract. Your organization still needs authority, transparency, access control, security, retention, and accountable human decisions. We document the technical boundary so your legal, privacy, security, and governance teams have evidence they can actually evaluate. This is technical orientation, not legal advice.

Request an on-premises AI assessment

Tell us the first workflow, the number of users, whether the system can reach the internet, the kind of data involved, and any hardware you already own. If on-prem is the wrong fit, we will say so and map a hybrid or Canadian-hosted route instead.

Frequently asked questions

Can the system run with no internet connection?

Yes, if the workload supports the trade-off. A true air gap needs a controlled method for importing models, packages, security updates, and documents; it is more than disabling Wi-Fi. We scope that operating procedure with the build.

Can staff use a familiar chat interface?

Yes. A private web interface can provide chat, file upload, citations, and model selection inside the network. We can also expose an internal API for approved business applications.

Can we connect our internal documents?

Yes. Private retrieval-augmented generation can index approved content locally, retrieve relevant passages, and attach citations to answers. Source permissions, deletion, refresh, and retrieval evaluation matter as much as the model.

Will one GPU support the whole company?

Sometimes, but user count alone is not enough. Prompt size, model size, simultaneous requests, output length, latency target, and availability expectations determine capacity. We measure a pilot instead of guessing.

Do we have to abandon cloud AI?

No. Many organizations should use a governed hybrid: local for sensitive and repeatable work, a hosted frontier model for approved tasks, and a model-neutral gateway so either side can change.

Related: local AI hardware guide; Ollama vs vLLM vs llama.cpp; local LLM security hardening; private RAG for Canadian organizations; AI sovereignty consulting.

Live Canadian SKUs (24 August 2026 prices): NVIDIA DGX Spark — $9,449 CAD · AMD Strix Halo 128 GB — $7,349 CAD · custom GPU inference rig (quote). Retired vapour names (Pleb AI Box, Workstation 24/48, Hashcenter AI Node 80+) are not for sale.