Skip to content

Technology · 07 AI Compute & Chipsets

AI hardware we size for training, inference and the edge.

Most AI projects never need a training cluster. We match each workload to the right silicon, rent before we buy, and confirm availability, region and cost at scoping, because chip names and supply change quickly.

Updated September 2026 You own everything Reply within one business day

Vision - line 3 camera

streaming

Warehouse detector v4 - 640x640 - centre aisle

pallet 0.97
forklift 0.93
person 0.99
pallet
97%
forklift
93%
person
99%

38ms

Inference

24/s

Frames read

1,284

Objects tracked

AI Compute & Chipsets, in short

OlDevs, a full-stack technology studio in Vancouver, sizes AI compute from the model outward.We measure the memory a model needs at its target precision, then choose between rented cloud accelerators, a workstation in the client's office, or an edge device. Inference is the usual workload and gets the smaller hardware; training and fine-tuning are scoped separately and rented by the hour until steady demand justifies owning anything.

Key facts

01Default cloud accelerator
NVIDIA H100 or H200, Blackwell when the region has it
02Sizing constraint
VRAM or HBM at the target precision, then interconnect
03Default posture
Rent by the hour; buy only after a year of steady load
04Local experiments
RTX PRO Blackwell or RTX 50-series workstation, Apple M4 or M5
05Edge default
NVIDIA Jetson Orin, or Hailo on Raspberry Pi 5 for low power
06Updated
September 2026

07 · AI Compute & Chipsets

Six classes of AI hardware, and when each one fits

Every option below is a trade between memory, throughput, availability and who has to keep it running after launch.

01

NVIDIA Hopper and Blackwell in the cloud

H100 and H200 are the safe pick for fine-tuning and serving large models: CUDA tooling is mature and both are bookable in Canadian and US regions. Blackwell (B200, GB200 NVL72, B300 Blackwell Ultra) suits the largest jobs but is still allocated, so we reserve at scoping. Later cost: rental that runs until switched off.

Training · Inference · CUDA · Cloud

02

AMD Instinct and Intel Gaudi 3

MI300X, MI325X and the MI350 series carry more HBM per card than the matching NVIDIA generation, so a large model fits on fewer GPUs. Gaudi 3 fits a client already committed to Intel or wanting a second source. Later cost: fewer engineers who know ROCm or Gaudi, and slower support for new models.

Inference · HBM · Second source

03

Google TPU and AWS Trainium or Inferentia

Trillium (v6e) and Ironwood (v7) on Google Cloud, and Trainium2, Trainium3 and Inferentia on AWS, fit a client who already lives on that cloud and runs JAX or PyTorch XLA. Queues are shorter than for NVIDIA parts. Later cost: portability, since leaving that silicon means re-testing the whole serving stack.

Cloud-native · Inference · Lock-in risk

04

Inference specialists: Groq and Cerebras

Groq and Cerebras serve open-weight models over an API with very low latency per token, ideal for assistants and agents where response time is the product, and for clients who do not want to run a GPU at all. Later cost: a limited model catalogue and one provider's roadmap, so we keep a standard API in front of it.

API · Low latency · Open models

05

Workstations: RTX PRO Blackwell, RTX 50-series, Apple M4 and M5

A desktop with an RTX PRO 6000 Blackwell or GeForce RTX 5090, or a Mac with an M4 Max or M5 and generous unified memory, runs mid-size open models for prototyping and private data work with no cloud bill. Later cost: one box cannot serve a production audience, so the cloud path is designed before the first demo.

Prototyping · Private data · On-premises

06

Edge AI: Jetson, Raspberry Pi 5 with AI HAT+, Hailo, Coral

Jetson Orin and Thor modules run vision and small language models on a machine, vehicle or kiosk; Raspberry Pi 5 with the AI HAT+ (Hailo) and Google Coral cover low-power detection. Right when the connection is unreliable or data must never leave the site. Later cost: updating models on devices you cannot walk to.

Edge · Offline · Computer vision

Stack

What we build on.

Concrete tools, current at the time of writing; confirmed for your project at scoping.

Cloud accelerators

NVIDIA H100NVIDIA H200NVIDIA B200 and GB200 NVL72AMD Instinct MI300X and MI350 seriesGoogle TPU v6e Trillium and v7 IronwoodAWS Trainium2 and Inferentia2Intel Gaudi 3

Serving and inference

vLLMSGLangNVIDIA TensorRT-LLMNVIDIA Triton Inference ServerHugging Face Text Generation Inferencellama.cppOllama

Training and fine-tuning

PyTorch 2.xCUDA (current release)ROCm (current release)JAXHugging Face Transformers and PEFTDeepSpeedRay

Quantisation and model formats

GGUFAWQGPTQbitsandbytesONNX RuntimeCore ML ToolsTensorRT

Workstations and edge devices

NVIDIA RTX PRO 6000 BlackwellNVIDIA GeForce RTX 5090Apple M4 Max and M5NVIDIA Jetson Orin Nano and AGX OrinNVIDIA Jetson AGX ThorRaspberry Pi 5 with AI HAT+Hailo-8 and Hailo-10HGoogle Coral

Where it runs

AWS Canada (Central) in MontrealGoogle Cloud Montreal and TorontoMicrosoft Azure Canada CentralOVHcloud BeauharnoisLambdaCoreWeaveRunPodKubernetes with the NVIDIA GPU Operator

What changed

2025–2026 updates.

What moved in this area and what it means for your build.

  1. Early 2025

    Blackwell arrived on the desk before the data centre

    GeForce RTX 50-series cards and the RTX PRO Blackwell workstation line shipped, with the RTX PRO 6000 carrying far more memory than earlier single workstation cards. A mid-size open model can now be prototyped and evaluated on one machine in the office before any cloud spend.

  2. 2025

    Blackwell went live in the cloud, Hopper got easier to book

    B200 instances and GB200 NVL72 racks came online at the hyperscalers and specialist GPU clouds through the year, and H100 and H200 queues shortened as a result. Fine-tuning on Hopper is now a routine booking; we hold Blackwell for jobs that need the interconnect.

  3. Mid 2025

    More HBM per GPU: Blackwell Ultra and AMD MI350

    NVIDIA announced B300 and GB300 Blackwell Ultra, and AMD shipped the MI350 series, both with a larger memory stack per accelerator. Larger models fit on fewer cards, but since a new SKU appears every year we specify by memory and interconnect rather than by part name.

  4. 2025

    Cloud-native silicon focused on inference

    Google introduced Ironwood (TPU v7) alongside Trillium, and AWS moved Trainium2 to general availability with Trainium3 following. Clients already on Google Cloud or AWS gain an inference option that avoids the NVIDIA queue, with a serving stack tied to that cloud.

  5. Late 2025

    Apple silicon for on-device models

    The M5 generation improved the neural accelerators inside Apple's GPU cores, and the Foundation Models framework introduced with iOS 26 gave apps a supported way to run small models on the device. An iOS or macOS product can now ship private, offline AI features without a server round trip.

  6. 2026

    Cheaper inference and more Canadian capacity

    The trend at the time of writing is falling cost per token through quantisation, speculative decoding and specialist providers, alongside a growing sovereign compute effort in Canada. In-country residency is more achievable than it was, though GPU choice in Canadian regions is still narrower than in the US; we confirm both at scoping.

How we choose

Six questions we ask before recommending anything.

01

Who will run it after launch

A rented H100 with a managed serving layer suits a client with no infrastructure team. A workstation or Jetson fleet is only right if someone on the client's side, or on ours under a support agreement, will patch drivers and roll out models.

02

Where the data is allowed to live

Health, public-sector and financial data often has to stay in Canada. That narrows the accelerator list before performance is discussed, and sometimes a slightly older GPU in Montreal beats a newer one in Oregon.

03

Memory first, throughput second

We compute the VRAM or HBM a model needs at its precision, with headroom for context and batch size. If it does not fit, no amount of compute helps. Only then do we look at tokens per second and interconnect.

04

Rent before you own

Hourly cloud GPUs cost nothing when idle and let us change chip generation freely. We recommend buying hardware only when utilisation has been steady for about a year and the client has somewhere secure and cooled to put it.

05

The budget of change

Every accelerator outside the NVIDIA and CUDA mainstream saves money now and costs portability later. If a client may need to move clouds within three years, we stay on portable formats and standard serving APIs and say so in the proposal.

06

Availability on the day, not in the brochure

Chip names, quotas and regional stock change month to month. Before any proposal we check what is actually bookable in the target region, hold the capacity, and write the plan around that rather than around a vendor slide.

Who it's for

AI Compute & Chipsets for organisations that have to get it right.

Whether the audience is a customer, a member, a citizen, or your own team, the choice has to hold up under real use.

Corporations

Fine-tuning on private documents, internal assistants and forecasting models. We usually start on rented Hopper GPUs in a Canadian region with a serving layer your IT team can audit, and plan reserved or owned capacity only once usage settles.

Associations and government

Residency and procurement rules come first. We favour in-country cloud regions, open-weight models and portable formats so a future vendor change does not strand the work, and we document the hardware choice in language an auditor can follow.

Franchises

Many locations, thin margins per site. Edge devices such as Jetson or a Raspberry Pi 5 with an AI HAT+ handle in-store vision locally, with one central cloud model for the analytics head office needs. Fleet updates are planned from day one.

Startups

Speed and burn rate. We prototype on a workstation or an inference API, quantise aggressively, and delay owning any GPU until there is a paying audience. When investors ask about compute, you will have a written sizing rationale.

FAQ

AI Compute & Chipsets — questions we hear first.

Almost never at the start. Rented cloud GPUs, an inference API or a single workstation cover prototyping and the first production release for most clients. Owning hardware only pays when utilisation is high and steady for many months, and it brings power, cooling, security and driver maintenance with it. We revisit the question once real usage data exists.

By memory first. If the model, its context window and the batch fit in an H100, that is the cheapest bookable option. H200 adds memory for larger models or longer contexts on the same software. Blackwell earns its place for multi-GPU training or very large serving jobs where the interconnect matters, and only where the region actually has stock.

Usually, yes. AWS, Google Cloud and Microsoft Azure all run Canadian regions with GPU instances, and OVHcloud has GPU capacity in Quebec. The catch is choice: the newest accelerators reach US regions first, so an in-country requirement can mean a previous-generation GPU or a longer wait. We confirm what is bookable in Canada during scoping.

Quantisation stores a model's weights in fewer bits, which shrinks memory use and speeds up inference. Going from 16-bit to 8-bit is usually invisible in output quality; 4-bit formats such as GGUF or AWQ trade a little accuracy for running on much smaller hardware. We measure the difference on the client's own evaluation set before committing.

Small and mid-size models can. Apple's M4 and M5 chips pair a Neural Engine with unified memory, so a MacBook can run a quantised open model locally, and iOS offers Core ML and the Foundation Models framework for on-device features. Larger models still need a server. We decide per feature: private and offline tasks on the device, heavy reasoning in the cloud.

When the device must act without a reliable connection, when video or sensor data is too large to stream, or when data must never leave the site. Jetson Orin and Thor handle real-time vision and small language models; Raspberry Pi 5 with an AI HAT+ or a Coral accelerator suits simpler detection at low power. Every edge project includes a plan for remote model updates.

Yes, when the fit is right. AMD Instinct offers more memory per card, Google TPU and AWS Trainium suit clients already on those clouds, and Groq or Cerebras give very fast inference over an API. NVIDIA stays the default because its software ecosystem is the broadest and hiring is easier. We keep the serving interface standard so a later switch is configuration, not a rewrite.

Still have a question? Ask us when you request a quote

Let’s connect

Let’s scope your build.

Tell us what you’re building. We’ll reply within one business day with recommended platforms, structure and a tailored quote — no obligation.

We’ll only use your details to prepare your quote. No lists, no spam.

Call us Request a quote