Diagram of a CPU with four large cores next to a GPU with hundreds of small cores, titled CPU, GPU and CUDA fundamentals for engineers
A few clever cores on one side, thousands of simple ones on the other.

A CPU is a few fast, flexible cores; a GPU is thousands of simple ones that all do the same arithmetic at once. Running a model is mostly multiplying large tables of numbers, which is exactly the work a GPU is built for. CUDA is how you tell an NVIDIA GPU what to compute.

This is the background to my earlier post, which measured half of a small model's layers on a CPU and on a GPU. It starts with how the hardware works and ends with three Rust programs you can read line by line.

How a CPU works

A CPU (central processing unit) runs a program one small instruction at a time, very fast. Each instruction is tiny: add two numbers, copy a value, compare, jump to another instruction.

The loop every CPU runs

  1. Fetch: read the next instruction from memory.
  2. Decode: work out what it asks for (an add? a jump?).
  3. Execute: do it, using a handful of fast storage slots called registers.
  4. Write back: store the result, then move to the next instruction.

A clock drives this loop. A 4 GHz CPU ticks 4 billion times a second, and a modern core finishes several instructions per tick.

What makes a CPU core clever

A CPU core is built to make one stream of instructions finish as soon as possible, even when the code is full of decisions.

  • Branch prediction: it guesses which way an if will go and starts working before it knows.
  • Out-of-order execution: it reorders instructions so slow ones do not block fast ones.
  • Caches: small, very fast memories on the chip that keep recently used data close.

All that cleverness takes chip area, so a CPU has few cores: 16 in a current high-end desktop chip, and up to 192 in a server chip.

Why caches matter so much

Main memory (RAM) is slow compared with the core. Fetching a value from RAM costs roughly 100 times more than fetching it from the nearest cache. So the CPU keeps layers of cache, each bigger and slower than the last.

StorageTypical sizeRough access time
Registersa few hundred bytesunder 1 ns
L1 cachetens of KB per coreabout 1 ns
L2 cacheabout 1 MB per corea few ns
L3 cachetens of MB, sharedabout 10 ns
RAMtens of GBabout 60 to 100 ns

These are approximate and vary by chip. The lesson holds everywhere: code that walks memory in order is fast, and code that jumps around is slow.

SIMD: a small taste of parallelism

SIMD (single instruction, multiple data) lets one instruction work on several numbers at once. With 256-bit AVX2, one add instruction adds 8 pairs of 32-bit floats. With 512-bit AVX-512 it adds 16 pairs.

This is the same idea a GPU takes to the extreme: do one operation on many numbers together.

A CPU loop in Rust

This adds two arrays on the CPU. Keep it in mind: Example 1 below does the same job on a GPU.

fn add_arrays(a: &[f32], b: &[f32], out: &mut [f32]) {
    for i in 0..out.len() {
        out[i] = a[i] + b[i]; // one core walks the arrays, element by element
    }
}

The compiler usually turns this loop into SIMD instructions for you. It is still one core doing all the work.

How a GPU works

A GPU (graphics processing unit) is built for the opposite goal to a CPU: finish a huge amount of simple work per second, not one task as soon as possible.

GPUs were made to draw images. A screen has millions of pixels, and each pixel needs the same small calculation with different inputs. That is the pattern GPUs are good at: the same operation on a very large number of values.

Many simple cores

A GPU drops most of the per-core cleverness (branch prediction, reordering) and spends the chip area on arithmetic units instead. The result is thousands of simple cores.

  • One GPU core is slower than one CPU core.
  • All of them together do far more arithmetic per second.
  • They work best when every core runs the same code on different data.

A kitchen analogy

A CPU is a few master chefs. Each can cook any dish, handle surprises, and switch tasks quickly. A GPU is a hall of thousands of line cooks who each know one step. Ask for one complicated dinner and the chefs win. Ask for a million identical sandwiches and the hall wins by a wide margin.

The GPU has its own memory

A GPU cannot read your program's normal memory directly. It has its own memory, called VRAM, sitting next to it on the card.

  • Data must be copied from RAM to VRAM before the GPU can use it, and copied back to read the result.
  • That copy travels over the PCIe link, which is slower than either memory.
  • Whatever the GPU works on must fit in VRAM. This is why VRAM size decides which models a card can run.

How the work is organised

You do not program each GPU core by hand. You write one small function and ask the GPU to run it many times at once, each copy with a different index. The GPU's hardware groups those copies and schedules them across its cores. The CUDA section below gives these pieces their names.

When a GPU is the wrong tool

  • Small jobs: copying data to the GPU and back costs more than the work saved.
  • Branchy logic: code full of different decisions per item wastes the lockstep hardware.
  • Step-by-step work: if each step needs the previous result, there is nothing to run side by side.

CPU vs GPU side by side

CPU: a few clever cores Core 1 predicts, reorders own L1 + L2 cache Core 2 predicts, reorders own L1 + L2 cache Core 3 predicts, reorders own L1 + L2 cache Core 4 predicts, reorders own L1 + L2 cache Shared L3 cache GPU: thousands of simple cores RAM (host memory) where your program keeps data VRAM (device memory) the only memory the GPU sees copy over PCIe the slow step
A CPU has a few clever cores; a GPU has thousands of simple ones. Each has its own memory.

The drawing is schematic. A real desktop CPU has around 16 cores, and a high-end GPU has over 20,000.

CPUGPU
Built toFinish one task as soon as possibleDo the most arithmetic per second
CoresFew and complex: 16 in a desktop Ryzen 9 9950X3D, 192 in a server EPYC 9965Thousands and simple: 21,760 in a GeForce RTX 5090
MemoryRAM: tens to hundreds of GBVRAM: 32 GB on the RTX 5090, 141 GB on a datacenter H200
Memory speedAbout 90 GB/s at best with two channels of DDR5-56004,800 GB/s on the H200
Strong atDecisions, varied tasks, running the operating systemThe same math over huge arrays
Weak atHuge arrays of identical mathSmall jobs and code full of decisions

The two are joined by a PCIe link. At best it moves about 32 GB/s (PCIe 4.0, 16 lanes) or 63 GB/s (PCIe 5.0), slower than either memory. So good GPU code copies data across once and keeps it there.

Memory speed matters as much as core count. The next section shows why it often decides how fast a model runs.

How a model runs its computations

A model is a very large set of numbers, called weights, plus a fixed recipe of arithmetic that combines them with your input. Running a model means doing that arithmetic, and almost all of it is matrix multiplication.

Tensors: numbers in a grid

  • A vector is a list of numbers. [1.0, 2.0, 3.0] has shape (3).
  • A matrix is a table of numbers. Three rows of two columns has shape (3, 2).
  • A tensor is the general word for either, with any number of dimensions.

Text becomes tensors too. Each token (a word or piece of a word) is looked up in a table and replaced by a vector of a few thousand numbers.

One layer: multiply, add, squash

A layer takes a vector, multiplies it by a matrix of weights, adds a second set of weights called the bias, and applies a simple function called an activation. Example 3 below computes exactly this for 3 inputs and 2 outputs, with the arithmetic written out by hand. Real layers do the same thing with thousands of inputs and outputs.

Why this suits a GPU

Each number in the result of a matrix multiply can be computed without knowing any of the others. Multiplying two 4,096 by 4,096 matrices produces about 16.8 million results, and each is its own sum of 4,096 products. That is the GPU's ideal job: identical, independent arithmetic, in huge quantity.

One forward pass of a language model

  1. Tokens in. The text is split into tokens, and each token is replaced by its vector.
  2. Layers. The vectors pass through a stack of layers, often dozens. Each layer mixes information between tokens (attention) and then transforms each token (a small network). Both parts are mostly matrix multiplies.
  3. Scores out. The last layer's output becomes one score for every token the model knows.
  4. Pick. One token is chosen from those scores.
  5. Repeat. The chosen token is added to the text and the pass runs again.

A 500-token answer is 500 passes. Models save work between passes with a cache of earlier results. My post on CPU, CUDA and running a language model picks up from there, with measurements.

Where the memory goes

The weights must sit in memory, in VRAM if a GPU runs the model. The size is the number of weights times the bytes used for each. Storing weights in fewer bits is called quantization.

FormatBytes per weightA 7-billion-weight model needs
32-bit float4about 28 GB
16-bit float2about 14 GB
8-bit1about 7 GB
4-bitabout 0.5 to 0.6about 4 GB

The 4 and 2 bytes rule is from Hugging Face's guide. The 7-billion column is arithmetic from it, and real files run a little larger.

Where the time goes

To produce one token, the model reads nearly every weight once. So when a model writes text, the limit is often how fast memory can deliver the weights, not how fast the cores can multiply. That delivery rate is called memory bandwidth.

tokens per second ≤ memory bandwidth ÷ model size

For a 14 GB model, desktop RAM at about 90 GB/s allows at most about 6 tokens per second. An H200's VRAM at 4,800 GB/s allows a ceiling near 340. These are upper bounds worked out from the formula, and real speeds are lower.

This is a large part of why GPUs generate text faster. An NVIDIA Research study of the limits reaches the same conclusion: "LLM serving will be memory bandwidth constrained." Reading a long prompt is different: all its tokens are processed together, so that phase is limited by compute.

CUDA: the ideas you need

CUDA is NVIDIA's system for running your own code on its GPUs. You write a small function, and CUDA runs thousands of copies of it at once, one per piece of data.

The vocabulary

TermMeaning
HostThe CPU and its memory (RAM). Your Rust program lives here.
DeviceThe GPU and its memory (VRAM).
KernelA function that runs on the GPU, as many copies at once.
ThreadOne running copy of the kernel. It handles one piece of the data.
BlockA group of threads, at most 1,024. Threads in a block can share fast memory.
GridAll the blocks of one launch.
Warp32 threads that the hardware runs in lockstep. You rarely handle warps yourself.
StreamA queue of GPU work that runs in order.

How a thread knows which data is its own

Every thread runs the same code, so each one needs a way to pick a different element. CUDA gives every thread three built-in values:

  • threadIdx: the thread's position inside its block.
  • blockIdx: the block's position inside the grid.
  • blockDim: how many threads each block has.

From these, a thread works out its own global index:

size_t i = blockIdx.x * blockDim.x + threadIdx.x;

A worked case: blocks hold 256 threads. Thread 5 of block 3 gets i = 3 * 256 + 5 = 773, so it handles element 773. No two threads get the same i.

Grid: one launch, here 4 blocks of 8 threads threadIdx.x = 0 threadIdx.x = 7 Block 0 blockIdx.x = 0 Block 1 blockIdx.x = 1 Block 2 blockIdx.x = 2 Block 3 blockIdx.x = 3 i = 0 i = 8 i = 16 i = 24 i = 30 i = 31 i = 26 Block 3, thread 2 handles element 26 i = blockIdx.x × blockDim.x + threadIdx.x = 3 × 8 + 2 Spare threads i is 30 and 31, but n = 30, so the bounds check skips them
A CUDA grid: 4 blocks of 8 threads covering 30 elements.

The drawing shows a small case: 4 blocks of 8 threads launched for 30 elements. Each box is one thread, and the last two have no element to handle.

Why kernels start with a bounds check

Threads come in whole blocks. For 1,000 elements with 256 threads per block you need 4 blocks, which is 1,024 threads. The last 24 threads have no element. The line if (i < n) makes them do nothing instead of writing past the end of the array.

The five steps of every CUDA program

  1. Allocate memory on the device.
  2. Copy the inputs from host to device.
  3. Launch the kernel over a grid of threads.
  4. Wait for the GPU to finish.
  5. Copy the results from device back to host.

A launch does not wait: the call returns as soon as the work is queued, and the GPU runs it in the background. The copy back in step 5 is what waits for the answer. This matters when you time GPU code, as covered under common mistakes below.

CUDA from Rust: which crate to pick

Start with cudarc to learn how CUDA works, and use candle when you want to run a model. The table shows where each option stood on 4 October 2026.

OptionWhat you writeRust toolchainStatus
cudarc 0.19.10Host code in Rust; kernels as CUDA C++ text compiled at run timeStableActively released. Wraps the CUDA driver, cuBLAS, cuDNN and more.
candle 0.11.0Tensor operations in Rust; no kernelsStableHugging Face's ML framework. Its CUDA backend is built on cudarc.
cuda-oxide (NVIDIA CUDA Rust)Kernels in Rust, one thread at a time, like CUDA C++Pinned nightlyEarly alpha, Linux only. Announced 8 September 2026.
cutile 0.4.0 (NVIDIA CUDA Rust)Kernels in Rust that work on whole tiles of dataStable 1.89+Published on crates.io, Linux only.
Rust-CUDA (cust)Kernels in Rust through a custom compiler backendPinned nightlyCommunity project. Last crates.io release is from February 2022, so it is used from git.

Why this post uses cudarc for the examples

  • It builds on stable Rust with one dependency line.
  • The kernel is ordinary CUDA C++, so every CUDA tutorial and the official NVIDIA guide apply directly.
  • You see each step yourself: copy in, launch, copy out. Higher-level libraries hide exactly the part a beginner needs to see.

What you need to run the examples

  • An NVIDIA GPU with a current driver, on Linux or Windows. Macs have no CUDA.
  • The CUDA toolkit installed. cudarc loads its libraries when the program starts, and uses its NVRTC compiler to build the kernel text.
  • This in Cargo.toml, with the feature matching your toolkit version (cuda-12080 means CUDA 12.8):
[dependencies]
cudarc = { version = "0.19.10", features = ["cuda-12080"] }

How far these examples were tested

The examples below were written on a Mac, which cannot run CUDA. Both cudarc programs compile against cudarc 0.19.10, and both kernels give correct results when replayed thread by thread on a CPU. They have not been run on an NVIDIA GPU yet. The candle example was run for real and its output is shown.

Example 1: add two arrays on the GPU

This program adds one million pairs of numbers, one GPU thread per pair. It is the GPU version of the CPU loop shown earlier, and it follows the five steps exactly.

use cudarc::driver::{CudaContext, LaunchConfig, PushKernelArg};
use cudarc::nvrtc::compile_ptx;

// The kernel: CUDA C++ source, compiled for the GPU when the program runs.
// Every GPU thread runs this same function, each with a different index `i`.
const KERNEL_SRC: &str = r#"
extern "C" __global__ void add_arrays(float *out, const float *a, const float *b, size_t n) {
    size_t i = blockIdx.x * blockDim.x + threadIdx.x; // which element is mine?
    if (i < n) {                                      // the last block may have spare threads
        out[i] = a[i] + b[i];
    }
}
"#;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // 1. Make the input on the CPU side (the "host").
    let n: usize = 1_000_000;
    let a_host: Vec<f32> = (0..n).map(|i| i as f32).collect();
    let b_host: Vec<f32> = (0..n).map(|i| 2.0 * i as f32).collect();

    // 2. Connect to GPU number 0 and get a stream (a queue of GPU work).
    let ctx = CudaContext::new(0)?;
    let stream = ctx.default_stream();

    // 3. Compile the kernel and load it onto the GPU.
    let ptx = compile_ptx(KERNEL_SRC)?;
    let module = ctx.load_module(ptx)?;
    let add_arrays = module.load_function("add_arrays")?;

    // 4. Copy the inputs from host memory to GPU memory (the "device").
    let a_dev = stream.clone_htod(&a_host)?;
    let b_dev = stream.clone_htod(&b_host)?;
    let mut out_dev = stream.alloc_zeros::<f32>(n)?;

    // 5. Launch: enough blocks of threads to cover all n elements.
    let mut launch = stream.launch_builder(&add_arrays);
    launch.arg(&mut out_dev);
    launch.arg(&a_dev);
    launch.arg(&b_dev);
    launch.arg(&n);
    unsafe { launch.launch(LaunchConfig::for_num_elems(n as u32)) }?;

    // 6. Copy the result back to the host and check it.
    let out_host: Vec<f32> = stream.clone_dtoh(&out_dev)?;
    for i in 0..n {
        assert_eq!(out_host[i], a_host[i] + b_host[i]);
    }
    println!("ok: added {n} pairs, out[10] = {}", out_host[10]);
    Ok(())
}

Expected output: ok: added 1000000 pairs, out[10] = 30 (element 10 is 10 + 20).

Reading it piece by piece

The kernel. The text inside KERNEL_SRC is CUDA C++, not Rust. __global__ marks a function as a kernel. extern "C" keeps its name plain, so Rust can find it by the string "add_arrays". The body has no loop: the loop is replaced by a million threads, each doing one addition.

Step 2, context and stream. The context is your program's connection to one GPU. The stream is a queue: work you put on it runs in the order you added it.

Step 3, compile. compile_ptx turns the kernel text into PTX, the GPU's assembly language. This happens while your program runs, which is why the CUDA toolkit must be installed.

Step 4, copy in. clone_htod means "host to device". a_host lives in RAM and a_dev lives in VRAM. They are separate copies, and the kernel can only see the device ones.

Step 5, launch. The arguments must be added in the same order as the kernel's parameters. LaunchConfig::for_num_elems picks 1,024 threads per block and enough blocks to cover n. For one million elements that is 977 blocks, or 1,000,448 threads, so 448 spare threads hit the if (i < n) check and do nothing.

Why unsafe. Rust cannot check the kernel text against your arguments. If you pass them in the wrong order, or pass too few, the GPU reads the wrong memory. The unsafe keyword is you taking responsibility for that match.

Step 6, copy back. clone_dtoh means "device to host". It is queued on the same stream as the launch, so it runs after the kernel finishes and returns once the numbers are in the Vec.

Would this beat the CPU?

Probably not. Adding a million numbers took one CPU core 0.16 ms when measured on an Apple M5 Pro, and here the data must also cross to the GPU and back. This example teaches the mechanics. The GPU pays off when there is far more arithmetic per byte copied, as in the next example.

Example 2: matrix multiply

Matrix multiply is the operation a model spends most of its time in, and it fits the GPU well: every cell of the answer can be computed independently. Here one GPU thread computes one cell.

To get the cell at row r, column c of the result, take row r of A and column c of B, multiply them pair by pair, and add the products.

use cudarc::driver::{CudaContext, LaunchConfig, PushKernelArg};
use cudarc::nvrtc::compile_ptx;

// One GPU thread computes one cell of the result: C[row][col].
// Matrices are stored flat, row after row, so cell (r, c) lives at index r * n + c.
const KERNEL_SRC: &str = r#"
extern "C" __global__ void matmul(float *c, const float *a, const float *b, int n) {
    int row = blockIdx.y * blockDim.y + threadIdx.y;
    int col = blockIdx.x * blockDim.x + threadIdx.x;
    if (row < n && col < n) {
        float sum = 0.0f;
        for (int k = 0; k < n; k++) {
            sum += a[row * n + k] * b[k * n + col]; // row of A times column of B
        }
        c[row * n + col] = sum;
    }
}
"#;

// The same math on one CPU core, used to check the GPU's answer.
fn matmul_cpu(a: &[f32], b: &[f32], n: usize) -> Vec<f32> {
    let mut c = vec![0.0f32; n * n];
    for row in 0..n {
        for col in 0..n {
            let mut sum = 0.0f32;
            for k in 0..n {
                sum += a[row * n + k] * b[k * n + col];
            }
            c[row * n + col] = sum;
        }
    }
    c
}

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let n: usize = 512;
    let a_host: Vec<f32> = (0..n * n).map(|i| (i % 7) as f32).collect();
    let b_host: Vec<f32> = (0..n * n).map(|i| (i % 5) as f32).collect();

    let ctx = CudaContext::new(0)?;
    let stream = ctx.default_stream();
    let module = ctx.load_module(compile_ptx(KERNEL_SRC)?)?;
    let matmul = module.load_function("matmul")?;

    let a_dev = stream.clone_htod(&a_host)?;
    let b_dev = stream.clone_htod(&b_host)?;
    let mut c_dev = stream.alloc_zeros::<f32>(n * n)?;

    // A 2D grid: blocks of 16 x 16 threads, and enough blocks to cover n x n cells.
    let block: u32 = 16;
    let blocks_per_side = (n as u32).div_ceil(block);
    let config = LaunchConfig {
        grid_dim: (blocks_per_side, blocks_per_side, 1),
        block_dim: (block, block, 1),
        shared_mem_bytes: 0,
    };

    let n_arg = n as i32;
    let mut launch = stream.launch_builder(&matmul);
    launch.arg(&mut c_dev);
    launch.arg(&a_dev);
    launch.arg(&b_dev);
    launch.arg(&n_arg);
    unsafe { launch.launch(config) }?;

    let c_host: Vec<f32> = stream.clone_dtoh(&c_dev)?;
    assert_eq!(c_host, matmul_cpu(&a_host, &b_host, n));
    println!("ok: {n} x {n} matrix multiply matches the CPU result");
    Ok(())
}

What is new compared with Example 1

A 2D grid. Blocks and grids can have up to three dimensions. Here each block is 16 by 16 threads (256 in all), and the grid is 32 by 32 blocks. That gives 512 by 512 threads, one per cell of the result.

Two indices per thread. The .x values give the thread's column and the .y values give its row. It is the same formula as before, used once per dimension.

Flat storage. GPU memory holds a plain run of numbers, not a table. A matrix is stored row after row, so the cell at row r, column c is at position r * n + c.

A loop inside the kernel. Each thread runs a short loop of n multiply-adds. The CPU version has three nested loops. The GPU version keeps only the innermost: the outer two became the grid of threads.

Matching argument types. The kernel declares int n, which is 32 bits. So the Rust side passes an i32, not a usize. Sizes must match exactly.

Why this one suits the GPU

The inputs are 2 × 512 × 512 numbers, about 2 MB to copy. The work is 512³, about 134 million multiply-adds. Lots of arithmetic for little copying is the pattern where a GPU wins.

Do not ship this kernel

This version is written to be read. Real code calls a tuned library instead: cuBLAS is NVIDIA's library of matrix routines, and its gemm function does this job far faster by using shared memory, careful memory access order and special hardware. cudarc wraps it, and candle calls it for you, which is the next example.

Example 3: one model layer with candle

Real model code does not write kernels. It uses a tensor library that already has tuned kernels for every common operation, and you describe the math. This example computes one layer of a neural network with candle, the library my project Sangama uses.

[dependencies]
candle-core = "0.11.0"

[features]
cuda = ["candle-core/cuda"]
use candle_core::{Device, Result, Tensor};

fn main() -> Result<()> {
    // Use the GPU when this build has CUDA and a GPU is present; otherwise the CPU.
    let device = Device::cuda_if_available(0)?;

    // One input row with 3 numbers, and a layer that turns 3 numbers into 2.
    let x = Tensor::new(&[[1.0f32, 2.0, 3.0]], &device)?; // shape (1, 3)
    let w = Tensor::new(&[[0.5f32, -1.0], [0.25, 0.0], [-0.5, 2.0]], &device)?; // shape (3, 2)
    let bias = Tensor::new(&[0.1f32, -0.1], &device)?; // shape (2)

    // A layer is: multiply by the weights, add the bias, apply an activation.
    let y = x.matmul(&w)?.broadcast_add(&bias)?.relu()?;

    // Bring the result back to host memory to print it.
    println!("y = {:?}", y.to_vec2::<f32>()?);
    Ok(())
}

Output from running it: y = [[0.0, 4.9]]

Checking the answer by hand

StepFirst outputSecond output
Multiply by weights1×0.5 + 2×0.25 + 3×(−0.5) = −0.51×(−1) + 2×0 + 3×2 = 5
Add bias−0.5 + 0.1 = −0.45 − 0.1 = 4.9
ReLU (negative becomes 0)04.9

How this maps to what you just learned

In candleWhat happens underneath on a GPU
Device::cuda_if_available(0)Creates the context for GPU 0 (Example 1, step 2).
Tensor::new(..., &device)Allocates VRAM and copies the numbers host to device (step 4).
matmulLaunches a tuned matrix multiply kernel from cuBLAS (Example 2, done properly).
broadcast_add, reluLaunch small kernels with one thread per element (Example 1's pattern).
to_vec2Copies the result device to host (step 6).

Run it with cargo run to use the CPU, or cargo run --features cuda on a machine with an NVIDIA GPU. The code does not change. Only the device does.

A full language model is this same pattern repeated: thousands of matrix multiplies, additions and activations over much larger tensors.

Common beginner mistakes

MistakeWhat goes wrongFix
Timing only the launch callThe launch returns before the GPU has done the work, so the time looks near zero.Call stream.synchronize() before stopping the clock.
Timing the first runThe first run includes compiling the kernel and waking the GPU.Run once to warm up, then time later runs.
Forgetting the bounds checkSpare threads write past the end of the array and corrupt other data.Start the kernel with if (i < n).
Copying data back and forthEach copy crosses the slow PCIe link, and the copies cost more than the math.Copy in once, do all the steps on the GPU, copy out once.
Using the GPU for small jobsLaunch and copy overhead outweighs the work saved.Keep small work on the CPU. Measure before moving it.
Mismatched argument typesA Rust usize passed to a kernel int can make the kernel read garbage.Match each Rust argument to the kernel parameter's exact size.
Running out of VRAMThe allocation fails. A GPU has no swap space to fall back on.Use smaller batches, or store the weights in fewer bits.
Expecting identical decimalsCPU and GPU may round long sums differently, so results differ in the last digits.Compare floats with a small tolerance, not ==.

The last row has one exception worth knowing. Examples 1 and 2 above compare with exact equality on purpose: their inputs are small whole numbers, which floats add and multiply without any rounding.

Timing GPU code honestly

use std::time::Instant;

let start = Instant::now();
unsafe { launch.launch(config) }?;
stream.synchronize()?; // wait until the GPU has really finished
println!("kernel took {:?}", start.elapsed());

Decide what you are measuring. Kernel time alone tells you how fast the math is. Time including both copies tells you whether moving the job to the GPU was worth it.

Glossary

TermPlain meaning
ActivationA simple function applied to each number after a layer's multiply and add. ReLU turns negatives into 0.
Bandwidth (memory)How many bytes per second a memory can deliver.
BlockA group of up to 1,024 GPU threads.
CacheSmall, fast memory on the chip that keeps recent data close to the cores.
ContextYour program's connection to one GPU.
CoreOne unit that executes instructions.
cuBLASNVIDIA's library of tuned matrix routines for the GPU.
CUDANVIDIA's system for running your own code on its GPUs.
DeviceThe GPU and its memory.
Forward passOne run of an input through all of a model's layers.
GridAll the blocks of one kernel launch.
HostThe CPU and its memory.
KernelA function that runs on the GPU as many threads at once.
PTXThe GPU's assembly language, produced by compiling a kernel.
QuantizationStoring weights in fewer bits to save memory.
SIMDOne CPU instruction working on several numbers at once.
StreamA queue of GPU work that runs in order.
TensorNumbers arranged in a grid of any number of dimensions.
Thread (GPU)One running copy of a kernel.
TokenA word or piece of a word, as a model sees text.
VRAMThe GPU's own memory.
Warp32 GPU threads that execute in lockstep.
WeightsThe numbers a model learned in training.

Sources and further reading

All pages were opened on 4 October 2026. Hardware figures are single examples, not surveys, and change with each chip generation.

CUDA, from NVIDIA

Rust crates

Hardware figures

Models