Notebook

GPU & inference

Serving, kernels, quantization and the cost of actually running the model.

What belongs here

GPU & inference is for serving, kernels, quantization and the cost of actually running a model. Notes here are about hardware, latency and money — not about writing prompts.

Titles, summaries and this definition are in the HTML. JavaScript only opens the mobile menu.

Nothing in this section yet.

This index is published even when empty so the section has a public address. New notes will appear below as HTML, not after JavaScript.