+ Field notes · product engineering

The machine
decides.

Running open models on hardware you already own — and why the constraint you are actually under determines which capabilities you can build on.
Starts at: what a model file physically contains. No ML background assumed.
Ends at: how to turn a hardware envelope into a product requirement.
Built around: one constrained workstation, eight models, and a piece of received wisdom that did not survive measurement.

+ Before we start

Almost everyone learning to build with AI starts at the API. That teaches you prompting. It does not teach you what a model is, what it costs, or what happens when the budget is not someone else's.

Running models on your own hardware teaches all three, quickly and unforgivingly. A cloud endpoint hides the weights, the memory, the arithmetic, and the bill. A machine on your desk hides nothing. When a model will not load, you learn what it weighs. When generation crawls, you learn which resource you ran out of. When a 30-billion-parameter model answers worse than a 7-billion one on your actual task, you learn that size is not capability.

This guide is about that education. It is not a tutorial for one stack, and it is deliberately light on model names, because those expire. It is about the physics underneath, which does not.

The claim it argues: for any given machine, a small number of hardware facts determine which AI capabilities you can realistically develop on it — and most published advice about models is derived in a regime that is not yours. The remedy is not better advice. It is measurement, on your hardware, against your tasks.

Claude appears throughout as a working tool. The dotted blue boxes contain prompts and workflows you can lift directly — for deriving your machine's limits, designing an evaluation set, and interpreting results without deceiving yourself.

Reading the plates Colour is a legend rather than decoration, and holds the same meaning across every plate. Amber is memory and bandwidth — the resource you run out of. Blue is compute. Magenta is stored weights. Green is useful output. Orange marks the one detail in a plate that matters most. Grey is capacity that exists but cannot be used.

+ Part I

What a model actually is.

Before anything about speed or capability, it helps to know precisely what you are downloading and what the machine does with it. Almost every confusion later resolves into a misunderstanding here.

01 — Chapter

A file of numbers, and a loop that reads it whole.

A language model is a file containing a very large list of numbers called parameters, plus a small description of how they are arranged into layers. Nothing in the file is text. Nothing in it is knowledge in any form you could inspect. It is a matrix of learned weights.

Generating text works like this. Your prompt is cut into tokens — roughly word fragments. The tokens are turned into vectors, and those vectors are pushed through every layer of the network, multiplied against the weights at each step. Out the other end comes a probability distribution over the vocabulary: how likely each possible next token is. One token is chosen, appended to the sequence, and the entire process runs again from the top.

That last sentence is the one to hold on to. Producing a 500-token answer means running the full network 500 times. There is no way to generate a sentence in one pass. The model has no memory of its own computation between steps — only the growing sequence of tokens it has already produced.

01 — One token per full pass through the weights
WEIGHTS IN RAM every layer, every time FULL PASS pass 1 FULL PASS pass 2 FULL PASS pass 3 · · · × 500 1 token 1 token 1 token The weights are re-read from memory on every single pass. That re-reading is the cost that dominates everything.
Memory traffic Compute Stored weights Output
Plate 01 — Generation is a loop, not a function call. The dashed amber paths are the part people forget: the whole weight file is streamed from memory once per token produced.
The consequence, stated early If producing one token requires reading every weight, then the speed at which your machine can read memory sets a hard ceiling on tokens per second. No amount of processor speed moves that ceiling. Chapters 5 and 6 do the arithmetic; it is worth carrying the intuition until then.
02 — Chapter

Quantization: why four bits is nearly free and three is not.

Models are trained with weights stored as 16-bit or 32-bit floating-point numbers. At 16 bits, a 7-billion-parameter model is about 14 GB, and a 70-billion one about 140 GB. On consumer hardware that is the end of the conversation.

Quantization is the compression that makes local inference possible at all. The idea is simple: instead of storing each weight at full precision, divide the range of values into a small number of buckets and store the bucket index. At 4 bits you have sixteen buckets per weight rather than 65,536 levels.

Naively this should be destructive, and done naively it is. The reason it works in practice is that modern formats do not use one bucket scheme for the whole model. Each weight matrix is divided into groups of 256, and every group subdivided again into eights of 32, with a scaling factor recorded at both levels — so the sixteen available bucket values are re-fitted to each small neighbourhood rather than stretched across the whole matrix. The bookkeeping for all those scaling factors adds roughly half a bit to every weight, which is why a nominally 4-bit format occupies closer to 4.5 — and the accuracy recovered by fitting locally exceeds what that extra storage costs.

The result is a quality curve that is not at all what intuition suggests.

02 — Quality against size · the cliff is not where you expect
100% 92% 84% QUALITY VS FULL PRECISION Q2_KQ3_K_MQ4_K_M Q5_K_MQ6_KQ8_0 the default worth defending THE CLIFF Q8 → Q4 costs about 2% quality and saves about 40% of memory. Q4 → Q3 costs several times more. Approximate community perplexity figures; exact values vary by model. The shape is what matters, not the decimals.
Quality curve The decision point
Plate 02 — The intuition that "more bits is proportionally better" is wrong. Quality is nearly flat from 8 bits down to 4, then falls away sharply. That flat region is what makes local inference viable.

Two refinements are worth knowing, because they turn this into an actual decision rule rather than a default.

the damage is uneven across skills: logical reasoning holds up well under compression, whereas numerical accuracy starts slipping once you drop below four bits. If your application does exact numeric work, spend the memory. If it does classification, summarisation, or discussion, do not.

And format choice interacts with your hardware. The more aggressively compressed variants trade decode speed for file size on processors, so the block-scaled family often generates faster on consumer machines despite the larger download. On a CPU box, the smallest file is often not the fastest.

Work with Claude — turn a spec sheet into a shortlist Rather than guessing which quantization to pull, hand over your constraints and make the trade-off explicit.
I have [RAM] of system memory, [channels]-channel [DDR4/DDR5-speed] memory,
and no CUDA GPU. I want to run a [size]B parameter model for [task].

Work out, showing your arithmetic:
1. Approximate file size at Q4_K_M, Q5_K_M and Q8_0
2. How much RAM remains for KV cache after each
3. Which of these my machine can actually hold at [context] tokens
4. Whether my task is quantization-sensitive, and why

Do not recommend a specific model. I want the envelope, not a pick.
Asking for the envelope rather than a recommendation is deliberate — you want the constraint reasoning to stay yours, and it does not expire when models change.
03 — Chapter

Inside the container, and why loading is instant.

Quantized weights ship in a container format that bundles the weights, the quantization scales, the tokenizer, and the architecture metadata into a single file. Older schemes split these across several files, which made distribution fragile.

One property of that packaging has an outsized effect on how local inference feels. Startup feels instantaneous because the file is mapped into the address space rather than read into it. Memory-mapping means pages are pulled from disk on demand and shared, so a model that is nominally 9 GB does not require a 9 GB read before the first token.

This is why a model can appear to load in a second and then be slow on its first answer: the loading you expected was deferred, not avoided.

03 — What is inside a quantized model file
THE FILE METADATA — LAYERS, DIMENSIONS TOKENIZER VOCABULARY QUANTIZED WEIGHT BLOCKS >99% of the file expanded at right ONE BLOCK, EXPANDED BLOCK OF 256 WEIGHTS carries its own scale and zero-point SUB-BLOCKS OF 32, EACH WITH ITS OWN SCALE SCALES COST ~0.5 BIT PER WEIGHT and buy back more accuracy than they cost
Adaptive structure Stored weights Why it works
Plate 03 — The nested block structure is the entire trick. Sixteen bucket values are useless across a whole matrix and remarkably sufficient across 32 neighbouring weights.
04 — Chapter

What actually occupies your memory.

People size their machine against the file on disk. That number is always too small, and the gap is where most failed loads come from.

Three things consume RAM during inference.

The weights. Roughly the file size, since the file is mapped rather than expanded.

The KV cache. Every token processed produces intermediate key and value vectors at every layer, and these are kept so that subsequent tokens do not have to recompute the whole history. The cache grows linearly with context length and is not part of the file size anyone quotes. At long context it can rival the weights themselves.

Runtime overhead. Activations, buffers, the framework, the operating system. Small but not zero.

04 — Memory budget at three context lengths
RAM CEILING 4K context 32K context 128K context WEIGHTS WEIGHTS KV CACHE WEIGHTS KV CACHE Same model, same file size. Only the context setting changed — and the third case no longer fits. Proportions are illustrative; the growth pattern is the point.
Weights KV cache Overhead Where it breaks
Plate 04 — Context length is a memory setting disguised as a capability setting. Doubling it does not double your abilities; it does inflate your footprint, and eventually it is what stops the model loading.
The most common self-inflicted failure Setting context to the maximum the model advertises, because it is free to type. It is not free. A configuration that runs comfortably at moderate context can become unstable at the advertised maximum, and the failure looks like a mysterious crash rather than a budget you overspent.

+ Part II

The two-phase machine.

This part contains the single most useful idea in the guide. Inference is not one workload — it is two, with opposite bottlenecks. Once you see the split, most puzzling behaviour on constrained hardware stops being puzzling.

05 — Chapter

Prefill and decode: one workload, two bottlenecks.

When you send a prompt, the machine does two very different jobs in sequence.

Prefill processes your entire prompt at once. Every token in it can be handled in parallel, because they are all already known. This is dense matrix work across many tokens simultaneously, and it saturates arithmetic units. Prefill is compute-bound.

Decode generates the answer one token at a time. Each token depends on the one before it, so there is nothing to parallelise. Each step reads the whole weight file to produce a single token. Decode is bandwidth-bound.

the arithmetic load concentrates almost entirely in prefill, while decode performs comparatively little computation and instead spends its time moving a great deal of data, and each weight is used exactly once before the next token requires it to be read again, and that single-use pattern is what pins latency to bandwidth rather than to processing power.

05 — The same request, two different machines
PREFILL — READING YOUR PROMPT all prompt tokens known — processed together ARITHMETIC UNITS SATURATED COMPUTE-BOUND scales with prompt length more cores genuinely help DECODE — WRITING THE ANSWER each token must wait for the one before it WHOLE WEIGHT FILE RE-READ, EACH TOKEN BANDWIDTH-BOUND independent of prompt length more cores do almost nothing Two phases, opposite constraints. Optimising the wrong one is the most common wasted afternoon in local AI.
Compute-bound Bandwidth-bound Output tokens Cannot start yet
Plate 05 — This split explains the characteristic rhythm of local inference: a long pause where nothing appears, then text arriving at a steady rate. The pause is prefill. The steady rate is decode.
Diagnostic value Next time a local model feels slow, ask which half is slow. A long wait before the first token with fluid output afterwards is a prefill problem — shorten the prompt. Output that trickles out word by word regardless of prompt size is a decode problem — use a smaller model or a machine with more memory bandwidth. These have completely different remedies, and neither fixes the other.
06 — Chapter

The bandwidth wall, in arithmetic.

You can predict your machine's generation speed to within a small factor without running anything. The calculation takes one line, and it is the most useful thing in this guide.

Producing one token requires reading every weight once. So:

tokens per second ≈ memory bandwidth ÷ model size in memory

A 9 GB model on a machine that can sustain 60 GB/s gives about 6.7 tokens per second as a theoretical ceiling. Real systems reach perhaps 60–80% of theoretical bandwidth, so expect 4 to 5. If you measure 4.5, your system is working correctly — that is not a configuration problem to solve.

Now the uncomfortable comparison. a consumer memory subsystem supplies something on the order of 100 GB/s, where datacentre accelerator memory reaches roughly ten times that. That order of magnitude is the difference between local and cloud inference. It is not driver tuning, thread count, or software maturity.

06 — Bandwidth divided by model size gives your ceiling
SUSTAINED MEMORY BANDWIDTH DUAL-CHANNEL DDR5 · ~60-80 GB/s QUAD-CHANNEL WORKSTATION · ~120 GB/s DATACENTRE GPU HBM · ~1000 GB/s RESULTING CEILING ON A ~60 GB/s MACHINE 4 GB MODEL ~15 tok/s — conversational 9 GB MODEL ~6.7 tok/s — usable 15 GB MODEL ~4 tok/s — patient
Bandwidth available Weights to move Resulting speed Patience required
Plate 06 — Model size and generation speed are inversely proportional, with the constant set by your memory subsystem. There is no software fix for the position of this line.
The upgrade that does not help how many channels the memory controller runs matters considerably more than the rated clock of the modules plugged into it. Buying faster RAM at the same channel count moves the ceiling slightly. Buying a faster processor moves it not at all. If you are choosing hardware for local inference, channel count and total bandwidth are the specification to read — and it is rarely on the marketing page.
Work with Claude — predict before you measure Doing this before you download anything builds the intuition properly, because you commit to a prediction and then find out whether your model of the machine was right.
My machine: [CPU], [RAM amount and type], [memory channels], [GPU or none].

Before I run anything, help me predict:
1. My approximate sustained memory bandwidth, and how you derived it
2. Expected decode speed for a 4 GB, 9 GB and 15 GB model
3. The largest model size that would still give me above 5 tokens/second
4. What measurement I should take to check whether your estimate was right

State your assumptions explicitly so I can tell which one was wrong
if the measurement disagrees.
Then measure. A prediction that misses by 3× usually means one assumption was badly off, and finding which one teaches you more than a correct answer would have.
07 — Chapter

Why long context costs you twice.

Context length is sold as a capability — "128K context" sounds like a feature you either have or lack. On constrained hardware it is better understood as a bill that arrives twice.

The first charge is memory. The KV cache grows linearly with the number of tokens held, as shown in plate 04. Long context consumes RAM that is then unavailable for weights.

The second charge is time. Prefill cost scales with prompt length, and the attention component scales worse than linearly. A prompt ten times longer takes more than ten times as long to process.

This is why pasting a large document into a local model produces a pause that feels disproportionate. It is disproportionate — that is the arithmetic working as designed.

07 — Time to first token against prompt length
TIME TO FIRST TOKEN 1K4K8K 16K32K PROMPT TOKENS linear, if it were the point where a workflow stops feeling interactive Shape is illustrative. Measure your own curve — chapter 11 shows how.
Measured cost Linear expectation Practical limit
Plate 07 — The gap between the solid line and the dashed one is the part that surprises people. Doubling your prompt more than doubles your wait.
08 — Chapter

Why autonomous agents stall on constrained hardware.

Everything so far converges on one practical conclusion, and it is the most valuable thing a constrained machine teaches you.

An agentic loop — a model that calls tools, reads results, and decides what to do next — works by resending the accumulated conversation on every step. Step one has a small prompt. By step ten the prompt contains the original instruction, ten tool calls, and ten sets of results.

Every step re-pays the prefill cost for the entire accumulated history. Not the new part. All of it.

08 — Where the time goes in a ten-step agent loop
EACH BAR = ONE AGENT STEP · WIDTH = TIME step 1 step 2 step 3 step 5 step 7 step 10 Blue = re-reading the whole conversation again. Green = the new tokens you actually wanted. The useful work per step is constant. The overhead grows every step. On a bandwidth-bound machine this is the difference between a demo that works and a workflow that does not.
Prefill re-paid New output Where it becomes unusable
Plate 08 — Agentic loops are not slow on constrained hardware because the model is weak. They are slow because the architecture repays a growing tax at every step, and prefill on a CPU is expensive.
The design lesson, arriving early A constrained machine does not make you worse at building with AI. It makes context economy a visible cost rather than an invisible one — and context economy is exactly the discipline that separates AI systems that scale affordably from those that do not. On a cloud endpoint this same waste appears only on the bill, months later. Here it appears in seconds, today.
Transfer to product engineering If you are designing a product feature around an LLM, the shape of plate 08 is your cost model in miniature. Ask early: does this feature accumulate context, and if so, is anything pruning it? A design that resends everything works fine in a demo of three turns and becomes unaffordable at thirty. Constrained hardware surfaces that flaw in an afternoon.

+ Part III

Your machine's envelope.

Three numbers describe what your hardware can do. Everything else — model choice, context settings, which capabilities are worth practising — follows from them.

09 — Chapter

Three numbers that decide everything.

NumberSetsHow to find it
Sustained memory bandwidthYour decode speed ceiling, for any model sizeChannels × transfer rate × 8 bytes; then assume 60–80% is achievable
Usable RAMThe largest model you can hold, and how much context you can affordTotal, minus operating system, minus whatever else must run
Parallel computePrefill speed — how long before the first token appearsCore count and vector instruction support; matters only for prefill

The asymmetry in that third row is the thing most people get wrong. A high core count feels like it should make everything faster. It makes prefill faster. It does close to nothing for decode, because decode is waiting on memory, not on arithmetic.

09 — From three numbers to a working envelope
BANDWIDTH GB per second USABLE RAM after the OS takes its share PARALLEL COMPUTE cores and vector width TOKENS PER SECOND bandwidth ÷ model size LARGEST MODEL RAM − KV cache − overhead TIME TO FIRST TOKEN for a given prompt size YOUR ENVELOPE which capabilities you can practise, which workflows are interactive, and which belong in the cloud Note the dashed line: RAM constrains model size, which then feeds back into your decode speed. The two are not independent.
Bandwidth Compute Capacity Derived limits The answer
Plate 09 — The envelope is not a judgment about your hardware. It is a specification, and knowing it precisely is what lets you choose work that will succeed rather than work that will disappoint you.
10 — Chapter

Deriving the ceiling before you download anything.

Work it in this order, on paper.

One — bandwidth. Channels × transfer rate × 8 bytes per transfer. Dual-channel DDR5-5600 gives 2 × 5600 × 8 = 89.6 GB/s theoretical. Take roughly 70% as realistically sustainable, so about 62 GB/s.

Two — memory budget. Subtract what the operating system and your other applications need. On a 32 GB machine, assume 6 to 8 GB is not yours. That leaves 24 to 26 GB for weights plus KV cache plus runtime.

Three — the speed you will accept. This is a judgment, not a calculation, and it should be made before you see any results. Roughly: above 10 tokens per second reads as conversational; 5 to 10 is usable with patience; below 3 is a batch job, not a chat.

Four — solve for model size. Divide bandwidth by your acceptable speed. Wanting 7 tokens per second at 62 GB/s means a model of about 8.9 GB or less — which lands you at a 13B-class model at 4-bit, or a 7B at higher precision.

What this produces Not a model recommendation — a size class, derived from your hardware and your patience. That class stays valid as models come and go, which is precisely why deriving it is worth more than any current list of names.
Work with Claude — build your envelope document Do this once, keep the output, and revisit it when your hardware changes rather than when the model landscape does.
Act as a systems engineer helping me characterise a machine
for local LLM inference. My hardware: [full specs].

Produce a one-page "inference envelope" covering:
- Sustained bandwidth estimate, with the arithmetic shown
- Usable RAM after OS overhead, stated as a range
- Maximum model file size for 5, 10 and 15 tokens/second
- Maximum practical context length at each of those sizes
- Two or three workload types this machine suits, and two it does not
- The three measurements that would confirm or refute your estimates

Where you are uncertain, say so and give a range rather than a point value.
Flag any assumption that, if wrong, would change the conclusion most.
That last instruction matters more than it looks. It surfaces the load-bearing assumption, which is the one to test first.
11 — Chapter

Measuring, rather than guessing.

Predictions are for calibration. Measurements are for decisions. Four numbers are worth capturing for every model you consider, and they take minutes to collect.

MetricWhat it tells youHow to read it
Load timeDisk and mapping costOne-off; ignore unless it is minutes
Time to first tokenPrefill cost at your prompt sizeMeasure at 500, 2K and 8K tokens — the curve matters more than any single point
Decode tokens/secondWhether you hit the bandwidth ceilingCompare against bandwidth ÷ model size. Far below means something is misconfigured
Peak resident memoryYour real headroomWatch it at your intended context, not at default

Record them in a table as you go. The table is the deliverable — not any individual number, and certainly not an impression.

Measure the second run The first run of any model includes disk reads, page faults, and cache warming. Run each measurement at least twice and use the second. A great many "this model is slow" conclusions are actually first-run artefacts, and the mistake is embarrassing precisely because it is so easy to avoid.

+ Part IV

Capability classes.

Model names expire. The categories they fall into have been stable for years and will likely outlast anything named in this document. Learning the taxonomy is what lets you re-evaluate a landscape in an afternoon instead of starting over.

12 — Chapter

A taxonomy that outlives the names.

Six classes cover essentially everything you would run locally. Each has a characteristic size, a characteristic cost profile, and a characteristic failure mode.

ClassDoesCostsFails by
General instructionChat, drafting, summarising, explanationModerate; scales with quality wantedConfident fluent wrongness on facts
ReasoningMaths, logic, multi-step derivationHigh — generates visible working before answeringLong confident chains to a wrong answer
CodeGeneration, explanation, review, translationLow to moderate at small sizesPlausible code that does not run
Vision-languageOCR, charts, screenshots, document imagesHigh per image; image tokens are expensiveInventing plausible text in unclear regions
EmbeddingTurns text into vectors for retrievalNegligibleRetrieving semantically near but useless passages
Small / edgeClassification, extraction, routingVery lowBreaking on anything outside its narrow job

Two observations about that table are worth more than the table itself.

The failure modes differ more than the capabilities do. Every class produces fluent output. What distinguishes them in practice is how they go wrong, and evaluating a model means learning to recognise its characteristic failure — not confirming it can produce something that looks right.

The embedding class is the one people skip and should not. It is tiny, nearly free to run, and it is the correct answer to "I have far more text than fits in a context window." Pushing a large corpus through a language model on constrained hardware is measured in days. Embedding it once and retrieving only relevant fragments is measured in seconds. Retrieval is not a lesser technique used because you lack resources; it is the right architecture, which constrained hardware forces you to learn properly.

10 — Capability classes against a constrained envelope
MEMORY FOOTPRINT → INTERACTIVE ←→ BATCH below this line, it stops feeling like a conversation EMBEDDING SMALL / EDGE CODE (SMALL) GENERAL (FAST) VISION REASONING GENERAL (BEST) Positions are indicative for a bandwidth-constrained machine. On a GPU box the whole cloud shifts upward and left.
Comfortable Workable Demanding Batch territory Interactivity line
Plate 10 — Classes below the dashed line are still useful; they are simply not conversational. Knowing which side of that line a task sits on before you start is most of the skill.
13 — Chapter

Matching class to constraint.

The selection rule, once you have an envelope, is short.

Start from the task, not the model. Write down what the output must be and how you will know it is correct. Only then ask which class produces that kind of output.

Pick the smallest model in that class that passes your acceptance test. Not the largest that fits. The difference is the whole discipline: a larger model that barely fits leaves no room for context, runs slowly enough to discourage iteration, and often does not measurably outperform a smaller one on a narrow task.

Match the interaction pattern to the phase profile. Single-shot work with a short prompt suits a bandwidth-bound machine well. Iterative work over a large accumulated context does not, for the reasons in chapter 8.

Reach for retrieval before reaching for a bigger model. If the problem is "too much text," the answer is almost never more parameters.

Transfer to product engineering This is requirements engineering with the names changed. Define the acceptance criterion first, choose the cheapest component that meets it, and verify against the criterion rather than against a datasheet. Engineers who have specified a sensor or a motor already know this discipline — the novelty in AI work is not the method, it is that the specification is behavioural rather than numeric, so you must build the test before you can trust the choice.
14 — Chapter

A snapshot, clearly dated.

Everything above is durable. This chapter is not, and it is separated deliberately so you can see where the expiry sits.

Valid as at August 2026 The specific families below were current when this was written. Treat them as worked examples of the classes in chapter 12, not as recommendations. If you are reading this more than six months after that date, use the taxonomy and re-derive the list — that exercise takes an afternoon and is more valuable than any list.

As of writing, open-weight families in wide local use include the Qwen series (strong dense models across general, code, and vision variants, at sizes from under a billion to tens of billions), the DeepSeek series (notable for mixture-of-experts coders and for distilled reasoning models), Google's Gemma line (efficient general models with vision), Meta's Llama line, and Mistral's small and medium models. Embedding models are a separate and slower-moving market, where a few hundred megabytes buys production-grade retrieval.

What matters for your purposes is the pattern rather than the roster: each family publishes several sizes of the same architecture, and the size you can run is set by your envelope, not by which family is currently best. Choosing a family is a minor decision. Choosing a size class is the one that determines whether the thing works on your desk.

Work with Claude — refresh this list yourself This is the prompt that makes the chapter self-renewing. Run it whenever you suspect the landscape has moved.
Search for the current state of open-weight language models
available for local inference.

For each of these classes — general instruction, reasoning, code,
vision-language, embedding, small/edge — tell me:
1. Which two or three families are currently well regarded
2. The size variants each publishes
3. Anything that changed materially in the last six months

Constraints: I can run models up to [X] GB and want at least
[Y] tokens/second. Filter to what fits.

Cite sources with dates. Where the evidence is thin or contested,
say so rather than picking a winner.

+ Part V

Evaluating without deceiving yourself.

This part is the one that transfers most directly into professional work. Choosing a model is easy. Knowing whether it is good enough for your purpose, and proving it to someone else, is the actual skill.

15 — Chapter

Why published benchmarks mislead you specifically.

Leaderboards are not dishonest. They are answering a different question from yours, and four gaps separate the two.

The precision gap. Published scores are almost always for full-precision weights. You are running a 4-bit quantization. As chapter 2 established, the loss is small on average but not uniform across capabilities — so a ranking established at full precision may not survive quantization in the specific dimension you care about.

The regime gap. This is the most consequential and the least discussed. Performance claims are measured on hardware with roughly ten times your memory bandwidth. An architectural advantage that is real there may vanish on your machine — chapter 20 documents a concrete case of exactly this.

The task gap. A benchmark measures average performance across a broad distribution. You have one narrow task. A model that ranks third overall may rank first on your task, and nothing in the leaderboard tells you which.

The contamination gap. Widely published benchmark sets leak into training data over time. Scores drift upward for reasons unrelated to capability.

The correct use of a leaderboard It is a shortlisting tool, not a decision tool. Use it to narrow a field of forty to a field of four. Then run your own evaluation on your own hardware against your own task, because that is the only measurement that answers your question.
16 — Chapter

Building a harness that gives the same answer twice.

An evaluation is worth having only if it is repeatable. Four properties make the difference between a harness and an impression.

Fixed prompts, written before you see any output. Write the whole set first. Prompts adjusted after seeing results measure your prompting, not the model.

Written acceptance criteria, per prompt. Not "is this good" but "does the output contain X, avoid Y, and satisfy Z." If you cannot state the criterion in advance, you do not yet know what you are testing.

Controlled sampling. Temperature and seed fixed. A model that produces different output each run cannot be compared against another model — you are measuring noise.

Recorded conditions. Quantization level, context setting, thread count, and the four metrics from chapter 11. A result without its conditions is not a result.

11 — The evaluation loop, and where it usually breaks
DEFINE THE TASK what must the output do WRITE PROMPTS all of them, first WRITE CRITERIA before seeing output RUN, FIXED SEED record conditions SCORE AGAINST THE CRITERIA next model THE THREE PLACES THIS LOOP BREAKS 1 · Criteria written after seeing output — you will unconsciously write criteria the output satisfies 2 · Prompt tweaked between models — you are now comparing two different tests 3 · One sample per prompt — sampling noise mistaken for a capability difference
Preparation Execution The step people skip
Plate 11 — The orange box is the one that gets skipped, and skipping it invalidates everything downstream. Criteria written after seeing output are not criteria; they are rationalisations.
Work with Claude — generate the eval set, then have it graded Two prompts. The first builds the harness; the second removes you from the scoring loop, which is where bias enters.
PROMPT 1 — BUILD THE SET

I need an evaluation set for [task] that I will run against several
local models. Produce 10 prompts that:
- vary in difficulty, from clearly easy to genuinely demanding
- each have a written, checkable acceptance criterion
- include at least two where the correct response is to refuse,
  say "not enough information", or flag a false premise

Output as a table: prompt | criterion | what a failure looks like.
Do not tell me which model to use.
PROMPT 2 — SCORE THE OUTPUT

Here is the criterion: [paste]
Here is the model's output: [paste]

Score pass or fail against the criterion only. Ignore style, tone,
and how confident it sounds. If it is borderline, say borderline and
state precisely which part of the criterion is unmet.

Do not be generous.
The instruction to include prompts where refusal is correct is the important one. A model that never says "I don't know" scores well on easy sets and fails in production.
17 — Chapter

The ways you will fool yourself.

Model evaluation is unusually vulnerable to self-deception, because outputs are fluent, subjective, and arrive one at a time. Five specific traps, each of which has a countermeasure.

TrapWhat happensCountermeasure
Fluency biasWell-written wrong answers score higher than awkward right onesScore correctness before reading for style; separate the two passes
AnchoringThe first model tested becomes the standard everything is judged againstRandomise order; score blind where you can
Cherry-pickingThe good run is remembered, the three bad ones are notRecord every run before scoring any
Prompt driftPrompts improve during testing, favouring whatever ran lastFreeze prompts; a change means restarting the whole set
Effort justificationA model that took an hour to set up gets graded generouslyWrite criteria before installing anything

The last one deserves emphasis because it is the least obvious and the hardest to resist. Having spent an evening getting a model running, you will want it to have been worth it. That desire is not a character flaw; it is a predictable bias, and the remedy is procedural rather than moral — decide what counts as success before you invest the effort, and the bias has nothing to act on.

The test that matters most Include tasks whose correct answer is "I don't know," "that is not in the document," or "the premise of your question is wrong." A model's willingness to decline is the single most important property for anything you intend to build on, and it is invisible in almost every benchmark. It is also the property most degraded by aggressive quantization, so test it at the precision you will actually deploy.
18 — Chapter

Thirty tasks, six classes.

What follows is not a curriculum to complete. It is a representative set — five tasks per capability class, spanning easy to demanding — designed so that working through it teaches you the failure modes of each class. Each carries an acceptance criterion, because a task without one teaches nothing.

General instruction

  1. Summarise a pasted article in two sentences. — No claim appears that is not in the source.
  2. Rewrite a paragraph for a non-expert reader. — Meaning preserved; no new jargon introduced.
  3. Convert rough notes into a structured message. — Every note represented; nothing invented.
  4. Summarise a transcript into action items. — All real actions captured, no imagined ones.
  5. Adapt one message for three audiences. — Register genuinely differs; substance identical across all three.

Reasoning

  1. Reverse a percentage discount. — Correct figure with working shown.
  2. A two-constraint logic puzzle. — Deduction chain valid at every step, not just the conclusion.
  3. A conditional probability question. — Correct answer and correct explanation of why intuition misleads.
  4. A multi-step word problem mixing rates and ratios. — Each step justified; arithmetic correct throughout.
  5. A problem with insufficient information to solve. — Says so, rather than inventing a value.

Code

  1. Write a small function with edge-case handling. — Runs; handles the stated edges.
  2. Explain an error message you paste. — Names the actual cause, not a generic one.
  3. Find the bug in a short broken function. — Locates the real defect; the fix works.
  4. Write tests for a function you provide. — Tests fail on a deliberately broken version.
  5. Review a module and prioritise the findings. — Findings are real and ordered by actual severity.

Vision-language

  1. Transcribe a clear screenshot. — Character-accurate.
  2. Extract a table image to structured data. — Rows and columns correct; parseable.
  3. Read values from a chart. — Within a stated tolerance; tolerance agreed in advance.
  4. Extract fields from a form or invoice to JSON. — Valid JSON; every field traceable to the image.
  5. Given a deliberately blurred or cropped region, report it. — Says the region is unreadable rather than guessing.

Retrieval

  1. Ask a question answered on one page of a document. — Correct and locatable in the source.
  2. Ask a question requiring two documents combined. — Synthesises correctly across both.
  3. Ask for a fact buried deep in a long document. — Retrieves the specific passage.
  4. Supply two conflicting sources and ask about the discrepancy. — Identifies the conflict; attributes each side.
  5. Ask something the corpus does not contain. — Says it is not there. This is the test that matters.

Small / edge

  1. Classify short texts into fixed categories. — Consistent labels across repeated runs.
  2. Extract structured fields from unstructured text. — Schema respected exactly.
  3. Route a request to one of five handlers. — Deterministic and correct.
  4. Handle an input outside every category. — Returns the fallback rather than forcing a match.
  5. Run the same input twenty times. — Identical output every time at temperature zero.
How to work the set Task five in each class is the one that discriminates. Tasks one to four confirm the model can do the obvious thing, which almost any current model can. The fifth tests whether it knows the limits of what it has been given — and that is the property that determines whether you can build on it.

+ Part VI

The case: one machine, eight models.

What follows is a real exercise, described without vendor identification because the specific machine does not matter — the shape of what it taught does.

19 — Chapter

The setup, and what was measured.

A desktop workstation: a recent 20-core consumer processor with 28 threads, 32 GB of dual-channel DDR5, integrated graphics with no CUDA support. Eight models were installed across six capability classes and worked through systematically over several weeks.

Derived envelope, using the method in chapter 10: roughly 60 GB/s sustainable bandwidth, about 25 GB usable after the operating system, which sets a practical ceiling around a 9 GB model for usable interactive speed.

Measured behaviour matched that prediction closely. Small models in the 4–6 GB range were conversational. Models around 9 GB were usable. A 15 GB model worked but demanded patience. Nothing in the measurements contradicted the arithmetic — which is itself the first useful finding, because it means the model of the machine was correct.

One measurement stood out and turned out to be the most instructive: an 11,000-token prompt took approximately five minutes before the first token appeared. Read against chapter 5, that is not a fault. It is prefill, compute-bound, on a machine without a GPU — behaving exactly as the two-phase model predicts.

Why that number decided everything else Five minutes of prefill per step makes agentic loops impossible and makes single-shot work perfectly pleasant. One measurement, correctly interpreted, determined which entire category of work the machine was suited for. That is what an envelope is for — and it took one measurement, not eight weeks of trial and error, once the model of the machine was in place.
20 — Chapter

Where received wisdom did not survive.

Here is the part of the exercise worth the whole exercise.

Two of the eight models used a mixture-of-experts architecture. The standard account of MoE goes like this: the model contains many specialised sub-networks, and a router activates only a small subset for each token. A model with 16 billion total parameters might activate only 2.4 billion per token: because only a fraction of the experts fire for any given token, the floating-point work per token resembles that of a far smaller dense network.

The natural conclusion — and the one stated in most write-ups, including the original notes for this exercise — is that MoE models run at roughly small-model speed while retaining large-model knowledge, making them ideal for constrained hardware.

On a bandwidth-bound machine, that conclusion does not hold.

A 2026 empirical study benchmarked an MoE model against dense baselines on consumer and edge hardware. On a laptop the sparse model trailed a dense one of equivalent active size by around a tenth, and on the more constrained device that shortfall widened to roughly a third while burning over twice the energy for each token produced. Instrumenting the decode path showed why: timing the decode path node by node put the router at under a tenth of the work inside the expert blocks — so the shortfall came from carrying every parameter in memory, dispatching between experts, and cache pressure, not from routing.

The conclusion the authors draw is the sentence to take away: where bandwidth is the binding limit, what you pay for is the full parameter count rather than the activated fraction, and sparsity returns nothing on the axis the device is actually short of.

12 — Where the mixture-of-experts argument breaks down
THE EXPECTATION — A FLOP ARGUMENT 16B TOTAL PARAMETERS 2.4B not activated this token "so it should run like a 2.4B model" THE REALITY — A BANDWIDTH ARGUMENT 16B TOTAL PARAMETERS ALL OF IT MUST BE RESIDENT AND REACHABLE the router may pick any expert, so none can be evicted Sparse activation reduces arithmetic. Arithmetic was not the constraint. Memory footprint and cache pressure were — and sparsity does nothing for either. The architecture claim was not wrong. It was measured in a regime where compute was the binding constraint. On hardware where bandwidth binds instead, the same design delivers a different result.
Compute saved Bandwidth unchanged Total weights The binding constraint
Plate 12 — Two correct arguments about the same model, reaching different practical conclusions because they optimise different resources. Which one applies to you is a property of your hardware, not of the model.
What this case is and is not It is one published study of one MoE model at one parameter scale on two devices, and the authors bound their claim accordingly. It is not proof that MoE never helps on consumer hardware. It is a demonstration that an architectural claim carries an implicit assumption about which resource is scarce — and that when your scarce resource differs, the claim may not transfer. That is the transferable lesson, and it applies well beyond this one architecture.
21 — Chapter

What the exercise actually taught.

Not which models are good. Those have already been superseded. What survived is a set of habits.

Predict, then measure, then explain the gap. Every time the prediction was wrong, an assumption was wrong, and finding it taught more than a correct prediction would have.

Read a claim's implicit regime. Every performance claim assumes something about what is scarce. Learning to ask "measured under what constraint" changed how every subsequent piece of advice was read.

Distinguish capability from interactivity. Several models could do things the machine could not do comfortably. Those are different findings and need recording separately.

Treat refusal as a first-class capability. Across every class, the tasks that discriminated between models were the ones where the correct answer was "not enough information."

Reach for architecture before scale. The retrieval setup — a 274 MB embedding model doing the work — answered "how do I handle far more text than fits" better than any larger model could have.

+ Part VII

Foundations for AI product engineering.

Everything above is transferable, and the transfer is the point. What a constrained machine teaches is not how to run models locally. It is how to reason about a system whose cost, capability, and failure modes are all coupled to a constraint you did not choose.

22 — Chapter

The envelope as a product requirement.

Every AI product has a deployment envelope, whether or not anyone has written it down. A cloud API hides it behind a bill; a phone or an embedded device makes it inescapable. In both cases it exists, and in both cases it determines what the product can be.

The discipline from Part III applies unchanged, with the terms renamed.

Local inferenceProduct engineering
Memory bandwidthLatency budget — what the user will wait for
Usable RAMCost ceiling per interaction
Model size classWhich tier of model the unit economics permit
Context length settingHow much history the design can afford to resend
Acceptance criteriaThe behavioural specification, and the regression suite

Writing this down before choosing a model is the same discipline as writing a mechanical specification before choosing a bearing. The failure mode is identical too: a component chosen from a datasheet, without a specification to test against, that works in the prototype and fails in the field.

13 — The same reasoning at three scales
YOUR WORKSTATION scarce: bandwidth budget: seconds of patience lever: model size test: your 30 tasks learning environment EMBEDDED DEVICE scarce: power and memory budget: milliwatts lever: quantization, task scope test: field conditions product CLOUD PRODUCT scarce: cost per call budget: unit economics lever: context discipline, tier test: regression suite product IDENTICAL METHOD · NAME THE SCARCE RESOURCE, SET A BUDGET, PICK THE LEVER, WRITE THE TEST only the scarce resource changes between columns
Learning Edge deployment Cloud deployment The common method
Plate 13 — The columns differ only in which resource is scarce. This is why a constrained workstation is a legitimate training ground for product work rather than a compromise you tolerate.
23 — Chapter

The evaluation set is the real deliverable.

In conventional software, tests are infrastructure supporting the product. In AI systems the relationship inverts, for one reason: the component you depend on will be replaced, repeatedly, by something you did not build and cannot inspect.

Models are deprecated. Providers change defaults. A quantization is re-released. A prompt that worked stops working. In every case the question is the same — is the system still doing its job? — and the only thing that answers it is a suite of tasks with written acceptance criteria.

An organisation with a good evaluation set can swap models in an afternoon. One without has to re-form an opinion from scratch every time, and will make that decision on impressions.

Transfer to regulated work Anyone who has worked under a quality system will recognise this immediately: it is design verification. You define intended use, derive acceptance criteria, and demonstrate the output meets them — with records. The novelty in AI systems is only that the specification is behavioural and probabilistic rather than dimensional, so acceptance must be defined statistically, over a set, rather than as a single pass or fail. The structure of the obligation is unchanged.
Work with Claude — convert an envelope into a specification This is the prompt that turns the whole exercise into an artefact someone else can act on.
I have characterised my hardware and evaluated several models
against a task. Here are my findings: [paste envelope + results table].

Draft a one-page component specification for this AI capability
as if it were any other engineered component:
- Intended use, stated narrowly
- Functional requirements with measurable acceptance criteria
- Operating envelope (latency, memory, context limits)
- Known failure modes and what each looks like in output
- Verification method — how a successor model would be qualified
- What would trigger re-qualification

Flag anything my findings do not actually support. I would rather
have gaps marked than filled with plausible text.
That final instruction is the one that makes the output usable. A specification with honest gaps is an engineering document; one with plausible filler is a liability.
24 — Chapter

What to carry forward.

Name the scarce resource first. Every performance claim, yours or anyone's, assumes something about what is scarce. Bandwidth, memory, cost per call, power, or patience. Identify it before you optimise, or you will optimise the wrong thing well.

Derive the envelope before choosing a component. A size class derived from your constraints outlives every model name, and it turns an open-ended search into a filter.

Predict, measure, then explain the difference. The gap between prediction and measurement is where the learning is. A correct prediction confirms your model; a wrong one improves it.

Write acceptance criteria before you build anything. Criteria written afterwards are rationalisations, and effort already spent will bias them.

Test for refusal, not just for capability. Whether a system knows the limits of what it has been given matters more than its ceiling, and it is nearly absent from published benchmarks.

Reach for architecture before scale. Retrieval, decomposition, and context economy solve problems that adding parameters does not.

Treat the evaluation set as the durable asset. Models are replaceable. The definition of what "working" means is the thing you actually own.

The through-line A constrained machine is not a compromised version of a better one. It is an instrument that makes visible the costs a well-resourced environment hides — and those costs are exactly what determines whether an AI product is affordable, reliable, and honest about its limits. Learning to work inside a tight envelope is not preparation for real engineering. It is the same discipline, made legible.
Sources

Where the numbers came from.

Technical claims in this guide are drawn from the following, checked in August 2026. Figures quoted are approximate and vary by model, harness, and hardware; verify against the primary source before relying on any specific number.

  • Prefill and decode bottlenecks. DigitalOcean, The Hidden Bottlenecks in LLM Inference (April 2026), and the TileLens memory-systems paper, arXiv:2607.04031.
  • CPU versus accelerator bandwidth. S. Tripathi, llama.cpp: CPU Inference for LLMs on Consumer Hardware (June 2026).
  • Memory channels over clock speed. Hardware Corner, Memory Bandwidth and Tokens per Second in Local LLM Inference (October 2025).
  • Quantization structure and the quality curve. Quantization in Practice: GPTQ vs AWQ vs GGUF, The AI Engineer (May 2026); community perplexity comparisons across K-quant levels (2026).
  • Mixture-of-experts on constrained hardware. Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study, arXiv:2606.21428 (July 2026). The authors bound their findings to one model at one scale on two devices; this guide does the same.