The machine
decides.
+ Before we start
Almost everyone learning to build with AI starts at the API. That teaches you prompting. It does not teach you what a model is, what it costs, or what happens when the budget is not someone else's.
Running models on your own hardware teaches all three, quickly and unforgivingly. A cloud endpoint hides the weights, the memory, the arithmetic, and the bill. A machine on your desk hides nothing. When a model will not load, you learn what it weighs. When generation crawls, you learn which resource you ran out of. When a 30-billion-parameter model answers worse than a 7-billion one on your actual task, you learn that size is not capability.
This guide is about that education. It is not a tutorial for one stack, and it is deliberately light on model names, because those expire. It is about the physics underneath, which does not.
The claim it argues: for any given machine, a small number of hardware facts determine which AI capabilities you can realistically develop on it — and most published advice about models is derived in a regime that is not yours. The remedy is not better advice. It is measurement, on your hardware, against your tasks.
Claude appears throughout as a working tool. The dotted blue boxes contain prompts and workflows you can lift directly — for deriving your machine's limits, designing an evaluation set, and interpreting results without deceiving yourself.
+ Part I
What a model actually is.
Before anything about speed or capability, it helps to know precisely what you are downloading and what the machine does with it. Almost every confusion later resolves into a misunderstanding here.
A file of numbers, and a loop that reads it whole.
A language model is a file containing a very large list of numbers called parameters, plus a small description of how they are arranged into layers. Nothing in the file is text. Nothing in it is knowledge in any form you could inspect. It is a matrix of learned weights.
Generating text works like this. Your prompt is cut into tokens — roughly word fragments. The tokens are turned into vectors, and those vectors are pushed through every layer of the network, multiplied against the weights at each step. Out the other end comes a probability distribution over the vocabulary: how likely each possible next token is. One token is chosen, appended to the sequence, and the entire process runs again from the top.
That last sentence is the one to hold on to. Producing a 500-token answer means running the full network 500 times. There is no way to generate a sentence in one pass. The model has no memory of its own computation between steps — only the growing sequence of tokens it has already produced.
Quantization: why four bits is nearly free and three is not.
Models are trained with weights stored as 16-bit or 32-bit floating-point numbers. At 16 bits, a 7-billion-parameter model is about 14 GB, and a 70-billion one about 140 GB. On consumer hardware that is the end of the conversation.
Quantization is the compression that makes local inference possible at all. The idea is simple: instead of storing each weight at full precision, divide the range of values into a small number of buckets and store the bucket index. At 4 bits you have sixteen buckets per weight rather than 65,536 levels.
Naively this should be destructive, and done naively it is. The reason it works in practice is that modern formats do not use one bucket scheme for the whole model. Each weight matrix is divided into groups of 256, and every group subdivided again into eights of 32, with a scaling factor recorded at both levels — so the sixteen available bucket values are re-fitted to each small neighbourhood rather than stretched across the whole matrix. The bookkeeping for all those scaling factors adds roughly half a bit to every weight, which is why a nominally 4-bit format occupies closer to 4.5 — and the accuracy recovered by fitting locally exceeds what that extra storage costs.
The result is a quality curve that is not at all what intuition suggests.
Two refinements are worth knowing, because they turn this into an actual decision rule rather than a default.
the damage is uneven across skills: logical reasoning holds up well under compression, whereas numerical accuracy starts slipping once you drop below four bits. If your application does exact numeric work, spend the memory. If it does classification, summarisation, or discussion, do not.
And format choice interacts with your hardware. The more aggressively compressed variants trade decode speed for file size on processors, so the block-scaled family often generates faster on consumer machines despite the larger download. On a CPU box, the smallest file is often not the fastest.
I have [RAM] of system memory, [channels]-channel [DDR4/DDR5-speed] memory, and no CUDA GPU. I want to run a [size]B parameter model for [task]. Work out, showing your arithmetic: 1. Approximate file size at Q4_K_M, Q5_K_M and Q8_0 2. How much RAM remains for KV cache after each 3. Which of these my machine can actually hold at [context] tokens 4. Whether my task is quantization-sensitive, and why Do not recommend a specific model. I want the envelope, not a pick.Asking for the envelope rather than a recommendation is deliberate — you want the constraint reasoning to stay yours, and it does not expire when models change.
Inside the container, and why loading is instant.
Quantized weights ship in a container format that bundles the weights, the quantization scales, the tokenizer, and the architecture metadata into a single file. Older schemes split these across several files, which made distribution fragile.
One property of that packaging has an outsized effect on how local inference feels. Startup feels instantaneous because the file is mapped into the address space rather than read into it. Memory-mapping means pages are pulled from disk on demand and shared, so a model that is nominally 9 GB does not require a 9 GB read before the first token.
This is why a model can appear to load in a second and then be slow on its first answer: the loading you expected was deferred, not avoided.
What actually occupies your memory.
People size their machine against the file on disk. That number is always too small, and the gap is where most failed loads come from.
Three things consume RAM during inference.
The weights. Roughly the file size, since the file is mapped rather than expanded.
The KV cache. Every token processed produces intermediate key and value vectors at every layer, and these are kept so that subsequent tokens do not have to recompute the whole history. The cache grows linearly with context length and is not part of the file size anyone quotes. At long context it can rival the weights themselves.
Runtime overhead. Activations, buffers, the framework, the operating system. Small but not zero.
+ Part II
The two-phase machine.
This part contains the single most useful idea in the guide. Inference is not one workload — it is two, with opposite bottlenecks. Once you see the split, most puzzling behaviour on constrained hardware stops being puzzling.
Prefill and decode: one workload, two bottlenecks.
When you send a prompt, the machine does two very different jobs in sequence.
Prefill processes your entire prompt at once. Every token in it can be handled in parallel, because they are all already known. This is dense matrix work across many tokens simultaneously, and it saturates arithmetic units. Prefill is compute-bound.
Decode generates the answer one token at a time. Each token depends on the one before it, so there is nothing to parallelise. Each step reads the whole weight file to produce a single token. Decode is bandwidth-bound.
the arithmetic load concentrates almost entirely in prefill, while decode performs comparatively little computation and instead spends its time moving a great deal of data, and each weight is used exactly once before the next token requires it to be read again, and that single-use pattern is what pins latency to bandwidth rather than to processing power.
The bandwidth wall, in arithmetic.
You can predict your machine's generation speed to within a small factor without running anything. The calculation takes one line, and it is the most useful thing in this guide.
Producing one token requires reading every weight once. So:
tokens per second ≈ memory bandwidth ÷ model size in memory
A 9 GB model on a machine that can sustain 60 GB/s gives about 6.7 tokens per second as a theoretical ceiling. Real systems reach perhaps 60–80% of theoretical bandwidth, so expect 4 to 5. If you measure 4.5, your system is working correctly — that is not a configuration problem to solve.
Now the uncomfortable comparison. a consumer memory subsystem supplies something on the order of 100 GB/s, where datacentre accelerator memory reaches roughly ten times that. That order of magnitude is the difference between local and cloud inference. It is not driver tuning, thread count, or software maturity.
My machine: [CPU], [RAM amount and type], [memory channels], [GPU or none]. Before I run anything, help me predict: 1. My approximate sustained memory bandwidth, and how you derived it 2. Expected decode speed for a 4 GB, 9 GB and 15 GB model 3. The largest model size that would still give me above 5 tokens/second 4. What measurement I should take to check whether your estimate was right State your assumptions explicitly so I can tell which one was wrong if the measurement disagrees.Then measure. A prediction that misses by 3× usually means one assumption was badly off, and finding which one teaches you more than a correct answer would have.
Why long context costs you twice.
Context length is sold as a capability — "128K context" sounds like a feature you either have or lack. On constrained hardware it is better understood as a bill that arrives twice.
The first charge is memory. The KV cache grows linearly with the number of tokens held, as shown in plate 04. Long context consumes RAM that is then unavailable for weights.
The second charge is time. Prefill cost scales with prompt length, and the attention component scales worse than linearly. A prompt ten times longer takes more than ten times as long to process.
This is why pasting a large document into a local model produces a pause that feels disproportionate. It is disproportionate — that is the arithmetic working as designed.
Why autonomous agents stall on constrained hardware.
Everything so far converges on one practical conclusion, and it is the most valuable thing a constrained machine teaches you.
An agentic loop — a model that calls tools, reads results, and decides what to do next — works by resending the accumulated conversation on every step. Step one has a small prompt. By step ten the prompt contains the original instruction, ten tool calls, and ten sets of results.
Every step re-pays the prefill cost for the entire accumulated history. Not the new part. All of it.
+ Part III
Your machine's envelope.
Three numbers describe what your hardware can do. Everything else — model choice, context settings, which capabilities are worth practising — follows from them.
Three numbers that decide everything.
| Number | Sets | How to find it |
|---|---|---|
| Sustained memory bandwidth | Your decode speed ceiling, for any model size | Channels × transfer rate × 8 bytes; then assume 60–80% is achievable |
| Usable RAM | The largest model you can hold, and how much context you can afford | Total, minus operating system, minus whatever else must run |
| Parallel compute | Prefill speed — how long before the first token appears | Core count and vector instruction support; matters only for prefill |
The asymmetry in that third row is the thing most people get wrong. A high core count feels like it should make everything faster. It makes prefill faster. It does close to nothing for decode, because decode is waiting on memory, not on arithmetic.
Deriving the ceiling before you download anything.
Work it in this order, on paper.
One — bandwidth. Channels × transfer rate × 8 bytes per transfer. Dual-channel DDR5-5600 gives 2 × 5600 × 8 = 89.6 GB/s theoretical. Take roughly 70% as realistically sustainable, so about 62 GB/s.
Two — memory budget. Subtract what the operating system and your other applications need. On a 32 GB machine, assume 6 to 8 GB is not yours. That leaves 24 to 26 GB for weights plus KV cache plus runtime.
Three — the speed you will accept. This is a judgment, not a calculation, and it should be made before you see any results. Roughly: above 10 tokens per second reads as conversational; 5 to 10 is usable with patience; below 3 is a batch job, not a chat.
Four — solve for model size. Divide bandwidth by your acceptable speed. Wanting 7 tokens per second at 62 GB/s means a model of about 8.9 GB or less — which lands you at a 13B-class model at 4-bit, or a 7B at higher precision.
Act as a systems engineer helping me characterise a machine for local LLM inference. My hardware: [full specs]. Produce a one-page "inference envelope" covering: - Sustained bandwidth estimate, with the arithmetic shown - Usable RAM after OS overhead, stated as a range - Maximum model file size for 5, 10 and 15 tokens/second - Maximum practical context length at each of those sizes - Two or three workload types this machine suits, and two it does not - The three measurements that would confirm or refute your estimates Where you are uncertain, say so and give a range rather than a point value. Flag any assumption that, if wrong, would change the conclusion most.That last instruction matters more than it looks. It surfaces the load-bearing assumption, which is the one to test first.
Measuring, rather than guessing.
Predictions are for calibration. Measurements are for decisions. Four numbers are worth capturing for every model you consider, and they take minutes to collect.
| Metric | What it tells you | How to read it |
|---|---|---|
| Load time | Disk and mapping cost | One-off; ignore unless it is minutes |
| Time to first token | Prefill cost at your prompt size | Measure at 500, 2K and 8K tokens — the curve matters more than any single point |
| Decode tokens/second | Whether you hit the bandwidth ceiling | Compare against bandwidth ÷ model size. Far below means something is misconfigured |
| Peak resident memory | Your real headroom | Watch it at your intended context, not at default |
Record them in a table as you go. The table is the deliverable — not any individual number, and certainly not an impression.
+ Part IV
Capability classes.
Model names expire. The categories they fall into have been stable for years and will likely outlast anything named in this document. Learning the taxonomy is what lets you re-evaluate a landscape in an afternoon instead of starting over.
A taxonomy that outlives the names.
Six classes cover essentially everything you would run locally. Each has a characteristic size, a characteristic cost profile, and a characteristic failure mode.
| Class | Does | Costs | Fails by |
|---|---|---|---|
| General instruction | Chat, drafting, summarising, explanation | Moderate; scales with quality wanted | Confident fluent wrongness on facts |
| Reasoning | Maths, logic, multi-step derivation | High — generates visible working before answering | Long confident chains to a wrong answer |
| Code | Generation, explanation, review, translation | Low to moderate at small sizes | Plausible code that does not run |
| Vision-language | OCR, charts, screenshots, document images | High per image; image tokens are expensive | Inventing plausible text in unclear regions |
| Embedding | Turns text into vectors for retrieval | Negligible | Retrieving semantically near but useless passages |
| Small / edge | Classification, extraction, routing | Very low | Breaking on anything outside its narrow job |
Two observations about that table are worth more than the table itself.
The failure modes differ more than the capabilities do. Every class produces fluent output. What distinguishes them in practice is how they go wrong, and evaluating a model means learning to recognise its characteristic failure — not confirming it can produce something that looks right.
The embedding class is the one people skip and should not. It is tiny, nearly free to run, and it is the correct answer to "I have far more text than fits in a context window." Pushing a large corpus through a language model on constrained hardware is measured in days. Embedding it once and retrieving only relevant fragments is measured in seconds. Retrieval is not a lesser technique used because you lack resources; it is the right architecture, which constrained hardware forces you to learn properly.
Matching class to constraint.
The selection rule, once you have an envelope, is short.
Start from the task, not the model. Write down what the output must be and how you will know it is correct. Only then ask which class produces that kind of output.
Pick the smallest model in that class that passes your acceptance test. Not the largest that fits. The difference is the whole discipline: a larger model that barely fits leaves no room for context, runs slowly enough to discourage iteration, and often does not measurably outperform a smaller one on a narrow task.
Match the interaction pattern to the phase profile. Single-shot work with a short prompt suits a bandwidth-bound machine well. Iterative work over a large accumulated context does not, for the reasons in chapter 8.
Reach for retrieval before reaching for a bigger model. If the problem is "too much text," the answer is almost never more parameters.
A snapshot, clearly dated.
Everything above is durable. This chapter is not, and it is separated deliberately so you can see where the expiry sits.
As of writing, open-weight families in wide local use include the Qwen series (strong dense models across general, code, and vision variants, at sizes from under a billion to tens of billions), the DeepSeek series (notable for mixture-of-experts coders and for distilled reasoning models), Google's Gemma line (efficient general models with vision), Meta's Llama line, and Mistral's small and medium models. Embedding models are a separate and slower-moving market, where a few hundred megabytes buys production-grade retrieval.
What matters for your purposes is the pattern rather than the roster: each family publishes several sizes of the same architecture, and the size you can run is set by your envelope, not by which family is currently best. Choosing a family is a minor decision. Choosing a size class is the one that determines whether the thing works on your desk.
Search for the current state of open-weight language models available for local inference. For each of these classes — general instruction, reasoning, code, vision-language, embedding, small/edge — tell me: 1. Which two or three families are currently well regarded 2. The size variants each publishes 3. Anything that changed materially in the last six months Constraints: I can run models up to [X] GB and want at least [Y] tokens/second. Filter to what fits. Cite sources with dates. Where the evidence is thin or contested, say so rather than picking a winner.
+ Part V
Evaluating without deceiving yourself.
This part is the one that transfers most directly into professional work. Choosing a model is easy. Knowing whether it is good enough for your purpose, and proving it to someone else, is the actual skill.
Why published benchmarks mislead you specifically.
Leaderboards are not dishonest. They are answering a different question from yours, and four gaps separate the two.
The precision gap. Published scores are almost always for full-precision weights. You are running a 4-bit quantization. As chapter 2 established, the loss is small on average but not uniform across capabilities — so a ranking established at full precision may not survive quantization in the specific dimension you care about.
The regime gap. This is the most consequential and the least discussed. Performance claims are measured on hardware with roughly ten times your memory bandwidth. An architectural advantage that is real there may vanish on your machine — chapter 20 documents a concrete case of exactly this.
The task gap. A benchmark measures average performance across a broad distribution. You have one narrow task. A model that ranks third overall may rank first on your task, and nothing in the leaderboard tells you which.
The contamination gap. Widely published benchmark sets leak into training data over time. Scores drift upward for reasons unrelated to capability.
Building a harness that gives the same answer twice.
An evaluation is worth having only if it is repeatable. Four properties make the difference between a harness and an impression.
Fixed prompts, written before you see any output. Write the whole set first. Prompts adjusted after seeing results measure your prompting, not the model.
Written acceptance criteria, per prompt. Not "is this good" but "does the output contain X, avoid Y, and satisfy Z." If you cannot state the criterion in advance, you do not yet know what you are testing.
Controlled sampling. Temperature and seed fixed. A model that produces different output each run cannot be compared against another model — you are measuring noise.
Recorded conditions. Quantization level, context setting, thread count, and the four metrics from chapter 11. A result without its conditions is not a result.
PROMPT 1 — BUILD THE SET I need an evaluation set for [task] that I will run against several local models. Produce 10 prompts that: - vary in difficulty, from clearly easy to genuinely demanding - each have a written, checkable acceptance criterion - include at least two where the correct response is to refuse, say "not enough information", or flag a false premise Output as a table: prompt | criterion | what a failure looks like. Do not tell me which model to use.
PROMPT 2 — SCORE THE OUTPUT Here is the criterion: [paste] Here is the model's output: [paste] Score pass or fail against the criterion only. Ignore style, tone, and how confident it sounds. If it is borderline, say borderline and state precisely which part of the criterion is unmet. Do not be generous.The instruction to include prompts where refusal is correct is the important one. A model that never says "I don't know" scores well on easy sets and fails in production.
The ways you will fool yourself.
Model evaluation is unusually vulnerable to self-deception, because outputs are fluent, subjective, and arrive one at a time. Five specific traps, each of which has a countermeasure.
| Trap | What happens | Countermeasure |
|---|---|---|
| Fluency bias | Well-written wrong answers score higher than awkward right ones | Score correctness before reading for style; separate the two passes |
| Anchoring | The first model tested becomes the standard everything is judged against | Randomise order; score blind where you can |
| Cherry-picking | The good run is remembered, the three bad ones are not | Record every run before scoring any |
| Prompt drift | Prompts improve during testing, favouring whatever ran last | Freeze prompts; a change means restarting the whole set |
| Effort justification | A model that took an hour to set up gets graded generously | Write criteria before installing anything |
The last one deserves emphasis because it is the least obvious and the hardest to resist. Having spent an evening getting a model running, you will want it to have been worth it. That desire is not a character flaw; it is a predictable bias, and the remedy is procedural rather than moral — decide what counts as success before you invest the effort, and the bias has nothing to act on.
Thirty tasks, six classes.
What follows is not a curriculum to complete. It is a representative set — five tasks per capability class, spanning easy to demanding — designed so that working through it teaches you the failure modes of each class. Each carries an acceptance criterion, because a task without one teaches nothing.
General instruction
- Summarise a pasted article in two sentences. — No claim appears that is not in the source.
- Rewrite a paragraph for a non-expert reader. — Meaning preserved; no new jargon introduced.
- Convert rough notes into a structured message. — Every note represented; nothing invented.
- Summarise a transcript into action items. — All real actions captured, no imagined ones.
- Adapt one message for three audiences. — Register genuinely differs; substance identical across all three.
Reasoning
- Reverse a percentage discount. — Correct figure with working shown.
- A two-constraint logic puzzle. — Deduction chain valid at every step, not just the conclusion.
- A conditional probability question. — Correct answer and correct explanation of why intuition misleads.
- A multi-step word problem mixing rates and ratios. — Each step justified; arithmetic correct throughout.
- A problem with insufficient information to solve. — Says so, rather than inventing a value.
Code
- Write a small function with edge-case handling. — Runs; handles the stated edges.
- Explain an error message you paste. — Names the actual cause, not a generic one.
- Find the bug in a short broken function. — Locates the real defect; the fix works.
- Write tests for a function you provide. — Tests fail on a deliberately broken version.
- Review a module and prioritise the findings. — Findings are real and ordered by actual severity.
Vision-language
- Transcribe a clear screenshot. — Character-accurate.
- Extract a table image to structured data. — Rows and columns correct; parseable.
- Read values from a chart. — Within a stated tolerance; tolerance agreed in advance.
- Extract fields from a form or invoice to JSON. — Valid JSON; every field traceable to the image.
- Given a deliberately blurred or cropped region, report it. — Says the region is unreadable rather than guessing.
Retrieval
- Ask a question answered on one page of a document. — Correct and locatable in the source.
- Ask a question requiring two documents combined. — Synthesises correctly across both.
- Ask for a fact buried deep in a long document. — Retrieves the specific passage.
- Supply two conflicting sources and ask about the discrepancy. — Identifies the conflict; attributes each side.
- Ask something the corpus does not contain. — Says it is not there. This is the test that matters.
Small / edge
- Classify short texts into fixed categories. — Consistent labels across repeated runs.
- Extract structured fields from unstructured text. — Schema respected exactly.
- Route a request to one of five handlers. — Deterministic and correct.
- Handle an input outside every category. — Returns the fallback rather than forcing a match.
- Run the same input twenty times. — Identical output every time at temperature zero.
+ Part VI
The case: one machine, eight models.
What follows is a real exercise, described without vendor identification because the specific machine does not matter — the shape of what it taught does.
The setup, and what was measured.
A desktop workstation: a recent 20-core consumer processor with 28 threads, 32 GB of dual-channel DDR5, integrated graphics with no CUDA support. Eight models were installed across six capability classes and worked through systematically over several weeks.
Derived envelope, using the method in chapter 10: roughly 60 GB/s sustainable bandwidth, about 25 GB usable after the operating system, which sets a practical ceiling around a 9 GB model for usable interactive speed.
Measured behaviour matched that prediction closely. Small models in the 4–6 GB range were conversational. Models around 9 GB were usable. A 15 GB model worked but demanded patience. Nothing in the measurements contradicted the arithmetic — which is itself the first useful finding, because it means the model of the machine was correct.
One measurement stood out and turned out to be the most instructive: an 11,000-token prompt took approximately five minutes before the first token appeared. Read against chapter 5, that is not a fault. It is prefill, compute-bound, on a machine without a GPU — behaving exactly as the two-phase model predicts.
Where received wisdom did not survive.
Here is the part of the exercise worth the whole exercise.
Two of the eight models used a mixture-of-experts architecture. The standard account of MoE goes like this: the model contains many specialised sub-networks, and a router activates only a small subset for each token. A model with 16 billion total parameters might activate only 2.4 billion per token: because only a fraction of the experts fire for any given token, the floating-point work per token resembles that of a far smaller dense network.
The natural conclusion — and the one stated in most write-ups, including the original notes for this exercise — is that MoE models run at roughly small-model speed while retaining large-model knowledge, making them ideal for constrained hardware.
On a bandwidth-bound machine, that conclusion does not hold.
A 2026 empirical study benchmarked an MoE model against dense baselines on consumer and edge hardware. On a laptop the sparse model trailed a dense one of equivalent active size by around a tenth, and on the more constrained device that shortfall widened to roughly a third while burning over twice the energy for each token produced. Instrumenting the decode path showed why: timing the decode path node by node put the router at under a tenth of the work inside the expert blocks — so the shortfall came from carrying every parameter in memory, dispatching between experts, and cache pressure, not from routing.
The conclusion the authors draw is the sentence to take away: where bandwidth is the binding limit, what you pay for is the full parameter count rather than the activated fraction, and sparsity returns nothing on the axis the device is actually short of.
What the exercise actually taught.
Not which models are good. Those have already been superseded. What survived is a set of habits.
Predict, then measure, then explain the gap. Every time the prediction was wrong, an assumption was wrong, and finding it taught more than a correct prediction would have.
Read a claim's implicit regime. Every performance claim assumes something about what is scarce. Learning to ask "measured under what constraint" changed how every subsequent piece of advice was read.
Distinguish capability from interactivity. Several models could do things the machine could not do comfortably. Those are different findings and need recording separately.
Treat refusal as a first-class capability. Across every class, the tasks that discriminated between models were the ones where the correct answer was "not enough information."
Reach for architecture before scale. The retrieval setup — a 274 MB embedding model doing the work — answered "how do I handle far more text than fits" better than any larger model could have.
+ Part VII
Foundations for AI product engineering.
Everything above is transferable, and the transfer is the point. What a constrained machine teaches is not how to run models locally. It is how to reason about a system whose cost, capability, and failure modes are all coupled to a constraint you did not choose.
The envelope as a product requirement.
Every AI product has a deployment envelope, whether or not anyone has written it down. A cloud API hides it behind a bill; a phone or an embedded device makes it inescapable. In both cases it exists, and in both cases it determines what the product can be.
The discipline from Part III applies unchanged, with the terms renamed.
| Local inference | Product engineering |
|---|---|
| Memory bandwidth | Latency budget — what the user will wait for |
| Usable RAM | Cost ceiling per interaction |
| Model size class | Which tier of model the unit economics permit |
| Context length setting | How much history the design can afford to resend |
| Acceptance criteria | The behavioural specification, and the regression suite |
Writing this down before choosing a model is the same discipline as writing a mechanical specification before choosing a bearing. The failure mode is identical too: a component chosen from a datasheet, without a specification to test against, that works in the prototype and fails in the field.
The evaluation set is the real deliverable.
In conventional software, tests are infrastructure supporting the product. In AI systems the relationship inverts, for one reason: the component you depend on will be replaced, repeatedly, by something you did not build and cannot inspect.
Models are deprecated. Providers change defaults. A quantization is re-released. A prompt that worked stops working. In every case the question is the same — is the system still doing its job? — and the only thing that answers it is a suite of tasks with written acceptance criteria.
An organisation with a good evaluation set can swap models in an afternoon. One without has to re-form an opinion from scratch every time, and will make that decision on impressions.
I have characterised my hardware and evaluated several models against a task. Here are my findings: [paste envelope + results table]. Draft a one-page component specification for this AI capability as if it were any other engineered component: - Intended use, stated narrowly - Functional requirements with measurable acceptance criteria - Operating envelope (latency, memory, context limits) - Known failure modes and what each looks like in output - Verification method — how a successor model would be qualified - What would trigger re-qualification Flag anything my findings do not actually support. I would rather have gaps marked than filled with plausible text.That final instruction is the one that makes the output usable. A specification with honest gaps is an engineering document; one with plausible filler is a liability.
What to carry forward.
Name the scarce resource first. Every performance claim, yours or anyone's, assumes something about what is scarce. Bandwidth, memory, cost per call, power, or patience. Identify it before you optimise, or you will optimise the wrong thing well.
Derive the envelope before choosing a component. A size class derived from your constraints outlives every model name, and it turns an open-ended search into a filter.
Predict, measure, then explain the difference. The gap between prediction and measurement is where the learning is. A correct prediction confirms your model; a wrong one improves it.
Write acceptance criteria before you build anything. Criteria written afterwards are rationalisations, and effort already spent will bias them.
Test for refusal, not just for capability. Whether a system knows the limits of what it has been given matters more than its ceiling, and it is nearly absent from published benchmarks.
Reach for architecture before scale. Retrieval, decomposition, and context economy solve problems that adding parameters does not.
Treat the evaluation set as the durable asset. Models are replaceable. The definition of what "working" means is the thing you actually own.
Where the numbers came from.
Technical claims in this guide are drawn from the following, checked in August 2026. Figures quoted are approximate and vary by model, harness, and hardware; verify against the primary source before relying on any specific number.
- Prefill and decode bottlenecks. DigitalOcean, The Hidden Bottlenecks in LLM Inference (April 2026), and the TileLens memory-systems paper, arXiv:2607.04031.
- CPU versus accelerator bandwidth. S. Tripathi, llama.cpp: CPU Inference for LLMs on Consumer Hardware (June 2026).
- Memory channels over clock speed. Hardware Corner, Memory Bandwidth and Tokens per Second in Local LLM Inference (October 2025).
- Quantization structure and the quality curve. Quantization in Practice: GPTQ vs AWQ vs GGUF, The AI Engineer (May 2026); community perplexity comparisons across K-quant levels (2026).
- Mixture-of-experts on constrained hardware. Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study, arXiv:2606.21428 (July 2026). The authors bound their findings to one model at one scale on two devices; this guide does the same.
Let's Talk