+ Field notes · systems programming

From power rail
to prompt.

How a computer gets from mains voltage to a running operating system — and what it stores along the way that can stop it.
Starts at: what firmware is. No prior boot knowledge assumed.
Ends at: the same chain on a microcontroller, for RTOS and embedded work.
Built around: a real fault — a desktop that halts at the vendor splash screen after update restarts, and recovers only when mains power is removed.

+ Before we start

There is a gap between pressing the power button and the existence of an operating system. That gap is where this guide lives.

Most explanations of computers begin after the operating system is running. That is a fair simplification for most purposes, but it makes a whole class of problems unintelligible. If a machine freezes before Windows has executed a single instruction, no amount of knowledge about Windows will help.

So we start earlier — at the power supply — and work up through the chips that wake first, follow the boot chain stage by stage, and reach the operating system near the end. Along the way we look hard at one thing that matters enormously and is rarely discussed: the small amount of writable non-volatile memory that firmware keeps for itself, and what happens when it fills.

Every concept here has a direct counterpart on a microcontroller. A PC boot chain and an STM32 boot chain are the same idea at different scales. Part VI makes the mappings explicit, and the green-edged boxes flag the connection as it arises.

Reading the plates Color carries meaning consistently throughout, and it is functional rather than decorative. Amber is always power. Blue is always control and reset signalling. Magenta is always stored, non-volatile data. Green is always executing code. Learn those four and every plate reads the same way.
On the case study The fault running through this guide is real and unresolved. The vendor is not named, and nothing here should be read as a finding about any manufacturer's product. It is used because it teaches the boundary between firmware and operating system better than an invented example would.

+ Part I

The physical layer.

Before any code runs there is a board with chips on it and voltage arriving from a wall socket. This part establishes what is physically present, what powers it, and in what order things come alive.

01 — Chapter

What firmware is, and why it has to exist.

A processor coming out of reset can do exactly one thing: fetch an instruction from a fixed address and execute it. It has no concept of a disk, a file, a partition, or an operating system. Those are abstractions built by software, and at power-on no software has run.

This creates a bootstrapping problem. The operating system lives on an SSD. To read the SSD you need a driver. Drivers are part of the operating system. To load the operating system you must read the SSD. Something has to break the circle.

Firmware is the code that breaks it. It lives in a memory chip the processor can address directly at power-on, needing no driver and no filesystem to reach. Its job is to bring up enough of the machine — memory, buses, storage controllers — that a real operating system can be located and loaded.

On a modern PC this firmware implements a specification called UEFI, which replaced the older BIOS design. People still say "the BIOS," and vendors still label the setup screen that way, so this guide uses both terms where no confusion results.

01 — Where the time goes between button and desktop
t = 0t ≈ 10–25 s POWER rails stabilise FIRMWARE — UEFI initialise CPU, RAM, buses, storage — then find a bootloader OPERATING SYSTEM kernel, drivers, desktop the fault studied in this guide happens here
Power Control Stored data Executing code The detail that matters
Plate 01 — The firmware window is the largest single block of time in a cold boot, and it is invisible to the operating system that follows. A hang inside it looks, to the user, like a frozen logo.
Embedded parallel On an STM32 or similar microcontroller the same problem is solved by a small mask-programmed boot ROM baked into silicon at manufacture. It is functionally the first stage of a UEFI firmware — fixed code at a known address whose only job is to get something larger running. The PC version is bigger and updatable; the principle is identical.

Firmware is not the operating system, and the boundary is sharp

It is worth being precise about the handover, because many troubleshooting mistakes come from blurring it. Firmware and the operating system do not run together. Firmware runs, does its work, loads the OS loader, and then — at a specific call named ExitBootServices — surrenders control and tears down most of its own machinery.

The consequence is practical. If a machine fails before that handover, nothing the operating system can do will change the outcome. Reinstalling Windows, repairing the bootloader, running sfc /scannow — all of these operate on the far side of a boundary the machine never reached.

02 — Chapter

The flash chip: where firmware physically lives.

Firmware is stored on a dedicated chip on the motherboard, typically an SPI NOR flash device of 16 or 32 megabytes — a small eight-pin package, socketed on some enthusiast boards and soldered on consumer ones.

Two properties of NOR flash shape everything that follows.

It is byte-addressable for reads. The processor can fetch instructions directly from it, which is precisely what is needed at reset. This distinguishes it from the NAND flash in an SSD, which reads only in large pages and needs a controller and a driver.

Writes are asymmetric and destructive. A bit can be cleared from 1 to 0 by programming. A bit cannot be set from 0 back to 1 individually — you can only erase an entire block, typically 4 KB, which returns every bit in it to 1. This asymmetry shapes the design of everything stored on the chip, and it is the root cause of the fault studied in Part V.

02 — SPI flash memory map, 32 MB typical
WHOLE CHIP FLASH DESCRIPTOR MANAGEMENT ENGINE separate processor, own firmware GIGABIT ETHERNET BIOS REGION what people mean when they say "the BIOS" expanded at right BIOS REGION — DETAIL NVRAM — UEFI VARIABLE STORE the only part writable at runtime · typically 128–512 KB DXE VOLUME — DRIVERS, PROTOCOLS largest code section PEI VOLUME — EARLY INIT, MEMORY TRAINING CPU MICROCODE UPDATES BOOT BLOCK — HOLDS THE RESET VECTOR
Power / first code Control regions Writable stored data Executable code
Plate 02 — Note the proportions. Almost the entire chip is read-only code that changes only during a firmware update. One small magenta band is writable while the system runs, and that band is where the fault lives.

The one writable region

Everything on that chip other than the magenta band is written once, at manufacture or during a firmware update, and read thereafter. The NVRAM region is different: it is written during normal operation, by the firmware itself and by the operating system, potentially many times per boot.

It is also small — a few hundred kilobytes against a 32 MB chip. Chapters 10 to 13 are devoted to it, because a region that is small, writable, subject to erase-block constraints, and written by two independent parties is exactly the kind of thing that fails in interesting ways.

Hold this thought A finite, small, writable store, shared between firmware and the operating system, on media that cannot rewrite in place. Everything in Part V follows from those four facts.
03 — Chapter

Power rails, and what "off" actually means.

A desktop power supply does not produce one voltage. It produces several, and it does not produce them all at the same times. Knowing which rails are live when is the single piece of knowledge that explains the fault in Part V.

RailVoltagePowersLive when
+12V12 VCPU, GPU, motors, fansSystem running only
+5V, +3.3V5 V, 3.3 VLogic, drives, most chipsSystem running only
+5VSB5 V standbyEmbedded controller, power button logic, wake circuits, some USB portsWhenever mains is connected
VBAT~3 V coin cellReal-time clock, a small amount of CMOS stateAlways, even unplugged

The standby rail is the important one. When you shut a desktop down, it is not off. Mains is still connected, the supply is still producing 5 V standby, and enough circuitry remains energised to watch the power button, listen for wake-on-LAN packets, charge a phone on some USB ports, and hold state.

03 — ACPI system states and live rails
MAIN RAILS+5VSBVBAT S0 — running normal operation S3 — sleep RAM kept refreshed S5 — soft off "shut down" · a power-button hold ends here G3 — mechanical off mains switched off or unplugged · the only state that clears standby ● rail energised ○ rail dead The gap between S5 and G3 is the whole story of the fault in Part V.
Rail live The critical distinction
Plate 03 — Holding the power button takes a machine to S5. Switching off at the wall takes it to G3. State held in standby-powered circuitry survives the first and not the second.
Why this is the crux On the machine that prompted this guide, a power-button hold does not clear the fault but a mains disconnection does. Read against plate 03, that single observation localises the problem to something living on +5VSB — before any other diagnostic step is taken.

The power button is a request, not a switch

On any machine built since the late 1990s the front power button is a momentary contact wired into standby-powered logic. Pressing it connects and disconnects nothing. It asserts a signal the embedded controller reads, and the controller decides what to do.

A short press is a request the operating system can handle gracefully. A four-second hold is an override: the controller drops the main rails regardless of what software wants. Even the override leaves +5VSB untouched, because the controller itself runs on that rail — cutting it would be cutting its own power.

Embedded parallel This is the structure of an STM32 backup domain: VBAT keeps the RTC and backup registers alive across a full VDD loss, and a system reset does not clear them. If you have ever chased a bug where state survived what you thought was a full reset, you have met this problem. Multi-domain power is why "reset" is an ambiguous word in hardware, and why you must always ask which reset.
04 — Chapter

Who is actually on the board.

A modern PC is not one processor. It is several independent processors that boot in sequence, and the main CPU is not the first of them.

ComponentRoleWakes
EC — embedded controllerSmall microcontroller. Watches the power button, sequences the rails, controls fans and thermalsFirst — runs on standby
PCH — chipsetOwns SPI flash access, USB, SATA, PCIe lanes, the real-time clockSecond
ME — Management EngineIndependent processor inside the PCH with its own firmwareBefore the main CPU
CPURuns UEFI, then the operating systemLast

That ordering has a consequence worth sitting with. By the time your CPU executes its first instruction, at least two other processors on the board have already booted and are running their own code. If one of them is stuck, the CPU may never be released from reset — and the screen shows nothing, or a logo, indefinitely.

04 — Board block diagram · who talks to whom
PSU 12V / 5V / 3.3V +5VSB always EC power sequencing runs on standby POWER BUTTON momentary signal PCH / CHIPSET MANAGEMENT ENGINE SPI CONTROLLER SPI FLASH firmware + NVRAM CPU held in reset until everything else is ready DDR5 must be trained NVMe SSD — EFI SYSTEM PARTITION + WINDOWS not reachable until PCIe is up +5VSB release reads firmware
Power Control Stored data Executing code
Plate 04 — The CPU does not fetch firmware from the flash chip directly. It goes through the chipset's SPI controller. That indirection matters: a chipset that has not initialised properly means a CPU that cannot fetch instructions at all.
Embedded parallel The EC does exactly what a supervisor MCU does in a larger embedded design — power sequencing, thermal management, and holding the application processor in reset until rails are in tolerance. If you have built a board with an i.MX or a Zynq alongside a small companion micro, you have built a PC in miniature. The design pressure is identical: something small and always-on must decide when it is safe to let something large and power-hungry start.

+ Part II

The boot chain.

Now we follow the machine forward in time, from the instant mains voltage is applied to the moment the operating system takes over. This is a strict sequence, and each stage depends on the one before it.

05 — Chapter

Power-on sequencing.

Before any instruction executes, a choreography of signals has to complete. Rails must come up in a defined order, settle within tolerance, and be confirmed stable — and only then is the CPU released from reset. Rushing this would mean executing instructions on unstable voltage, which produces failures that look random and are nearly impossible to debug.

SignalMeaning
PS_ON#Asserted by the EC to tell the supply to bring up the main rails. The # means active-low: pulling it to ground turns the supply on.
PWR_OKAsserted by the supply once the main rails are in tolerance. The supply's own statement that it is ready.
RSMRST#Resume reset. Releases the standby-powered portion of the chipset. De-asserts once +5VSB is stable.
PLTRST#Platform reset. The big one. When this de-asserts, the CPU begins fetching instructions.
05 — Power-on sequencing · time runs left to right
AC MAINS +5VSB RSMRST# PS_ON# 12V / 5V / 3.3V PWR_OK PLTRST# plug inbutton rails okreset released CPU fetches first instruction
Power rail Control signal Code begins
Plate 05 — Note how much happens before the orange dot. Everything left of it is hardware sequencing with no software involved. A machine that hangs here shows a black screen, not a logo — which is how you tell this stage apart from the next one.
Diagnostic value Because +5VSB and RSMRST# come up as soon as mains is applied and stay up through a soft-off, any state established in that window persists across a power-button hold. Removing mains forces the whole sequence to restart from the left edge of plate 05.
06 — Chapter

The reset vector, and running without memory.

When PLTRST# de-asserts, the CPU begins executing at a hardwired address called the reset vector. On x86 this is 0xFFFFFFF0 — sixteen bytes below the top of the 32-bit address space. That address is not RAM. The chipset maps it onto the top of the SPI flash chip, so the very first fetch lands in the boot block shown in amber in plate 02.

There is an immediate problem: DRAM does not work yet. Memory controllers need configuration, and DDR5 needs elaborate calibration before a single byte can be reliably stored. So the earliest firmware code runs with no usable RAM — no stack, no heap, no variables in the ordinary sense.

The solution is a trick called Cache-as-RAM. The CPU's own cache is placed in a mode where it behaves as a small block of ordinary memory, disconnected from the DRAM behind it. That gives early firmware a few hundred kilobytes of scratch space — enough to run a stack and get the memory controller working.

06 — Cache-as-RAM · executing before memory exists
CPU L2/L3 CACHE locked as scratch RAM stack lives here ~256 KB usable SPI FLASH instructions fetched from 0xFFFFFFF0 DDR5 DRAM not yet usable needs training first reads code no access yet RESULT firmware can run a normal C program with a stack, long enough to bring up RAM
Stored data Executing code Not yet available
Plate 06 — Cache-as-RAM is a bootstrapping trick with no equivalent in ordinary programming: the CPU repurposes a performance feature as its only working memory for a few hundred milliseconds, then abandons it.

The UEFI phase model

UEFI organises all of this into named phases. These abbreviations appear in any firmware documentation, so they are worth learning.

07 — UEFI phases, and where the splash appears
SEC cache-as-RAM PEI train memory DXE drivers, devices BDS choose boot device OS kernel takes over vendor splash drawn here NVRAM VARIABLE STORE read and written from DXE onward
Control handoff Stored data Executing code What the user sees
Plate 07 — The vendor logo is drawn early in DXE, once a graphics output protocol exists. Everything after that point happens with the logo already on screen, which is why "frozen at the logo" covers a wide range of underlying stages.
07 — Chapter

Memory training, and why it is cached.

DDR5 runs at signalling rates where the electrical delay along a circuit-board trace is a significant fraction of a clock period. Making it work is not a matter of setting a frequency register. The memory controller must empirically determine, for every data lane, the exact timing at which signals should be sampled — sweeping delays, writing test patterns, reading them back, and building a map of what works.

This is memory training, run by the Memory Reference Code during the PEI phase. On a DDR5 system with several ranks it can take from a few seconds to well over half a minute.

Half a minute of black screen on every boot would be unacceptable, so firmware caches the result. The trained timings are written to non-volatile storage, and on later boots the firmware checks whether the configuration still matches — same modules, same slots, same settings — and if so loads the saved values instead of retraining.

08 — Memory training · fast path and slow path
PEI PHASE STARTS no usable DRAM CACHED PARAMS VALID? LOAD SAVED TIMINGS ≈ 1 second yes FULL RETRAIN sweep every lane 5–40 s, screen blank no SAVE TO NVRAM for next boot DRAM USABLE continue to DXE
Slow path Decision Stored data Executing code
Plate 08 — Training results are cached in non-volatile storage, the first of several things that quietly accumulate there. If that cache is invalidated or unreadable, the machine takes the slow path and appears to hang.
A common misdiagnosis A machine that sits on a blank or logo screen for thirty seconds after a firmware change is very often not hung. It is retraining memory because the cached parameters were invalidated. Waiting five minutes before declaring a hang is good practice, and costs nothing.
Embedded parallel Calibrate-and-cache is a pattern you will reuse constantly: ADC offset trims, oscillator tuning values, touchscreen calibration, sensor zero points. The structure is always the same — an expensive characterisation at first run, results stored in non-volatile memory, a validity check on later runs. The failure mode is always the same too: if the validity check is wrong, you either recalibrate needlessly every boot or, worse, trust stale values.
08 — Chapter

DXE and BDS: building a working machine.

With real memory available, firmware can behave like ordinary software. The DXE phase — Driver Execution Environment — loads a large collection of drivers from the flash chip and executes them. These enumerate the PCIe bus, initialise USB controllers, bring up the storage controllers, set up graphics output, and start the TPM.

This is where the vendor logo appears, because it is the first point at which a display exists to draw on.

DXE is also where the largest number of things can go wrong, precisely because it is where the machine touches the outside world. Every USB device is enumerated. Every PCIe device is probed. A device that responds incorrectly, or does not respond at all, can stall the enumeration loop — and since the logo is already on screen, the symptom is a frozen logo.

The BDS phase — Boot Device Selection — then decides what to boot. It reads a list of boot options from the NVRAM variable store, tries them in order, and hands control to the first that works.

09 — From variable to bootloader
NVRAM VARIABLES BootOrder 0003, 0001, 0000 Boot0003 Windows Boot Manager Boot0001 USB device stored on the SPI flash chip points to EFI SYSTEM PARTITION on the SSD · FAT32 \EFI\ Microsoft\Boot\ bootmgfw.efi Boot\ bootx64.efi loads WINDOWS BOOT MANAGER spinning dots start here A hang left of the first arrow means firmware. A hang right of the second means the operating system.
Storage / control Stored data Executing code
Plate 09 — The boot path is defined by variables in NVRAM, not by anything on the disk. This is why a corrupted variable store can make a perfectly healthy Windows installation unbootable, and why "reinstall Windows" is the wrong instinct.
09 — Chapter

The handoff, and warm versus cold.

The bootloader loads the kernel, then calls ExitBootServices. At that moment firmware tears down its own drivers and memory allocations and hands ownership of the hardware to the operating system. A small residue called runtime services survives — and one of the things it provides is the interface through which the operating system reads and writes UEFI variables. That becomes important in chapter 13.

Not all restarts are equal

This distinction is essential for understanding the fault in Part V.

TypeWhat happensDevice state
Warm reset
a normal restart
PLTRST# is asserted and released. Rails stay up throughout.Devices are not power-cycled. Many retain internal state.
Cold boot
shutdown, then start
Main rails drop, then return. Standby stays up.Most devices re-initialise. Standby-powered ones do not.
Mechanical off
mains removed
Every rail except the coin cell collapses.Everything re-initialises from scratch.

A device or controller that mishandles a warm reset — that does not correctly re-initialise when its power was never interrupted — will fail on restarts while working perfectly on cold boots. This is a real and common class of fault, and it produces exactly the selective symptom pattern we are investigating.

Embedded parallel On an MCU you meet the same taxonomy as distinct reset sources, readable from a status register: power-on reset, brown-out reset, external pin reset, watchdog reset, software reset. Good embedded firmware reads that register on startup and behaves differently depending on the answer, because how you got here determines what you can safely assume about peripheral state. PC firmware does the same thing; it just does not tell you about it.

+ Part III

The store that fills up.

This part is the heart of the guide. Everything so far has been context for one small region of flash that firmware and the operating system both write to — and that behaves in ways neither of them entirely controls.

10 — Chapter

What a UEFI variable actually is.

A UEFI variable is a named blob of bytes that survives power loss. It is the firmware's equivalent of a configuration file, except there is no filesystem — just a region of flash and a set of conventions for carving it up.

Each variable has four parts.

PartPurpose
Namespace GUIDA 128-bit identifier that scopes the name. Two vendors can both define a variable called Config without collision, because their GUIDs differ.
NameA UTF-16 string — BootOrder, Boot0003, SecureBoot, dbx.
AttributesA bitfield. The important flag is NON_VOLATILE: set, the variable is written to flash; clear, it lives only in RAM until reset. Others control whether the operating system can access it after ExitBootServices, and whether writes must be cryptographically authenticated.
DataThe payload. Anywhere from four bytes to tens of kilobytes.
10 — Anatomy of a stored variable
ONE VARIABLE RECORD AS STORED IN FLASH STATE NAMESPACE GUID NAME (UTF-16) ATTRIBUTES DATA THE STATE BYTE — WHY DELETION IS NOT DELETION 0x3F ADDED · the live record 0x3D DELETED · still occupying space 0x3E IN_DELETED · superseded by a newer copy WHY IT WORKS THIS WAY Flash bits go 1 → 0 freely, but 0 → 1 only by erasing a whole block. Clearing bits in the state byte marks a record dead without erasing anything.
Payload Stored data The mechanism that matters
Plate 10 — Deleting a UEFI variable does not free its space. It clears bits in a state byte, marking the record dead while leaving every byte of it in place. This one detail drives everything in the next three chapters.

Who writes these

Two independent parties write to the same store, and neither has full visibility of the other.

The firmware writes boot entries, boot order, setup options, memory training results, error records, and Secure Boot key databases.

The operating system writes through the runtime services interface that survived ExitBootServices. On Windows this happens through SetFirmwareEnvironmentVariable. The operating system creates and updates boot entries, stages firmware update requests, and — depending on platform and version — records diagnostic and event data.

The shared-resource problem Neither party knows how much room the other is using, and neither is obliged to clean up after itself promptly. A finite resource with two uncoordinated writers is a classic engineering hazard, and it is the structural reason this region fails more often than its size would suggest.
11 — Chapter

Inside the store: an append-only log.

Because flash cannot rewrite a byte in place, the variable store is not organised as a table that gets edited. It is organised as an append-only log.

Updating a variable does not modify the existing record. It writes a complete new record at the end of the used space, then clears bits in the old record's state byte to mark it superseded. The old record stays exactly where it is, consuming exactly as much room as before.

The consequence is direct and slightly alarming: every write to a UEFI variable consumes free space, including writes that merely change a value that already exists. Update BootOrder a thousand times and you have a thousand records, of which one is live.

11 — Free space over successive writes
FRESH STORE AFTER NORMAL USE AFTER MANY UPDATE CYCLES LIVE FREE SPACE LIVE SUPERSEDED — DEAD BUT OCCUPIED FREE LIVE SUPERSEDED RECORDS ACCUMULATED OVER MONTHS free space now too small for the next write The live data barely grows. What grows is the debris behind it.
Live records Superseded records Exhaustion point
Plate 11 — The amount of genuinely live configuration in a variable store is small and roughly constant. What grows without bound is the trail of superseded records behind it.

Fault-tolerant write

There is a further constraint. Power can be lost at any instant, including halfway through updating a variable. If that left the store in an inconsistent state, the machine might never boot again.

Firmware therefore uses a fault-tolerant write protocol: the new record is staged in a dedicated working block, a marker is set, the record is committed to the main store, and only then is the working block released. If power fails mid-sequence, the next boot inspects the marker and either completes or discards the operation.

This is correct and necessary. It also means every write touches more than one flash block, and consumes more room and more time than the size of the data would suggest.

Embedded parallel This is a journalling write, the same structure as a filesystem journal or an EEPROM emulation library's transfer buffer. If you have implemented power-fail-safe non-volatile storage on an MCU, you have written this protocol — probably as a two-page scheme with a valid marker. The PC version is larger and standardised, not different in kind.
12 — Chapter

Garbage collection, and when it does not run.

Debris cannot accumulate forever, so firmware implements reclaim — a garbage collection pass. It walks the store, copies every live record into a clean area, erases the old blocks, and starts again with the dead records gone.

The question that matters is not how reclaim works. It is when it runs, and reclaim is typically triggered only when free space falls below a threshold and a write is attempted. It is not scheduled, not periodic, and not something the user can invoke.

12 — The reclaim cycle, and the window where it fails
WRITE REQUESTED by firmware or the OS ENOUGH FREE SPACE? APPEND RECORD mark old one superseded yes no RECLAIM copy live records, erase blocks retry reclaim cannot complete THE FAILURE WINDOW live records exceed a clean block · a stale marker blocks the pass · the write neither succeeds nor returns
Reclaim path Decision Stored data Where it hangs
Plate 12 — Reclaim runs lazily, at the moment of greatest pressure, on the boot where least room remains. If it cannot complete, the write that triggered it does not return, and the boot does not continue.
The design smell Garbage collection that runs only under pressure will be exercised least during testing and most in the field, years after release, on machines nobody is watching. This is a general lesson worth carrying into your own designs: the recovery path that runs rarely is the path that is least likely to work.
13 — Chapter

How the store actually fails.

Four distinct failure modes, with different symptoms.

ModeMechanismTypical symptom
ExhaustionSuperseded records fill the region; reclaim cannot free enoughFirmware updates fail; boot hangs at the splash; settings do not persist
FragmentationTotal free space is adequate but no contiguous run is large enough for a big recordLarge writes fail while small ones succeed — appears intermittent
CorruptionPower lost at the wrong instant, or a bad block, leaves an inconsistent headerFirmware refuses to parse; may fall back to defaults or hang
Stale transactionA fault-tolerant write marker left set from an operation that never completedEvery boot retries the same operation and stalls the same way

The last two share a defining property, and it is the one that matters for Part V. They persist across resets, because the flawed state is stored in non-volatile memory. A reset does not clear them; the machine reads the same bad state back on the next boot and fails identically. Only rewriting the store fixes it.

The vendor tooling tell A useful diagnostic signal, generalisable well beyond one manufacturer: when a vendor ships a dedicated NVRAM repair utility inside its firmware update package, that platform has a known history of variable store problems. Nobody writes and maintains such a tool speculatively. If you are diagnosing a machine and find one in the update bundle, you have learned something real about what the vendor expects to go wrong.

+ Part IV

Updating firmware.

Firmware updates are the one routine operation that rewrites the flash chip. Understanding what they touch — and what they quietly reset as a side effect — is necessary to interpret the case in Part V.

14 — Chapter

Capsule updates: the two-phase handshake.

The modern way to update firmware from a running operating system is the UEFI capsule. The operating system cannot write the flash chip directly — that region is locked while the OS runs, deliberately, because unrestricted write access would be a catastrophic security hole.

Instead the update is handed over in two phases across a reboot.

In phase one the operating system writes the new firmware image, wrapped in a signed capsule, to the EFI System Partition, and sets a UEFI variable telling firmware that an update is pending. In phase two, on the next boot, firmware sees the variable, locates the capsule, verifies its signature, and performs the flash before the operating system loads.

13 — Capsule update · two phases across a reboot
REBOOT PHASE 1 — OPERATING SYSTEM RUNNING WRITE SIGNED CAPSULE TO ESP an ordinary file on an ordinary partition SET "UPDATE PENDING" VARIABLE a write into NVRAM — the handoff token REQUEST RESTART PHASE 2 — FIRMWARE, BEFORE THE OS LOADS READ THE PENDING VARIABLE if unreadable, nothing happens VERIFY SIGNATURE, FLASH THE CHIP screen typically blank or on the vendor logo CLEAR THE VARIABLE, CONTINUE BOOT IF THE FINAL STEP NEVER HAPPENS the variable stays set, and every subsequent boot retries the same update
Firmware control Stored data OS execution The failure loop
Plate 13 — The whole mechanism hinges on a single variable in NVRAM. If it can be set but not cleared, the machine attempts the same update on every boot — a self-reinforcing loop that no operating system action can break.
Note the coupling A capsule update is triggered by an operating system update, requires a restart, and depends entirely on a variable stored in NVRAM. Every element of the symptom pattern in Part V is present in this one mechanism.
15 — Chapter

What a flash actually resets.

There is a second update path. Vendor tools can flash the chip directly from Windows using a utility such as AMI's AFUWIN, driven by command-line switches that select which regions to program. A typical vendor invocation looks like this:

AFUWINx64.EXE image.cap /p /b /n /r /sp

The switches matter more than they appear to. /p programs the main firmware body, /b the boot block, and — the significant one — /n programs the NVRAM region.

That last switch means the variable store is not preserved across the update. It is rewritten. Every superseded record, every stale transaction marker, every accumulated fragment is erased and replaced with a clean store.

14 — Which regions each recovery action actually clears
RAM STANDBY STATE NVRAM STORE FIRMWARE CODE Restart (warm) rails never drop cleared kept kept kept Shutdown / button hold reaches S5 cleared kept kept kept Mains removed reaches G3 cleared cleared kept kept Firmware flash with /n rewrites the chip cleared cleared cleared replaced Only the bottom row clears the variable store — and it does so regardless of what the new firmware version changed.
Cleared Retained Clears standby or beyond
Plate 14 — This table is the analytical core of Part V. A firmware update resets the variable store as a side effect, which means a flash can appear to fix a problem it did not address.
The inference trap When a firmware update makes a fault disappear, there are two explanations and they are not equally likely. Either the new version contains a fix, or the act of flashing reset accumulated state. Distinguishing them requires reading the changelog. If the release notes list only security patches, the second explanation is the stronger one — and the fault will return.

+ Part V

Reading a real fault.

Everything so far has been machinery. Here we use it. The case is a consumer desktop tower that has behaved the same way since it was delivered, and the reasoning below is the reasoning any competent diagnosis follows: constrain first, hypothesise second.

16 — Chapter

Symptoms as evidence.

The observed behaviour, stated without interpretation:

ObservationWhat it rules out
Halts at the vendor splash screen; no operating system progress indicatorEverything after ExitBootServices. This is firmware, not the OS.
Occurs only on restarts triggered by an operating system updateGeneric hardware failure. A failing component would not select for one restart type.
Ordinary restarts and cold boots complete normallyPersistent damage to code or storage. The same firmware executes fine most of the time.
A power-button hold does not recover itAnything cleared by dropping the main rails.
Removing mains supply does recover itAnything that survives loss of standby power — narrowing to one region of plate 14.
Present since the machine was deliveredDegradation, wear, and user-induced causes.

Read together, those six lines are more diagnostic than any tool. They place the fault after firmware has started executing and drawn a logo, before the operating system loads, in a domain powered by standby, triggered by a mechanism specific to operating system updates, and present from day one.

Only one thing described anywhere in this guide satisfies all six constraints at once.

17 — Chapter

Why two ways of switching off give different answers.

The recovery asymmetry is the strongest single piece of evidence, and it is worth stating precisely why.

A power-button hold takes the machine to S5. The main rails collapse; +5VSB stays energised because the embedded controller that performed the shutdown runs on it. Switching off at the wall takes the machine to G3, where standby collapses too.

If the fault survived the first and not the second, the flawed state must live somewhere that standby power maintains. Two candidates fit: volatile state inside a standby-powered device — the embedded controller's own RAM, or a USB controller on an always-on port — and non-volatile state in the flash chip, which survives everything but is re-read on each boot and can be corrected by a full power cycle only if the firmware's recovery path is itself reached.

The trigger discriminates between them. A fault that occurs specifically after operating system updates, and not on ordinary restarts, points at the mechanism that operating system updates use and ordinary restarts do not: a variable written into NVRAM to request work at the next boot.

The working hypothesis A pending request stored in the variable store — or a store too full to service the write that would clear it — leaves firmware attempting the same operation on every boot and stalling in the same place. Removing mains does not erase the store, but it does force a full re-initialisation of every controller involved, which can allow the stalled operation to complete on the following boot. The condition then rebuilds, and the fault returns.

That hypothesis is testable, which is what makes it worth stating. It predicts recurrence on a timescale of weeks. A hypothesis that predicted nothing would be a story, not an analysis.

18 — Chapter

Testing it, and the honest limits.

In this case a firmware update was applied, and the next operating system update restart completed normally. It would be easy to record that as a fix. It is not one, for two reasons drawn directly from Part IV.

First, the changelog for that release listed security patches only, with nothing under problem fixes. There was no documented change capable of correcting a boot hang.

Second, the flash was performed with the /n switch, which rewrites the variable store. The bottom row of plate 14 applies: the update cleared exactly the state the hypothesis blames, whether or not the new firmware differed in any relevant way.

One clean boot after an intervention that resets the suspected cause is not evidence that the cause was addressed. It is the outcome both explanations predict. Only recurrence, or its sustained absence over many update cycles, separates them.

If this happensConclude
Fault returns within weeks or a few monthsAccumulation confirmed. The flash reset a counter, not a defect.
Fault absent across many update cycles and a feature updateHypothesis weakened. Consider a genuine firmware fix or an unrelated coincidence.
Fault returns immediatelySomething is regenerating the condition quickly — a repeatedly failing update, not slow accumulation.
What this guide does not claim This is inference from observed behaviour and published vendor documentation. It has not been confirmed by instrumenting the machine, and no claim is made about any manufacturer's product or its quality. The value here is the method — constrain with observations, hypothesise only what the constraints permit, and state what would prove you wrong.

+ Part VI

The same chain, one scale down.

Everything in this guide has a microcontroller equivalent. Not an analogy — the same engineering problem, solved with the same structures, at a size you can hold in your head and inspect with a debugger. This part makes the mapping explicit.

19 — Chapter

Boot ROM, bootloader, application.

An STM32 coming out of reset does exactly what an x86 does: fetch from a fixed address. The address differs, and the mechanism for choosing what sits there differs, but the shape is identical.

On Cortex-M the vector table sits at the base of the boot region. The first word is the initial stack pointer, the second is the reset handler's address. The core loads the stack pointer, jumps to the handler, and execution begins. Where that boot region maps — internal flash, system memory containing the mask-programmed boot ROM, or SRAM — is selected by BOOT pins or option bytes sampled at reset.

15 — The same chain at two scales
PC — UEFI MCU — CORTEX-M EC sequences rails, releases PLTRST# standby-powered supervisor POR / BOR circuit releases reset on-die brown-out detector Reset vector at 0xFFFFFFF0 mapped to SPI flash boot block Vector table at boot region base SP then reset handler SEC / PEI — cache-as-RAM, train DDR no usable RAM at entry Startup code — clocks, .data, .bss SRAM works from reset; no training DXE / BDS — drivers, read BootOrder boot target chosen from NVRAM Bootloader — check image, read config slot chosen from emulated EEPROM ExitBootServices — OS owns hardware runtime services survive Jump to application, start scheduler relocate VTOR first
Power domain Initialisation Handoff First code
Plate 15 — Read across the rows. Every stage on the left has a counterpart on the right. The PC version is larger, standardised, and written by several vendors who must interoperate; that is the whole of the difference.

One asymmetry is worth naming. The MCU has no memory training stage, because on-chip SRAM works from the instant power is valid. That single difference removes the longest and most fragile step in the PC boot chain — and it is why an MCU can be running application code microseconds after reset while a PC takes twenty seconds.

20 — Chapter

EEPROM emulation: the identical bug.

Most modern microcontrollers have no real EEPROM. They have flash, with the same constraint the PC has: bits clear individually, but set only by erasing a whole page. To store configuration that changes at runtime, you emulate EEPROM in flash — and every vendor library that does this converges on the same design as the UEFI variable store, because the constraint forces it.

Two pages. Writes append a record to the active page. Superseded records are marked dead, not removed. When the active page fills, live records are transferred to the spare page, the old page is erased, and the roles swap.

16 — Two-page EEPROM emulation · page transfer
STATE A — NORMAL OPERATION PAGE 0 — ACTIVE FREE PAGE 1 — ERASED, SPARE all bits 1 STATE B — ACTIVE PAGE FULL, TRANSFER RUNS PAGE 0 — FULL SUPERSEDED RECORDS — NO FREE SPACE PAGE 1 — RECEIVING FREE live only If power fails during the transfer, both pages hold partial state — which is why the page header carries a status word. If live records outgrow one page, the transfer can never complete. The device hangs or loses configuration.
Live records Superseded Transfer / failure point
Plate 16 — Compare with plates 11 and 12. This is the same algorithm at a thousandth of the scale, with the same two failure modes: a transfer interrupted by power loss, and live data that has outgrown a single page.
The bug you will actually write The failure that bites hardest in production is not corruption — it is a slow leak. A task writes a counter to emulated EEPROM every minute "because it is only a few bytes." Two years later the page transfer runs during a brown-out and the device will not start. The write frequency was the defect; the storage layer merely reported it.
21 — Chapter

Power domains and reset sources.

Chapter 3 established that a PC has several power domains and that "off" is ambiguous. An MCU has exactly the same structure, exposed more honestly.

PC conceptMCU equivalentShared consequence
+5VSB standby railVBAT backup domainState survives what looks like a full power-off
Embedded controllerPOR/BOR supervisor circuitSomething always-on decides when the main core may run
PLTRST# platform resetSystem reset (NRST)Resets the core; does not reset the backup domain
CMOS clear jumperBackup domain reset (BDRST)The only way to clear persistent state deliberately
Warm versus cold restartRCC_CSR reset flagsHow you arrived determines what you may assume

The right-hand column is the transferable lesson. Reset is not one operation, and "the device restarted" is not a complete description of what happened. An MCU tells you which reset occurred if you read the flag register before clearing it. Firmware that skips this is firmware that will eventually make an assumption that is false exactly once, on one customer's unit, in a way you cannot reproduce.

Practice Read the reset cause register in your startup code, before anything clears it, and record it. On a watchdog reset, treat peripheral state as unknown and re-initialise fully rather than resuming. On a brown-out reset, treat any non-volatile write that was in flight as suspect. This costs a few dozen bytes and saves entire debugging weeks.
22 — Chapter

What to carry into your own designs.

Eight lessons, each earned by something in this guide.

Budget non-volatile writes as a resource. Every write consumes erase-cycle life and free space. Decide at design time how often each item may be written, and enforce it. Unbounded write frequency is the root cause of most field failures in this class.

Make the garbage collector run early, not late. Reclaim triggered only at exhaustion is exercised least in testing and most in the field. Trigger at a comfortable threshold, and provide a way to force it.

Instrument the store. Expose free space, record count, and dead-record ratio through a diagnostic interface. The PC platform's inability to report this is precisely why the fault in Part V is hard to diagnose. Do not reproduce that limitation.

Never let a stored request become unclearable. If a flag says "do this work at next boot," the code that clears it must not depend on the work succeeding, or on there being room to write. Bound the retries and record the failure.

Ask which reset, always. Read the reset cause, branch on it, and never assume peripheral state after a warm start.

Separate power domains deliberately, and document them. Anything on an always-on rail is state that survives what your users will call "switching it off."

Distinguish a fix from a reset. When an intervention that clears accumulated state makes a symptom vanish, you have learned nothing yet. Look for what the change actually contained before concluding.

Make the diagnosis falsifiable. The hypothesis in Part V is worth stating only because it predicts recurrence. State what would prove you wrong, then wait for it. An explanation compatible with every outcome is not an explanation.

The through-line Every problem in this guide comes from the same place: a finite resource, written by more than one party, on media that cannot be rewritten in place, with a cleanup path that runs only under pressure. That description fits a UEFI variable store, an emulated EEPROM, a log-structured filesystem, and a flash translation layer equally well. Learn the shape once and you will recognise it wherever it appears.