More cores, more nearby copies
Suppose you assign a data-processing job to more cores. Each can retain frequently used data nearby, yet the job fails to speed up proportionally. Sometimes a core waits for main memory; sometimes it waits for another core to give up write permission. Should you enlarge the cache, rearrange data, or reduce shared writes? This is a teaching scenario, not a product benchmark.
The twelve existing plates begin with data distance and then follow the handoff of cached copies. The prose connects nearby shops, status labels, and directories to real mechanisms. Replacement policies, atomic operations, and protocol details appear in folded notes. On a first pass, follow who waits and what they are waiting for.
The Memory Wall: a fast CPU can still wait
Follow a data-processing job through one core. It loads a value, computes, then loads another value. Nearby data lets dependent work continue. A request that must reach main memory can leave that work waiting while the clock keeps ticking. Raising the clock rate does not necessarily deliver the missing data sooner.
The Memory Wall names the gap between computation and data delivery. Separate latency from bandwidth: latency is the time before a requested value arrives; bandwidth is the sustained amount transferred per second. A wider road carries more vehicles, but each vehicle still needs time to complete its trip. The analogy leaves out controller queues and the memory operations themselves. [1]
The L1D, L2, LLC, and DRAM ladder in plate 1 illustrates different orders of access cost. Its cycle ranges depend on the processor, clock, and load. An LLC miss cannot be converted into a fixed number of wasted instructions. An out-of-order core may execute independent work while waiting; dependent work or exhausted resources can make the delay visible in execution time. [4]
That distinction matters when diagnosing our job. Faster arithmetic may simply consume the available data sooner if each iteration waits for another load. Cache takes a different approach: retain likely-to-be-reused data near the core so the job makes fewer expensive trips.
Cache keeps frequently used data nearby
The convenience-store analogy explains why a small cache helps. A nearby shop cannot stock everything in a large warehouse, but it can keep frequently requested items. The core checks a nearby copy first. Finding it is a hit; a miss sends the request farther down the hierarchy. Cache occupies storage on the chip without increasing the program's addressable main-memory capacity.
Average memory access time, or AMAT, can be modeled as AMAT = Hit Time + Miss Rate × Miss Penalty. Every lookup pays the hit time. The miss rate is the fraction of lookups that miss, and the miss penalty is their additional cost. Use compatible units for all terms. This simple single-level model compares average access cost; it does not directly predict an entire program's runtime. [1]
Plate 2 assumes a one-cycle hit and an additional 300-cycle miss penalty. With 99% hits, 1 + 0.01 × 300 = 4 cycles. With 95% hits, 1 + 0.05 × 300 = 16 cycles. That four-percentage-point gap gives four times the average cost in this model. Changing the penalty to 100 produces 2 and 6 instead. Both are teaching examples, not measurements of a particular CPU. [1]
The remedy depends on why requests miss. More capacity may prevent some misses while making lookup slower. Prefetching moves data ahead of demand and can turn a later request into a hit. Memory-level parallelism overlaps independent requests to hide some waiting. Overlap does not make an individual DRAM request faster, so it cannot simply be substituted as a smaller penalty for every access in AMAT.
Locality makes cache lines worth fetching
Consider how the same job reads its data. Summing an array usually touches one element and then its neighbor. A loop may also reuse a value it just loaded. Nearby access is spatial locality; reuse over time is temporal locality. These patterns let a small cache cover data the job uses frequently. [1]
Hardware normally transfers a cache line at a time. Plate 3 uses a common 64-byte line, but the actual size depends on the processor. Loading one value brings neighboring bytes along. The transfer pays off if those bytes are used soon. A program that touches only a small part before jumping far away may waste much of that traffic.
A cache must also locate its copies. The index chooses a set, the tag identifies which memory block occupies an entry, and the offset selects a position within the line. Multiple ways give a set more places to hold competing blocks. Extra ways can reduce conflicts while adding comparison and selection hardware.
Different misses need different remedies. First access produces a compulsory miss. A repeatedly used set of data that cannot fit produces capacity misses. Blocks competing for one set produce conflict misses. Other cores can invalidate a useful copy and cause coherence misses. A larger cache alone does not eliminate ownership exchanges caused by shared writes. [1][3]
L1, L2, and L3 balance lookup time and capacity
A larger nearby store is harder to search with extremely low latency. The core also wants more data to stay on the chip, so processors commonly use several cache levels. A small L1 serves urgent accesses. Larger levels catch data that L1 cannot retain. Most requests can stay close while a larger working set still has somewhere to go.
Plate 4 shows one common organization. L1 may have separate instruction and data caches, giving instruction fetch and data access separate paths. L2 is often private, while the last-level cache, or LLC, may be shared across cores and divided into slices. This is an example architecture, not a promise that every CPU has the same levels or sharing arrangement. [1]
Writes introduce further choices. Write-through sends updates to the next level. Write-back can modify a cached copy first and record the change with a dirty bit. The next level may therefore lack the newest data. Write-allocate obtains the line on a write miss. Repeated updates can stay nearby rather than traveling down the hierarchy each time. [1][5]
Delayed write-back creates a responsibility. The newest data may exist only beside one core. When that copy is replaced or another participant requests it, the controller must supply it according to the protocol rather than discard it. The hierarchy that reduced waiting now raises another question: how do different cores manage their copies?
Engineer extension: inclusion and replacement policies
Inclusive, exclusive, and non-inclusive/non-exclusive policies define different inclusion relationships between levels. An inclusive LLC can help track upper-level copies, but evicting its line may require invalidating those copies. Non-inclusive designs lack the same inclusion guarantee and commonly need separate metadata for tracking. Product-specific rules matter.
A set-associative cache must choose which way to evict. Approximate LRU tries to retain recently used data; RRIP estimates how long it may be before data is reused. Better hit rates still require state, update logic, and verification. Capacity alone cannot describe the design.
Private copies need a coherence protocol
Let two cores work on shared data. Both first read X = 10 and keep it in private caches. One core then wants to write X = 20. Without rules for managing those copies, the other could keep using its old value. This teaching example shows why hardware needs cache coherence. [2]
Coherence lets participants obtain writes and agree on their order for one address. It does not require every cache to contain an identical value at every instant. Some copies can become invalid and fetch data again on their next use. What matters is that an invalid old copy cannot continue serving as valid data.
Start with an access-permission rule. During a permitted access phase, a line may have one holder with write permission or multiple holders with read-only permission. This is SWMR. The writer can also read its own data. Another core must obtain permission before writing, and the handoff must preserve the preceding write phase's data. [2]
Hardware manages the line's permissions and data transfer. A program that lets two threads update a counter still needs valid synchronization or atomic operations. Coherence does not automatically combine read, add one, and write back into one indivisible update. We first need to see how the hardware records each copy's permissions.
MESI records validity, sharing, and modification
A controller must know whether a copy can be read or written and whether the next level still has the newest value. MESI uses four states. Modified, or M, is unique and dirty. Exclusive, or E, is unique and clean. Shared, or S, is clean and may be shared. Invalid, or I, means the copy cannot be used. [2]
Return to X. A core can hold a clean, unique copy in E and move to M when it modifies the data. Another core's read then requires the protocol to rearrange data and permissions. The M holder may need to supply the newest value because main memory can still contain an older one. Cache-to-cache transfer describes a path that delivers data directly between caches. [2][6]
S permits sharing; it does not prove that a second copy currently exists. I is not a dirty-data state. Physical storage may retain old bits, but those bits are unusable as a valid copy. Plate 6's matrix helps distinguish exclusive and shared access, but I must not be interpreted as shared-and-dirty. Validity determines whether the copy can serve a request.
The four letters describe stable outcomes. Real transactions also include time spent waiting for data or invalidation acknowledgments. Beginners can follow who may read, who may write, and who supplies the newest value. A complete controller must additionally handle requests that interleave during those handoffs.
Engineer extension: MOESI, MESIF, and transient states
MOESI adds Owned so a designated owner can supply shared data that remains dirty relative to memory. MESIF uses Forward to select a responder among shared copies. These are different protocols; basic MESI's clean Shared state cannot simply be assumed for every extension.
Transient states represent unfinished transactions. Verification needs simultaneous reads and writes, different orders of data and acknowledgments, and resource-exhaustion behavior. Four stable states do not describe a complete RTL controller.
Reads may share; writes must acquire ownership
Assume neither core has X. Core 0 misses and issues the read request called BusRd in plate 7. If no other core holds the line, it may obtain E. When Core 1 reads later, the protocol can leave both copies in S. The readers can use clean nearby data without taking turns acquiring write permission. [2]
Core 0 now wants to write. It already has S data but lacks exclusive write permission. Its controller asks other sharers to invalidate their copies and waits for the acknowledgments required by the protocol before completing the handoff. BusUpgr illustrates this upgrade: the requester already has the data and needs a permission change. It applies when the requester itself holds a valid S copy. [2]
A requester lacking the data needs both data and write permission. That request is commonly described as RFO, Read For Ownership, or BusRdX in the teaching protocol. RFO and BusUpgr are therefore not interchangeable. Modern interconnects may use different packet names, but still need to transfer data and manage permissions.
An invalidation protocol removes old copies, letting future readers fetch the updated line when needed. An update-based protocol instead sends new values to other copies. These approaches produce different traffic and waiting costs. A job that repeatedly alternates writes between cores can pay for a handoff even when each write changes only a few bytes. [2]
Engineer extension: atomic operations, LOCK, and LL/SC
An atomic read-modify-write prevents another conflicting operation from intervening within the update. Coherence ownership provides a foundation, while the core, cache controller, and architectural rules establish atomicity together. Modified state alone is insufficient.
x86 locked operations on cacheable write-back memory contained within one line can normally use cache locking. Split or uncacheable locked accesses may need bus locking. LL/SC uses a reservation; store-conditional can fail because of interference or other architecturally permitted causes. Software must handle retries according to the architecture.
Snooping broadcasts a request to possible holders
Before granting write permission, a controller must find other copies. Snooping resembles a classroom announcement. Core 0 issues a transaction, other cache controllers observe it and check their tags, and matching holders respond, downgrade, or invalidate according to their state. Software does not call each other core to clear its copy. [2]
Broadcast reaches possible holders conveniently, but many recipients have no relevant data and still check. Tag lookup also consumes resources. A snoop and the core's own lookup may compete for access. Extra ports or duplicate tag information can reduce that competition at the cost of more circuitry.
A traditional shared bus also provides a common ordering point for transactions. Small systems can use that order to coordinate permissions. More participants can turn the shared channel and broadcast checks into bottlenecks. The plate's traffic-growth sketch illustrates scaling pressure; its actual rate depends on request generation, message replication, and filtering.
Software can amplify this cost through false sharing. Two threads may each write a different counter, yet both counters occupy one 64-byte line. Write permission then moves back and forth. The shared object is the hardware's management unit, even though the program uses separate variables. Reading common immutable data does not create that same write-ownership contention. [3]
Engineer extension: diagnosing false sharing
Use a profiler or appropriate hardware events to confirm line bouncing, then inspect field addresses and the actual line size. Separate hot writable fields, add padding, or accumulate per thread before combining results. The plate's alignas(64) needs an object-size and layout check: an aligned start does not separate all fields inside an object.
Padding increases footprint and can cause capacity misses. If threads really modify one shared value, that is true sharing; rearranging fields alone cannot fix it. Compare equal workloads before and after a change to establish whether the cost fell.
A directory tracks the relevant holders
If only two cores hold X, why ask the whole system every time? A directory records the nodes that may hold copies. Each line's address maps to a home node. A requester contacts the home, which can notify relevant sharers or locate the owner of the newest data. Uninvolved nodes avoid many unnecessary probes.
The record takes space. A full-bit vector reserves a bit for every tracked node. Limited pointers record only a small number of nodes and need another strategy when that space runs out. A coarse vector represents a group with one bit, saving storage while potentially probing members that have no copy. Each choice trades tracking precision against traffic. [2]
Plate 9 assumes 64 tracked nodes and a 64-byte line. Its full sharer vector needs 64 bits, or 8 bytes. Relative to 64 data bytes, that is 12.5%. The formula is P ÷ (8 × B), with P measured in bits and B in bytes. It counts only the sharer vector, excluding tags, state, and ECC. It also does not specify how many lines the entire directory tracks. [2]
Reduced broadcast can require a longer route. If another owner has the data, a request may travel requester → home → owner → requester. Routing, congestion, and replies determine the total cost. A directory is not guaranteed to be faster on every access; its scaling benefit is greater control over which participants receive a request.
Filters reduce probes; synchronization orders communication
The two methods can operate within one system. A small shared bus may snoop directly. A larger interconnect can track copies in a directory and probe selected nodes. Distributed home agents and snoop filters divide tracking from message delivery to reduce unnecessary broadcast.
A filter must not omit a node that may hold a valid copy. An extra probe usually costs traffic, but a missed holder can break correctness. Some designs back-invalidate private copies when tracking metadata runs out or is evicted. Data can therefore lose a cache hit because the tracking record ran out of space, even if the data itself could still fit. [2]
Now consider a program that writes a result and then announces completion through a flag. A consumer hopes that seeing the flag means it can read the result. These are two addresses. Coherence governs each address's copies; the memory-consistency model and program synchronization govern visibility across addresses. MESI alone does not establish a correct handoff. [10]
Use synchronization supplied by the language and platform, such as a lock or a correctly paired atomic release/acquire operation. It must constrain permitted compiler and processor behavior. The plate's loads and stores illustrate ordering, not usable C/C++ code with unsynchronized ordinary shared variables.
Engineer extension: coherence and memory consistency
x86 TSO, Arm's memory model, and RISC-V RVWMO have specific rules. Weaker ordering does not make different architectures identical. Store buffers, write combining, and out-of-order execution affect observable order; software must select atomic operations or fences according to the language and ISA.
In a valid C/C++ publication example, a producer writes the data before releasing an atomic flag. A consumer must acquire the corresponding published value to establish the intended synchronization. Acquiring an arbitrary value, or inserting an unrelated barrier, does not establish that handoff.
The interconnect must deliver the coherence messages
The directory tells us whom to contact, but messages still need to travel. A shared bus offers one channel. A ring connects nodes around a loop, while a mesh provides multiple routes. A network-on-chip, or NoC, carries these messages inside the chip. Topology changes path lengths and shared bandwidth as well as connectivity.
Return to Core 0's write to X. The directory selects holders involved in the handoff; the NoC carries requests, data, and acknowledgments. Sending a packet does not prove that all other copies have become invalid. Packets may wait or interleave with other transactions. The protocol decides when the permission transfer is complete. [7][9]
Distance affects performance too. Accesses to different LLC slices on one chip may have different costs, known as NUCA. Main-memory access across sockets commonly introduces NUMA differences. Placing a thread near its frequently used data may help more than adding another core, but the improvement needs measurement.
Multicore work therefore depends on both cache permissions and network resources. Hop count, congestion, and whether replies can make progress affect the original job. A fast cache beside each core does not remove all waiting between cores.
Engineer extension: progress and deadlock
Packets can hold buffers while waiting for credits. A cycle of resource dependencies can create deadlock. Virtual channels, separate request/response networks, and rules that avoid cyclic dependencies are common design tools. Adding another channel alone is not a deadlock-freedom proof.
Retries, transient states, and reordered replies need verification alongside the coherence state machine. Arm CHI specifies scalable coherent-interconnect transactions; gem5 Ruby provides models for studying controllers and networks. A few passing simulation examples do not replace full hardware verification.
CXL extends memory roles beyond one chip
The problem continues when participants sit outside the chip. CPUs and accelerators may need to access shared data, and a system may want additional memory. CXL provides different protocol roles for these accesses. It does not imply that every device has the same cache or that all devices share identical memory semantics.
CXL.io handles discovery, configuration, and general I/O. CXL.cache lets a device cache host memory. CXL.mem lets the host access device-attached memory. In the plate's introductory classification, Type 1 uses io+cache, Type 2 uses io+cache+mem, and Type 3 uses io+mem. These roles clarify who accesses whose memory; verify the specification and product for actual capabilities. [8]
An accelerator using the same data needs to consider how its copies participate in coherence. A memory expander first raises capacity and access-path questions. More addressable memory does not mean faster access. The existing plate gives 150–300 ns without test conditions; this is neither a CXL specification guarantee nor an established universal measurement range. Experimental results depend on the device and measurement method; also check the host and topology. [11]
CXL 3.x fabric and back-invalidation capabilities also depend on version and implementation. Putting them on one plate organizes mechanisms to investigate; it does not mean every Type 3 device implements them all. Beyond the chip, designers still track copies, access permissions, and data paths across greater distances and more participants. [8]
Engineer extension: bias has a version and device boundary
The plate's Host Bias and Device Bias sketch simplifies a particular CXL Type 2 protocol context involving access to device-attached memory. It is not a universal data-owner switch for every CXL device, or a rule for every CXL.cache access. Check the specific version's bias, coherence, and back-invalidation rules when studying newer systems.
Return to the job: change data layout or hardware?
- Repeated trips for the same data suggest examining reuse and cache hits. Alternating writes to one line suggest examining ownership handoffs. A fixed input and timing boundary are needed to test either hypothesis.
- Line bouncing between independent writable variables can be false sharing. Changing layout may help; genuinely shared updates still need a correct synchronization design.
- If every core only reads unchanged data, shared copies need no repeated write-permission handoff. Capacity misses, bandwidth contention, and remote accesses can still limit performance.
This lesson does not derive a complete memory model, enumerate every transient state, or deliver controller RTL. The curriculum's Memory ordering and synchronization and NUMA and interconnects lessons are still planned. For a published application, continue with the AI inference architecture classroom to examine how data delivery affects accelerators.
References
- Hennessy and Patterson, Computer Architecture: A Quantitative Approach: Memory hierarchy, locality, AMAT, and coherence fundamentals.
- Nagarajan, Sorin, Hill, and Wood, A Primer on Memory Consistency and Cache Coherence, Second Edition (2020): Coherence, consistency, SWMR, snoops, and directories.
- Ulrich Drepper, What Every Programmer Should Know About Memory: Cache-line behavior and false-sharing consequences.
- Agner Fog, The microarchitecture of Intel, AMD, and VIA CPUs: Measured cache and microarchitecture behavior across processor generations.
- Intel 64 and IA-32 Architectures Software Developer's Manuals: Locked operations, cache locking, memory types, and architectural behavior.
- AMD64 Architecture Programmer's Manual: AMD64 memory-system and MOESI-related architectural reference.
- Arm AMBA CHI Architecture Specification: Scalable coherent interconnect and transaction concepts.
- Compute Express Link 3.1 Specification: CXL.io, CXL.cache, CXL.mem, device types, bias, and fabric features.
- gem5 Ruby Cache Coherence Models: Executable coherence-state-machine and NoC simulation context.
- Linux Kernel Memory Barriers: Software-visible memory ordering, barriers, and multicore synchronization.
- Yan Sun et al., Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices, MICRO 2023: Device and measurement-method dependencies; not a universal latency guarantee.
Learning guide
CPU and Memory Systems
Prerequisites
- Basic CPU, load/store, and memory hierarchy concepts
What I learned
- Explain how the Memory Wall leads to cache hierarchies
- Use the SWMR invariant to explain MESI
- Compare snooping, directories, and hybrid filters
- Distinguish coherence from memory consistency
- Recognize false sharing and CXL device roles