Menu Close

How the CPU Talks to Memory: Loads, Stores, Addresses, and the Memory Bus

Posted in Computer Architecture

A CPU can perform arithmetic, compare values, execute branches, and manipulate bits.

How the CPU Really Talks to Memory: Loads, Stores, Cache & Memory Bus Explained

But most useful programs need much more than registers and an ALU.

They need data.

That data may include variables, arrays, program instructions, images, files being processed, network buffers, and many other kinds of information.

Most of this information cannot remain inside CPU registers.

It must be stored somewhere else.

That is where memory comes in.

The CPU constantly communicates with memory to retrieve data and write results back.

At the instruction level, two of the most important operations are:

load

and

store

A load moves data from memory toward the CPU.

A store moves data from the CPU toward memory.

Behind these simple instructions is a carefully organized hardware system involving addresses, caches, buses, interconnects, memory controllers, and DRAM.

Understanding this path is an important step toward understanding how the CPU connects to the rest of the computer.

1. Registers Are Fast but Small

The CPU contains registers that hold values currently being used by instructions.

For example, suppose a program needs to add two numbers.

The CPU may execute something conceptually similar to:

add r1, r2, r3

The operation means:

r1 = r2 + r3

The ALU can perform this operation very quickly because both input values are already inside registers.

But registers are limited.

A CPU may have dozens or hundreds of architecturally visible registers, depending on the instruction set.

A program, however, may work with megabytes or gigabytes of data.

Clearly, all of that information cannot fit inside registers.

Most program data therefore lives in memory.

The CPU must move data between memory and registers whenever necessary.

2. Memory Is Organized by Address

Main memory can be viewed as a large collection of storage locations.

Each location has an address.

A simplified memory might look like this:

Address        Data

0x1000         42
0x1004         17
0x1008         93
0x100C         21

The CPU does not normally ask memory:

“Give me the variable called counter.”

Instead, hardware works with an address.

Conceptually, the CPU asks:

Read data from address 0x1000

or:

Write this value to address 0x1000

Addresses are therefore one of the fundamental mechanisms connecting software to hardware.

3. A Load Reads Data from Memory

A load instruction transfers data from memory into a CPU register.

Consider a simplified instruction:

load r1, [0x1000]

Conceptually, this means:

r1 = memory[0x1000]

If memory location 0x1000 contains the value 42, then after the load:

r1 = 42

The CPU can now use that value in arithmetic or other operations.

In a real instruction set, the address is often calculated from registers rather than written directly inside the instruction.

For example, a MIPS-style load might look like:

lw $t0, 8($t1)

The CPU calculates the memory address:

address = $t1 + 8

Then it loads the word stored at that address into $t0.

So a load usually involves both:

address calculation

and

memory access.

4. A Store Writes Data to Memory

A store instruction performs the opposite operation.

It moves data from a register toward memory.

Conceptually:

store r1, [0x1000]

means:

memory[0x1000] = r1

If r1 contains:

42

the CPU writes that value to the requested memory location.

A MIPS-style example might be:

sw $t0, 8($t1)

Again, the CPU first calculates:

address = $t1 + 8

Then it stores the contents of $t0 at that memory address.

Loads and stores are therefore the basic bridge between:

CPU registers
        ↕
memory

5. The ALU Often Calculates Memory Addresses

The ALU is not used only for arithmetic such as addition and subtraction.

It also plays an important role in memory access.

Consider:

lw $t0, 8($t1)

Suppose:

$t1 = 0x1000

The CPU must calculate:

0x1000 + 8

The result is:

0x1008

That becomes the address used for the memory access.

The datapath therefore looks roughly like:

Base Register
      +
    Offset
      ↓
     ALU
      ↓
Memory Address
      ↓
    Memory

This is why load and store instructions use the ALU even though they may not appear to be ordinary arithmetic instructions.

The ALU is calculating the effective address.

6. Addresses and Data Are Different Things

When the CPU communicates with memory, two kinds of information must be transferred.

The first is the address.

This tells the memory system where the CPU wants to access.

The second is the data.

This is the actual information being read or written.

For a load:

CPU → Address → Memory

CPU ← Data ← Memory

For a store:

CPU → Address → Memory

CPU → Data → Memory

The CPU must also indicate whether the operation is a read or a write.

So memory communication involves three basic concepts:

Address
Data
Control

These ideas historically correspond to what are often called:

Address Bus
Data Bus
Control Bus

7. What Is the Memory Bus?

Traditionally, a bus is a shared group of electrical connections used to transfer signals between hardware components.

A simplified CPU-memory system can be imagined like this:

CPU
 │
 ├── Address
 ├── Data
 └── Control
 │
Memory

The address signals identify a location.

The data signals carry the actual value.

The control signals indicate what kind of operation should happen.

For example:

READ
WRITE

Older computer architectures often exposed these buses more directly.

Modern processors are considerably more complex.

The CPU may contain integrated memory controllers, multiple cache levels, high-speed internal interconnects, and specialized memory channels.

So the term memory bus is useful for understanding the basic concept, but modern systems do not necessarily contain one simple shared bus directly connecting the CPU core to DRAM.

8. The CPU Usually Does Not Talk Directly to DRAM

A simple diagram might show:

CPU ↔ RAM

But real computers normally contain several layers between a CPU core and main memory.

A more realistic path is:

CPU Core
   ↓
L1 Cache
   ↓
L2 Cache
   ↓
L3 Cache
   ↓
Memory Controller
   ↓
DRAM

When the CPU performs a load, it does not immediately go all the way to DRAM.

First, the processor checks whether the requested data is already available in a cache.

If it is, the CPU can receive the data much more quickly.

Only when the required cache line is not available does the request need to travel farther through the memory hierarchy.

9. Why Caches Matter

Modern CPUs can execute instructions extremely quickly.

Main memory is much slower.

If every load instruction had to wait for DRAM, the CPU would spend enormous amounts of time waiting.

Caches help reduce this problem.

A cache keeps copies of recently used or nearby memory data close to the CPU.

The typical hierarchy is:

Registers
↓
L1 Cache
↓
L2 Cache
↓
L3 Cache
↓
Main Memory

The closer a storage level is to the CPU, the faster it generally is.

But it is also smaller.

A successful cache lookup is called a:

cache hit

If the requested data is not found, the processor experiences a:

cache miss

The request must then continue to another cache level or eventually to main memory.

10. Memory Is Usually Transferred in Blocks

A CPU may request only a few bytes.

But modern memory systems usually do not move one isolated byte from DRAM every time.

Caches organize memory into blocks called cache lines.

A common cache-line size in modern systems is 64 bytes, although this is architecture-dependent.

Suppose the CPU wants data from:

0x1008

The memory hierarchy may fetch an entire block containing nearby addresses.

This is useful because programs often access nearby memory locations.

For example, when processing an array:

for (int i = 0; i < 1000; i++)
    sum += array[i];

after reading one array element, the CPU is likely to need the next one soon.

Bringing nearby data into cache therefore improves performance.

This behavior takes advantage of spatial locality.

11. The Memory Controller Talks to DRAM

When the requested data cannot be satisfied by the caches, the request eventually reaches the memory controller.

Modern CPUs often integrate the memory controller directly into the processor package or chip.

The memory controller translates processor memory requests into commands that DRAM can understand.

DRAM is organized internally into structures such as:

channels
ranks
banks
rows
columns

The memory controller decides how to access these structures and manages timing requirements.

The CPU core therefore does not directly manage individual DRAM rows and columns.

Instead, it issues memory requests through the processor’s memory subsystem.

The memory controller handles the lower-level communication with physical memory.

12. What Happens During a Load?

Now we can follow a load instruction through the system.

Suppose the CPU executes:

load r1, [address]

A simplified sequence is:

1. Decode the load instruction

2. Read the base register if needed

3. Calculate the effective address

4. Send the memory request into the cache hierarchy

5. Check L1 cache

6. If necessary, check lower cache levels

7. If the data is not cached, request it from main memory

8. The memory controller accesses DRAM

9. The requested cache line returns toward the CPU

10. The required value is extracted

11. The value is written into the destination register

Conceptually:

Instruction
    ↓
Address Calculation
    ↓
Cache Lookup
    ↓
Memory if Necessary
    ↓
Data Returns
    ↓
Register

A single load instruction may therefore trigger a surprisingly large amount of hardware activity.

13. What Happens During a Store?

A store follows a related but slightly different path.

Suppose:

store r1, [address]

The CPU must determine:

where the data goes

and:

what data should be written

A simplified sequence is:

1. Decode the store instruction

2. Read the source register

3. Calculate the effective address

4. Update the appropriate cache entry

5. Eventually propagate the modified data through the memory hierarchy

Modern CPUs often use a technique called write-back caching.

With write-back caching, a store does not necessarily update DRAM immediately.

The CPU can modify the cached copy first.

The changed cache line is marked as dirty.

It may be written back to lower memory levels later.

This avoids forcing the CPU to wait for DRAM after every store.

14. Loads and Stores Connect Programs to Memory

Consider a simple C statement:

c = a + b;

At the source-code level, this looks like one operation.

But suppose a, b, and c are stored in memory.

The processor may need operations conceptually similar to:

load r1, [a]

load r2, [b]

add r3, r1, r2

store r3, 

The actual sequence depends on the compiler, architecture, optimization level, register allocation, and whether values are already cached or held in registers.

But the fundamental pattern is important:

Memory
  ↓
Load
  ↓
Registers
  ↓
ALU
  ↓
Register
  ↓
Store
  ↓
Memory

This is one of the most important data flows inside a computer.

15. Memory Access Is Much More Expensive Than Register Access

From the programmer’s point of view, accessing a variable may look simple.

From the processor’s point of view, the cost can vary enormously.

A value already in a register is immediately available to the CPU execution units.

A value in L1 cache takes longer.

L2 and L3 take longer still.

Going to DRAM is much more expensive.

So two load instructions that look almost identical in assembly language may have very different execution times depending on where the requested data is found.

This is one reason memory behavior has such a large effect on software performance.

Modern processors therefore invest enormous amounts of chip area and engineering effort in:

caches
prefetchers
memory controllers
load/store units
store buffers
memory scheduling
high-speed interconnects

The challenge is not simply performing arithmetic quickly.

It is keeping the CPU supplied with data.

16. The Simple Bus Model and the Modern CPU

When learning computer architecture, it is useful to begin with this model:

CPU
 ↕
Address Bus
Data Bus
Control Bus
 ↕
Memory

This explains the essential concepts clearly.

The CPU chooses an address.

Control signals indicate a read or write.

Data moves between the processor and memory.

But modern processors add many layers:

CPU Core
   ↓
Load/Store Unit
   ↓
L1 Cache
   ↓
L2 Cache
   ↓
Shared Cache
   ↓
Internal Interconnect
   ↓
Memory Controller
   ↓
Memory Channel
   ↓
DRAM

The basic principle remains the same.

The CPU must identify where the data is located, request the appropriate operation, and transfer the data through the memory system.

The hardware implementing that process has simply become much more sophisticated.

17. Conclusion

The CPU cannot do useful work with registers alone.

Programs contain far more information than registers can hold.

That is why processors constantly communicate with memory.

A load moves data from memory toward a CPU register.

A store moves data from a CPU register toward memory.

The CPU calculates an address, the memory system identifies the requested location, and data travels through the memory hierarchy.

At the simplest level, this communication can be understood through:

Address
Data
Control

or the traditional idea of:

Address Bus
Data Bus
Control Bus

Modern computers add caches, load/store units, interconnects, memory controllers, and DRAM channels between the CPU core and physical memory.

But the fundamental idea remains simple:

The CPU performs computation, while memory holds the larger working state of the program.

Loads and stores connect the two.

Once we understand how the CPU communicates with memory, we are ready to move one level higher.

The next question is no longer just how hardware accesses data.

It is:

Who manages the CPU, memory, devices, files, and running programs?

That is where the operating system begins.

Leave a Reply

Your email address will not be published. Required fields are marked *