Processes, Address Spaces, and Memory
The same program can run twice at once, with different variable values in each run. Understanding why means separating execution state, the addresses a program sees, and the memory the machine actually uses. It also explains why restarting a notebook kernel loses variables, while setting a container memory limit addresses a different problem.
Programs, processes, threads, and the operating system
OSTEP’s Processes chapter defines a process as a running program: instructions together with current execution state and open resources. Concurrency: An Introduction distinguishes a process from its threads.
Separate thread stacks track separate function calls. They still belong to one address space; they are not memory isolation barriers between threads. Concurrent updates to shared data need coordination, or execution order can change the result.
The operating system schedules execution, manages memory, and provides services such as files. Its kernel runs with privileges; ordinary applications request kernel services through system calls. OSTEP’s Introduction to Operating Systems explains the controlled transition between user mode and kernel mode. A “kernel” in Jupyter Architecture is instead a user process executing Python or R, distinct from the OS kernel.
From creation to exit
For Unix interfaces, the Process API chapter separates several actions:
fork()creates a child process. Parent and child continue from the call’s return point, with separate process identities and private memory state.exec()replaces the current process’s program image with another program, rebuilding code, data, heap, and stack. Success never returns to the old program;exec()itself creates no new process.- A parent can use
wait()orwaitpid()to wait for a child to finish. A blocking wait suspends the calling thread; other threads in the process can still run. If the child has already exited and its status remains uncollected, the call can return immediately. Signals and errors can also end a wait, so returning does not always mean the child has exited. - An exited child stops executing. Unix normally retains a small exit record that the parent collects with
wait()orwaitpid(). An exited process whose record has not been collected is a zombie. Linux’swait()manual also explains that explicitly ignoringSIGCHLDor settingSA_NOCLDWAITprevents exited children from becoming zombies.
The Processes chapter uses three states to explain execution: ready means able to run but waiting for CPU time; running means executing; blocked means waiting for an event such as I/O. The scheduler can take a running task off the CPU and run another ready task. Source-code order alone cannot tell you which process will run first.
Address spaces, stacks, and heaps
In Address Spaces, an address space is a process’s view of memory addresses and their mappings. Pointers generally hold virtual addresses. The same numeric address in two processes can refer to different physical memory. A private address space provides access isolation; it does not give every process a whole exclusive block of RAM.
An address space contains code, static data, stacks, and a heap, and can include mapped files and shared regions. A stack holds call frames, including return locations and local state; a heap holds dynamically allocated data. Memory API distinguishes two levels of management: the allocator manages blocks within a process, while the OS manages its memory mappings. malloc() is a library call that may reuse existing space rather than request memory from the kernel every time. Freeing an object’s memory need not immediately return its pages to the OS; for example, glibc releases free space at the top of the heap subject to the threshold conditions in its mallopt() manual.
Two processes can deliberately map shared memory. Linux’s mmap() manual, for example, specifies that updates to a MAP_SHARED mapping are visible to other processes mapping the same region. Their address spaces remain separate, with some mappings referring to common storage. Shared writes still need coordination.
Virtual memory, paging, and page faults
Virtual memory separates address spaces from physical memory. Address Translation explains the division of work: the OS sets up mappings and access permissions, and the hardware memory management unit (MMU) translates addresses during access.
Paging divides virtual address spaces into fixed-size pages and physical memory into equally sized page frames. A page table records the physical frame for a virtual page and its access conditions. A contiguous range of virtual addresses can therefore occupy noncontiguous physical frames.
Work out an address translation
Assume a teaching model with 4096-byte pages, or 4 KiB. This is a chosen model parameter, not the local machine’s page size. Divide virtual address 12345 by the page size: the quotient is virtual page number 3 and the remainder is offset 57. If process A maps page 3 to frame 7, the physical address is 7 * 4096 + 57 = 28729. If process B maps it to frame 19, the same virtual address translates to 77881.
The following also calculates a 10000-byte buffer starting at a page boundary and exclusively occupying the pages it needs. It requires 3 pages, giving a capacity of 12288 bytes and 2288 unused bytes in the last page. Run it in Python 3:
page_size = 4096
va = 12345
vpn, offset = divmod(va, page_size)
print(f"VPN={vpn}, offset={offset}")
for process, frame in [("A", 7), ("B", 19)]:
print(f"{process}: physical address={frame * page_size + offset}")
size = 10000
pages = (size + page_size - 1) // page_size
print(f"buffer: pages={pages}, capacity={pages * page_size}, slack={pages * page_size - size}")
print(f"two resident pages: {2 * page_size} bytes")
Output:
VPN=3, offset=57
A: physical address=28729
B: physical address=77881
buffer: pages=3, capacity=12288, slack=2288
two resident pages: 8192 bytes
If only two buffer pages are resident in RAM, that resident data occupies 8192 bytes. Here 10000 is the requested data size, 12288 is the mapped region’s capacity in the model, and 8192 is its assumed resident size. The calculation excludes page-table and allocator overhead. It is also not a measurement of real Python objects. Requested bytes, virtual mappings, and physical residency answer different questions.
What happens on a page fault
Accessing a valid but nonresident page triggers a page fault and transfers control to the kernel. Beyond Physical Memory: Mechanisms describes the case where a page was swapped to disk: read it back, update the page table, and retry the instruction. While disk I/O is pending, the OS can run other ready tasks. An access without a valid mapping, or one that violates permissions, cannot be handled simply by reading a page back.
A fault may require no disk read. Complete Virtual Memory Systems describes demand zeroing, which allocates and zeroes a physical page on first access, and copy-on-write (COW), which initially shares a page and copies it for a process when that process writes. Linux’s fork() manual confirms its use of COW: creating a child need not immediately copy all private physical pages. A COW write does not become a shared update visible to both parent and child.
File descriptors, pipes, and inheritance
On Unix, file descriptors are nonnegative integers a process uses to refer to open resources. Standard input, standard output, and standard error conventionally use 0, 1, and 2 at program startup. After fork(), the child’s descriptor copies refer to the same open file descriptions as the parent’s, including a shared file offset. Private memory and shared open resources can coexist.
Program replacement has inheritance rules too. Linux’s execve() manual says open descriptors normally survive, while those marked close-on-exec are closed. A child created by fork() inherits a copy of its parent’s environment and its working directory. execve() preserves the working directory, but the caller supplies the new program’s environment. Launchers such as Python’s subprocess can also close unwanted descriptors or specify a different environment and working directory. Starting a separate process does not itself revoke access granted by these resources.
A pipe connects a read end to a write end with a one-way byte stream. It is not a variable both sides manipulate, and has no built-in message boundaries; structured data needs an agreed format. Once buffered data is exhausted, a reader sees end-of-file only when all write-end descriptors have closed. One unnecessary writer left open can keep it waiting.
Run a child and collect data and exit status
This example uses only the Python standard library. The parent encodes a list as JSON and sends it through a standard-input pipe to a new Python process. The child decodes its own list, appends 7, writes its result to standard output and the example environment value to standard error, then deliberately exits with status 3. This transfers serialized data; it does not directly share the parent’s list.
The subprocess documentation specifies that communicate() sends input, reads output, and waits for termination. It suits this small payload; it buffers output in the parent’s memory and is unsuitable for unlimited streams.
import os
import subprocess
import sys
import json
numbers = [2, 3, 5]
child_code = r'''
import json, os, sys
numbers = json.load(sys.stdin)
numbers.append(7)
print(json.dumps({"numbers": numbers, "sum": sum(numbers)}))
print("mode=" + os.environ["DEMO_MODE"], file=sys.stderr)
raise SystemExit(3)
'''
env = dict(os.environ, DEMO_MODE="worker")
child = subprocess.Popen(
[sys.executable, "-c", child_code],
stdin=subprocess.PIPE, stdout=subprocess.PIPE,
stderr=subprocess.PIPE, text=True, env=env,
)
out, err = child.communicate(json.dumps(numbers))
print("child stdout:", out.strip())
print("child stderr:", err.strip())
print("exit status:", child.returncode)
print("parent numbers:", numbers)
Output (Python 3.13.15, macOS):
child stdout: {"numbers": [2, 3, 5, 7], "sum": 17}
child stderr: mode=worker
exit status: 3
parent numbers: [2, 3, 5]
The result 17 is 2 + 3 + 5 + 7; the parent’s list keeps its original values. Status 3 is a deliberately chosen nonzero status. A result on stdout does not make the run successful. The parent collects data and waits for termination before printing the four lines, so their displayed order does not depend on scheduling order.
Calling only wait() without reading the pipes can deadlock on large output: the child fills a pipe and waits for a reader while the parent waits for the child. communicate() handles the pipes together to avoid this wait. Python may launch a child using mechanisms such as posix_spawn(); this example is not a trace proving that fork() was called underneath.
Process, container, and virtual-machine boundaries
Docker’s explanation of containers and VMs distinguishes containers sharing a kernel from virtual machines running their own. Linux containers share the Linux kernel that hosts them. If that Linux environment runs inside a VM, the containers share that VM’s kernel.
Docker’s architecture documentation explains the use of namespaces to establish isolated resource views. Its resource-constraints documentation separately describes memory and CPU limits. By default, Docker containers have no resource limits, although actual use remains subject to the host kernel; the corresponding limits also need kernel support. A separate filesystem view does not automatically bring an exclusive or limited memory pool.
Choose the boundary according to the problem:
- Notebook variables belong to the process executing them. Jupyter Architecture explains the relationship between files and runtime state.
- For service containers and persistent storage, see Docker Engine and Compose. A memory budget must also say whether it limits one process or the entire group in a container.
- An agent’s tool command may run in a separate process while retaining inherited environment values, open resources, and account permissions. Agent Harness explains how an external system constrains execution scope.
- Computer Science Fundamentals connects these concepts to the wider learning path through representation, execution, and resources.