Clause
← All documentation

Execution model: frames and continuations

Decision of plan 11 step 17 (2026-10-05). It fixes how generated Erlang functions call, return, suspend and fail once recursion, tail calls and processes arrive. Step 19 implemented calls, returns, tail calls and frames (Implementation lists what is still open); steps 23, 24, 26 and 43 build on it.

Decision

Every Erlang process runs on its own flat stack of explicit frames, and generated code moves between functions only by guaranteed tail transfers (LLVM musttail). A non-tail call stores the caller's continuation in the caller's frame and jumps to the callee; a return jumps back to that continuation. The native stack therefore stays one call deep above the scheduler, whatever the Erlang recursion depth, and a process can stop at any transfer and resume later on any thread.

Process state

FieldMeaning
stack, capacityOne growable word array; moves when it grows
frame, topWord offsets of the current frame header and of the first free word
x[], liveArgument/result registers; the first live words are roots at a transfer
reductionsCalls left in the time slice
resume_atEntry a suspended process continues at
failure channelToday's revision-2 channel: reason, payload, stack trace, halted

Frames

A frame is a fixed header followed by the function's slots (BEAM Y registers). Slots are zeroed at push, as root frames are today.

Header wordMeaning
previousOffset of the caller's header (frames link by offsets, never pointers)
functionThe function's descriptor
resumeContinuation index the body switches on when control returns here
handlerContinuation index of the innermost active handler, 0 when none

The descriptor extends today's abi::v1::FrameDescriptor (module descriptor, module and function atom slots, arity) with the entry and body code pointers and the slot count. Header words are not terms; a walker follows previous and reads slot counts from descriptors. A runtime-owned bottom frame sits under the first call of every process: continuation 1 is a normal exit (result in x[0]), continuation 2 the handler for an uncaught exception.

Operations

A function that makes no non-tail call and keeps no slot across a safepoint may skip its frame and return straight to its caller's body. This is an optimization the compiler may add later, not part of the contract.

Root visibility

At every transfer and every safepoint the roots of a process are: all slots of all frames on its stack, x[0..live), and the failure channel's payload, argument list and stack term (already process roots). The result handoff words of today's root stack disappear; x[0] takes their role.

Values never survive a transfer in native registers or SSA values. Each body reloads its frame address from stack + frame after it is entered, and reads live values from slots. Within one continuation, a service that can push a frame or move the heap invalidates every slot pointer and heap pointer held in SSA values; the reload rule for collections is fixed in collection in generated code.

Successor of the 8F root stack

The segmented stack exists only because generated code holds absolute frame pointers across calls. Under this model no frame pointer survives a transfer, so the stack becomes one flat block per process:

Targets

musttail with the uniform void (Process *) signature is accepted by every required target at O0 and O2 (prototype below, clang 23.1.2):

TargetWordResult
x86_64-pc-windows-msvc, x86_64-unknown-linux-gnu64tail jumps
i686-pc-windows-msvc, i686-unknown-linux-gnu32tail jumps
aarch64-unknown-linux-gnu, arm64-apple-macosx14.064tail jumps
armv7-unknown-linux-gnueabihf32tail jumps

The backend reports an error when it cannot honour a musttail call, so a successful compile is the guarantee. Fallback for a future target that rejects it: a trampoline. Each code returns the next code pointer to the scheduler loop instead of jumping (null suspends); frames, roots and continuation indices are unchanged. No required target needs it.

Implementation

Step 19 (2026-10-05) implements the model with these choices and gaps:

Alternatives compared

Prototype in tests/prototypes/execution_model: the same Erlang functions (sum/1 body recursion, loop/2 tail recursion, fail/1 raising boom at depth N, catcher/1 catching it) hand-lowered three ways. Host: Windows x64, clang 23.1.2; times are single runs at O2. Run python tests/prototypes/execution_model/run.py.

Explicit frames + musttail (chosen)Native calls + root frames (today)LLVM coroutines (C++20)
1M-deep body recursionok, 17 ms; 5 words per frame; 17 stack moves80 B native stack per level: about 13,000 levels in a 1 MiB threadok, 60 ms; one heap allocation of 64–80 B per call
10M tail callsok, 5 ms; constant stackO0 grows 80 B per call; O2 only by sibling-call luckno tail calls: 1M iterations keep 1M frames
Yield / resumeevery entry; two processes interleaveimpossible without a native stack per processsymmetric transfer (itself musttail)
Native stack at depth 1M136–144 Bgrows per level144–520 B
Roots visible to GCslots in known framesslots in known framescoroutine frame layout chosen by LLVM; terms would need a second rooted copy
Exceptionsunwind to handler framechannel check per returnchannel check per return

The trampoline variant of the chosen model passes the same runs (20 ms recursion, 16 ms for 10M tail calls at O2). Rejected:

Accepted costs: one indirect jump per return plus a switch dispatch; values live across calls are reloaded from slots (they are stored there already); native debugger backtraces show only the current function, while Erlang stack traces come from the frame chain.

Prototype evidence

run.py builds the chosen model (musttail and trampoline), the native baseline and the coroutine model for the host at O0 and O2, runs them, and compiles the hand-lowered functions (generated.cpp, freestanding) for each target above at O0 and O2, counting musttail calls in the IR and tail jumps in the assembly. 2026-10-05 result: PASS. For the chosen model at both levels: return (42), 1,000,000-deep body recursion, 10,000,000 tail calls, an error raised at depth 100,000 caught by a handler frame, the same error uncaught (bottom frame, trace of 8 fail frames), and two processes interleaving in 4,000-reduction slices (1,002 slices). All 15 transfer sites are musttail at O0 on every target (14 at O2 after inlining).

The sum/1 body for i686-pc-windows-msvc (O2 IR, names shortened). The x86_64-pc-windows-msvc IR is the same with i64 words and doubled offsets.

%2 = load ptr, ptr %0, align 4                      ; stack base
%3 = getelementptr inbounds nuw i8, ptr %0, i32 8
%4 = load i32, ptr %3, align 4                      ; current frame offset
%5 = getelementptr inbounds nuw [4 x i8], ptr %2, i32 %4
%6 = getelementptr inbounds nuw i8, ptr %5, i32 16  ; slot 0 (N)
%7 = getelementptr inbounds nuw i8, ptr %5, i32 8   ; header: resume index
%8 = load i32, ptr %7, align 4
%9 = icmp eq i32 %8, 0                              ; switch on resume
...
16:                                                 ; N > 0: call sum(N - 1)
  store i32 1, ptr %7, align 4                      ; resume = 1
  ...                                               ; x0 = N - 1
  %19 = tail call ptr @call(ptr %0, ptr @SUM)       ; push frame or park
  musttail call void %19(ptr nonnull %0)
  ret void
20:                                                 ; resume 1: x0 += N
  ...
  %24 = tail call ptr @leave(ptr %0)                ; pop, caller body
  musttail call void %24(ptr nonnull %0)
  ret void