VRAM hit
If a selected expert is already resident in the layer’s managed GPU arena, MER uses it directly while autoregressive state remains on the GPU path.
Execution path
A sparse Mixture-of-Experts model activates only part of its expert pool for each token. MER uses that selectivity to keep active experts in GPU memory while bounded RAM and local NVMe back a larger routed expert namespace.
The model router selects the expert subset. MER then checks residency, runs hits from the GPU tier, and stages misses from RAM or NVMe into the managed GPU arena before execution continues.
The learned model router selects the expert subset for this token
MER checks the layer’s managed expert residency
Resident experts execute directly from the GPU tier
A missing expert comes from bounded RAM or local NVMe
MER stages the expert into the managed GPU arena
Routed expert execution completes and token progression continues
Route-aware expert movement is being tested; no public speedup claim
Validated today
Attention, routing, native Q4_0 routed expert execution, residency checks, recovery, combine, and token progression now participate in a real autoregressive GPU-native path.
That path has been exercised on physical NVIDIA L4 hardware with MER’s current Qwen3-Coder qualification workload. Current engineering focuses on performance as the managed expert-residency budget becomes more constrained.
Hit or miss
If a selected expert is already resident in the layer’s managed GPU arena, MER uses it directly while autoregressive state remains on the GPU path.
MER sources the missing expert from bounded RAM or local NVMe, stages it into the managed GPU arena, and continues execution. A miss is slower than a resident hit.
Maximum safe occupancy
MER first reserves headroom for non-expert weights, context and active-request memory, temporary GPU work, residency operations, and allocator or driver safety margin.
The remaining safe budget becomes managed routed-expert residency. More expert VRAM can retain a larger working set, reduce misses, and improve throughput without changing the larger NVMe-backed capacity architecture.
Route-aware movement
MER is testing route-aware expert movement: using observed routing behavior to move likely-needed expert data earlier within the available resource budget. Selected experts remain on the critical path.
The mechanism is under active measurement. Earlier movement is not yet a public speedup claim or a performance guarantee.
Correctness
Frozen qualification workloads check deterministic generated token and text output while residency capacity and execution architecture change. This is a strict engineering check, not a claim of formal mathematical equivalence across every platform or model.