Execution path

How MER runs beyond VRAM.

A sparse Mixture-of-Experts model activates only part of its expert pool for each token. MER uses that selectivity to keep active experts in GPU memory while bounded RAM and local NVMe back a larger routed expert namespace.

  • Sparse MoE routing
  • VRAM / RAM / NVMe
  • GPU-native on NVIDIA L4

From token to selected expert

The model router selects the expert subset. MER then checks residency, runs hits from the GPU tier, and stages misses from RAM or NVMe into the managed GPU arena before execution continues.

Conceptual token-to-expert path
  1. 01Route

    The learned model router selects the expert subset for this token

  2. 02Check

    MER checks the layer’s managed expert residency

  3. 03Hit

    Resident experts execute directly from the GPU tier

  4. 04Source

    A missing expert comes from bounded RAM or local NVMe

  5. 05Stage

    MER stages the expert into the managed GPU arena

  6. 06Continue

    Routed expert execution completes and token progression continues

  7. 07Under measurement

    Route-aware expert movement is being tested; no public speedup claim

Validated today

The full GPU-native qualification loop is operational.

Attention, routing, native Q4_0 routed expert execution, residency checks, recovery, combine, and token progression now participate in a real autoregressive GPU-native path.

That path has been exercised on physical NVIDIA L4 hardware with MER’s current Qwen3-Coder qualification workload. Current engineering focuses on performance as the managed expert-residency budget becomes more constrained.

Hit or miss

Residency changes underneath GPU execution.

VRAM hit

If a selected expert is already resident in the layer’s managed GPU arena, MER uses it directly while autoregressive state remains on the GPU path.

Expert miss

MER sources the missing expert from bounded RAM or local NVMe, stages it into the managed GPU arena, and continues execution. A miss is slower than a resident hit.

Maximum safe occupancy

Keep useful expert weights in the VRAM that remains.

MER first reserves headroom for non-expert weights, context and active-request memory, temporary GPU work, residency operations, and allocator or driver safety margin.

The remaining safe budget becomes managed routed-expert residency. More expert VRAM can retain a larger working set, reduce misses, and improve throughput without changing the larger NVMe-backed capacity architecture.

Route-aware movement

Test earlier movement of likely-needed experts.

MER is testing route-aware expert movement: using observed routing behavior to move likely-needed expert data earlier within the available resource budget. Selected experts remain on the critical path.

The mechanism is under active measurement. Earlier movement is not yet a public speedup claim or a performance guarantee.

Correctness

Every faster path must preserve expected generated output.

Frozen qualification workloads check deterministic generated token and text output while residency capacity and execution architecture change. This is a strict engineering check, not a claim of formal mathematical equivalence across every platform or model.

What MER is not

  • Not an LLM wrapper, hosted model API, or new foundation model.
  • Not a claim that SSD performs like VRAM or that every miss is hidden.
  • Not a claim that a complete 30B model uses only 2 GiB of total GPU memory.
  • Not generally available, production-ready, or universally hardware-compatible.