Technical overview

A GPU-owned loop over virtualized expert memory.

MER is an experimental GPU-native inference runtime written in Rust for sparse Mixture-of-Experts models. Its operational qualification path combines GPU-owned autoregressive execution with managed expert residency across VRAM, bounded RAM, and local NVMe.

  • Validated capability
  • Active performance work
  • Future productization

Execution architecture

The compute, residency, and control planes have different ownership boundaries. The labels below separate validated behavior from active optimization and later productization.

MER architecture with current status
GPU owned compute planeRun the autoregressive token loop and routed expert work on the accelerator.
  • Validated

    Attention, routing, native Q4_0 experts, combine, and token progression

  • Active optimization

    Kernel efficiency, synchronization reduction, and constrained-residency throughput

  • Qualification

    NVIDIA L4 through WGPU and Vulkan on the current Qwen3-Coder workload

MER memory / residency planeVirtualize the routed expert namespace below the GPU working set.
  • Validated

    Bounded GPU residency, RAM staging/cache, and NVMe-authoritative backing

  • Miss path

    Measured NVMe → RAM → GPU staging into the managed expert arena

  • Under measurement

    Route-aware expert movement; no public speedup claim

CPU control planeService storage and recovery without owning normal warm-layer arithmetic.
  • Current

    Startup, loading, storage service, orchestration, telemetry, and status

  • Correctness

    Frozen output checks across qualification changes

  • Productization

    Broader device coverage, deployment tooling, and operational packaging

These labels separate validated capability, active performance work, and future productization. They do not imply general production readiness or universal hardware support.

Warm GPU path
Attention → router → resident Q4_0 experts → combine → token progression.
Physical miss path
A missing routed expert is sourced from bounded RAM or local NVMe and staged into the managed GPU arena.
Validated today
GPU-native autoregressive execution, bounded residency, measured miss recovery, and deterministic output checks on NVIDIA L4.
Active optimization
Performance under constrained expert residency, route-aware expert movement, storage and transfer overlap, kernel and dispatch efficiency, and portability across hardware environments.
Future productization
Broader hardware qualification, automated hardware and model profiling where appropriate, packaging and installability, operator-facing diagnostics, controlled private alpha, and eventual commercial support.

VRAM budget policy

Spend the remaining safe VRAM on routed-expert residency.

Reserve first

Non-expert GPU requirements

Account for dense state, context and active requests, temporary GPU work, in-flight residency operations, and allocator or driver safety margin.

Allocate next

Managed expert arena

Give the remaining safe budget to layer-qualified routed-expert residency. This budget is distinct from total GPU memory and total application VRAM use.

Back the rest

RAM and local NVMe

RAM stages and caches expert payloads while local NVMe provides authoritative backing for the larger routed expert namespace.

I/O path

Storage joins the routed workload.

MER uses direct-I/O-capable local storage as authoritative backing for routed expert data. RAM provides bounded cache and staging capacity between that backing store and the GPU arena.

Routed inference is shaped by random expert access, routing locality, queue contention, filesystem behavior, and concurrent GPU work. Sequential drive specifications do not predict this workload by themselves.

Validated

Current hardware-qualified capabilities.

  • GPU-native autoregressive token loop
  • Model attention, router, top-k selection, and token progression
  • Native Q4_0 routed expert execution and weighted combine
  • Bounded, layer-qualified expert residency in a managed GPU arena
  • NVMe-authoritative expert backing with RAM cache and staging
  • Measured physical miss service and recovery
  • NVIDIA L4 qualification through WGPU and Vulkan
  • Deterministic generated-output checks on frozen workloads

Capability status

Working engine, active qualification.

Current MER capability status
StatusScopeMeaning
ValidatedGPU-native Qwen3-Coder qualification on NVIDIA L4Full autoregressive workloads complete with expected deterministic output.
ValidatedVirtualized routed-expert memoryExpert staging spans local NVMe, bounded RAM, and managed GPU residency.
Active optimizationSeverely constrained expert residencyClose the gap to the warm zero-miss compute-side reference.
Active qualificationModels and hardwareExpand beyond the current Qwen3-Coder and NVIDIA L4 validation scope.
Future productizationDeployment and supportPrepare packaging, operator diagnostics, and a controlled private alpha.