Non-expert GPU requirements
Account for dense state, context and active requests, temporary GPU work, in-flight residency operations, and allocator or driver safety margin.
Technical overview
MER is an experimental GPU-native inference runtime written in Rust for sparse Mixture-of-Experts models. Its operational qualification path combines GPU-owned autoregressive execution with managed expert residency across VRAM, bounded RAM, and local NVMe.
The compute, residency, and control planes have different ownership boundaries. The labels below separate validated behavior from active optimization and later productization.
Attention, routing, native Q4_0 experts, combine, and token progression
Kernel efficiency, synchronization reduction, and constrained-residency throughput
NVIDIA L4 through WGPU and Vulkan on the current Qwen3-Coder workload
Bounded GPU residency, RAM staging/cache, and NVMe-authoritative backing
Measured NVMe → RAM → GPU staging into the managed expert arena
Route-aware expert movement; no public speedup claim
Startup, loading, storage service, orchestration, telemetry, and status
Frozen output checks across qualification changes
Broader device coverage, deployment tooling, and operational packaging
These labels separate validated capability, active performance work, and future productization. They do not imply general production readiness or universal hardware support.
VRAM budget policy
Account for dense state, context and active requests, temporary GPU work, in-flight residency operations, and allocator or driver safety margin.
Give the remaining safe budget to layer-qualified routed-expert residency. This budget is distinct from total GPU memory and total application VRAM use.
RAM stages and caches expert payloads while local NVMe provides authoritative backing for the larger routed expert namespace.
I/O path
MER uses direct-I/O-capable local storage as authoritative backing for routed expert data. RAM provides bounded cache and staging capacity between that backing store and the GPU arena.
Routed inference is shaped by random expert access, routing locality, queue contention, filesystem behavior, and concurrent GPU work. Sequential drive specifications do not predict this workload by themselves.
Validated
Capability status
| Status | Scope | Meaning |
|---|---|---|
| Validated | GPU-native Qwen3-Coder qualification on NVIDIA L4 | Full autoregressive workloads complete with expected deterministic output. |
| Validated | Virtualized routed-expert memory | Expert staging spans local NVMe, bounded RAM, and managed GPU residency. |
| Active optimization | Severely constrained expert residency | Close the gap to the warm zero-miss compute-side reference. |
| Active qualification | Models and hardware | Expand beyond the current Qwen3-Coder and NVIDIA L4 validation scope. |
| Future productization | Deployment and support | Prepare packaging, operator diagnostics, and a controlled private alpha. |