MER plans around the whole system: total GPU memory, usable application VRAM, a managed routed-expert residency budget, system RAM, local NVMe, and the request shape.
GPU and usable VRAM
Bounded system RAM
Local NVMe
Request shape
The system is one connected path.
The target memory and control path
GPU / VRAMHot transformer compute, request state, and managed expert residencySystem RAMExpert staging, cache, miss buffers, and host headroomNVMe / SSDCapacity backing the larger routed-expert namespaceCPU controlStartup, storage path, miss orchestration, and telemetryRequest shapeContext length, active requests, concurrency, and temporary GPU work
What matters
Plan around the full hierarchy.
01
GPU / VRAM
Backend and compute support
Total and usable application VRAM
Dense state, context, and temporary GPU work
Managed routed-expert residency budget
02
System RAM
Expert staging and cache
Miss buffering and upload headroom
Model and operating-system headroom
A bounded ceiling for predictable operation
03
Local NVMe / SSD
Authoritative routed-expert backing
Direct-I/O-capable storage path
Random-read latency and contention
Sustained routed-workload behavior
04
CPU / host control
Startup and model loading
Storage and miss orchestration
Telemetry and status
Headroom for reliable service
05
Request shape
Context length
Memory for active requests
Concurrency
Temporary GPU work and safety reserves
Current qualification device
NVIDIA L4 validation, not a universal minimum.
Accelerator
NVIDIA L4
Execution backend
WGPU / Vulkan
Current model
Qwen3-Coder-30B-A3B-Instruct, pure Q4_0
Qualified scope
Full GPU-native autoregressive execution with routed expert residency, staging, and deterministic output checks
Still expanding
Portability beyond the current NVIDIA L4 / GCP environment, including AWS, on the path toward controlled private alpha. AWS inference and performance qualification remain ahead.
One qualified device establishes the current hardware result. It does not define a universal minimum GPU or imply that every accelerator is supported.
Constrained qualification
2 GiB of expert residency is not 2 GiB of total GPU use.
MER managed a 15.21 GiB routed-expert pool with a 2 GiB expert-residency VRAM budget and a 12 GiB host-memory ceiling while recording zero swap and zero OOM events. The workload measured NVMe-to-RAM-to-GPU expert staging as inference continued.
VRAM planning
Reserve, then fill the safe remainder.
1 / Non-expert state
Dense and other model requirements
2 / Requests
Context, KV state, and active requests
3 / Workspaces
Kernels and temporary GPU buffers
4 / Residency operations
In-flight staging and replacement
5 / Safety
Allocator, runtime, and driver margin
6 / Expert budget
The remaining safe VRAM for managed routed-expert residency