Hardware

VRAM is the hot tier, not the capacity ceiling.

MER plans around the whole system: total GPU memory, usable application VRAM, a managed routed-expert residency budget, system RAM, local NVMe, and the request shape.

  • GPU and usable VRAM
  • Bounded system RAM
  • Local NVMe
  • Request shape

The system is one connected path.

The target memory and control path
GPU / VRAMHot transformer compute, request state, and managed expert residencySystem RAMExpert staging, cache, miss buffers, and host headroomNVMe / SSDCapacity backing the larger routed-expert namespaceCPU controlStartup, storage path, miss orchestration, and telemetryRequest shapeContext length, active requests, concurrency, and temporary GPU work

What matters

Plan around the full hierarchy.

01

GPU / VRAM

  • Backend and compute support
  • Total and usable application VRAM
  • Dense state, context, and temporary GPU work
  • Managed routed-expert residency budget
02

System RAM

  • Expert staging and cache
  • Miss buffering and upload headroom
  • Model and operating-system headroom
  • A bounded ceiling for predictable operation
03

Local NVMe / SSD

  • Authoritative routed-expert backing
  • Direct-I/O-capable storage path
  • Random-read latency and contention
  • Sustained routed-workload behavior
04

CPU / host control

  • Startup and model loading
  • Storage and miss orchestration
  • Telemetry and status
  • Headroom for reliable service
05

Request shape

  • Context length
  • Memory for active requests
  • Concurrency
  • Temporary GPU work and safety reserves

Current qualification device

NVIDIA L4 validation, not a universal minimum.

Accelerator
NVIDIA L4
Execution backend
WGPU / Vulkan
Current model
Qwen3-Coder-30B-A3B-Instruct, pure Q4_0
Qualified scope
Full GPU-native autoregressive execution with routed expert residency, staging, and deterministic output checks
Still expanding
Portability beyond the current NVIDIA L4 / GCP environment, including AWS, on the path toward controlled private alpha. AWS inference and performance qualification remain ahead.

One qualified device establishes the current hardware result. It does not define a universal minimum GPU or imply that every accelerator is supported.

Constrained qualification

2 GiB of expert residency is not 2 GiB of total GPU use.

MER managed a 15.21 GiB routed-expert pool with a 2 GiB expert-residency VRAM budget and a 12 GiB host-memory ceiling while recording zero swap and zero OOM events. The workload measured NVMe-to-RAM-to-GPU expert staging as inference continued.

VRAM planning

Reserve, then fill the safe remainder.

1 / Non-expert state
Dense and other model requirements
2 / Requests
Context, KV state, and active requests
3 / Workspaces
Kernels and temporary GPU buffers
4 / Residency operations
In-flight staging and replacement
5 / Safety
Allocator, runtime, and driver margin
6 / Expert budget
The remaining safe VRAM for managed routed-expert residency