Evidence

Current qualification evidence

What MER has demonstrated—and what remains open.

The current evidence covers a frozen Qwen3-Coder Q4_0 workload on NVIDIA L4: deterministic output, GPU-native autoregressive execution, out-of-core expert staging, residency-performance behavior, and a measured GPU dispatch optimization.

  • Qwen3-Coder-30B-A3B-Instruct
  • Pure Q4_0
  • NVIDIA L4 / WGPU / Vulkan

Current qualification workload

These figures describe one fixed model, artifact, workload, and device—not product-wide performance.

15.21 GiB

Routed expert footprint

Pure Q4_0 expert weights in the current qualification artifact

6,144

Routed experts

48 MoE layers × 128 experts per layer

Top-8

Experts per token and layer

Learned routing for the fixed Qwen3-Coder workload

NVIDIA L4

Qualification accelerator

Physical hardware using MER’s WGPU / Vulkan backend

Model
Qwen3-Coder-30B-A3B-Instruct
Architecture
qwen3_moe
Quantization
Pure Q4_0
Geometry
48 MoE layers, 128 routed experts per layer, top-k 8, d_model 2,048, expert FFN dimension 768
Engine
Rust
Backend
WGPU / Vulkan

A / Correctness

Deterministic generated output is checked.

Frozen qualification workloads preserve the expected generated token and text output as execution and residency behavior change. This provides a deterministic engineering guardrail for the current workload.

B / GPU execution

Full autoregressive qualification now runs on the L4.

Attention, routing, routed Q4_0 expert execution, residency checks, recovery, combine, and token progression participate in an operational GPU-native autoregressive path. The result establishes a hardware-qualified prototype, not general production readiness.

C / Out-of-core proof

The expert pool exceeded both managed VRAM and bounded host memory.

Routed expert footprint
Approximately 15.21 GiB
Managed expert-residency VRAM
2 GiB
Host-memory ceiling
12 GiB
Swap events
Zero
OOM events
Zero
Measured tier movement
Local NVMe → RAM → GPU staging

D / Residency and performance

Throughput rises as residency misses collapse.

Decode throughput by expert-residency budget

Single frozen Qwen3-Coder Q4_0 qualification workload on NVIDIA L4.

Frozen NVIDIA L4 residency-capacity qualification points
Managed expert VRAMApproximate decode throughputQualification scope
2 GiB~1.10 tokens/sConstrained residency
4 GiB~1.25 tokens/sConstrained residency
8 GiB~1.71 tokens/sConstrained residency
12 GiB~2.67 tokens/sConstrained residency
14 GiB~3.90 tokens/sConstrained residency
15.21 GiB~5.59 tokens/sWarm zero-miss reference

E / GPU optimization

Fewer routed dispatches improved the measured L4 workload.

16 → 2

Dispatches per layer

Qualified routed expert dispatch count before and after the optimization

87.5%

Dispatch reduction

One-eighth as many routed dispatches per layer

+7.36%

Decode throughput

Measured improvement on the fixed NVIDIA L4 qualification workload

+7.32%

End-to-end throughput

Measured improvement on the same workload

This is one configuration-specific engineering result. It is not an industry benchmark or a claim about other models and hardware.

Methodology

Frozen workload, explicit boundaries.

Model and format
Qwen3-Coder-30B-A3B-Instruct, pure Q4_0
Hardware
Physical NVIDIA L4
Backend
WGPU / Vulkan
Correctness
Expected deterministic token and text output checked on frozen runs
Residency curve
One fixed workload with predictive movement disabled
Publication boundary
Configuration-specific internal qualification; no universal benchmark

Current engineering direction

Constrained-residency performance is the frontier.

The basic architecture is established. Current work aims to improve constrained-residency performance, move the right expert data earlier when useful, qualify additional environments, and turn the research runtime into a reproducible private-alpha package.

  • Reduce the performance penalty of constrained expert residency
  • Measure route-aware expert movement without making a public speedup claim
  • Improve storage and transfer overlap where the architecture permits
  • Improve GPU kernel and dispatch efficiency
  • Reduce synchronization and recovery overhead
  • Expand hardware portability beyond the current L4 environment, including AWS
  • Prepare reliable packaging and controlled private-alpha evaluations

MER remains pre-release systems research. AWS inference and performance qualification remain ahead. There is no general availability or production SLA.