Routed expert footprint
Pure Q4_0 expert weights in the current qualification artifact
Evidence
Current qualification evidenceThe current evidence covers a frozen Qwen3-Coder Q4_0 workload on NVIDIA L4: deterministic output, GPU-native autoregressive execution, out-of-core expert staging, residency-performance behavior, and a measured GPU dispatch optimization.
These figures describe one fixed model, artifact, workload, and device—not product-wide performance.
Pure Q4_0 expert weights in the current qualification artifact
48 MoE layers × 128 experts per layer
Learned routing for the fixed Qwen3-Coder workload
Physical hardware using MER’s WGPU / Vulkan backend
qwen3_moeA / Correctness
Frozen qualification workloads preserve the expected generated token and text output as execution and residency behavior change. This provides a deterministic engineering guardrail for the current workload.
B / GPU execution
Attention, routing, routed Q4_0 expert execution, residency checks, recovery, combine, and token progression participate in an operational GPU-native autoregressive path. The result establishes a hardware-qualified prototype, not general production readiness.
C / Out-of-core proof
D / Residency and performance
Single frozen Qwen3-Coder Q4_0 qualification workload on NVIDIA L4.
| Managed expert VRAM | Approximate decode throughput | Qualification scope |
|---|---|---|
| 2 GiB | ~1.10 tokens/s | Constrained residency |
| 4 GiB | ~1.25 tokens/s | Constrained residency |
| 8 GiB | ~1.71 tokens/s | Constrained residency |
| 12 GiB | ~2.67 tokens/s | Constrained residency |
| 14 GiB | ~3.90 tokens/s | Constrained residency |
| 15.21 GiB | ~5.59 tokens/s | Warm zero-miss reference |
E / GPU optimization
Qualified routed expert dispatch count before and after the optimization
One-eighth as many routed dispatches per layer
Measured improvement on the fixed NVIDIA L4 qualification workload
Measured improvement on the same workload
This is one configuration-specific engineering result. It is not an industry benchmark or a claim about other models and hardware.
Methodology
Current engineering direction
The basic architecture is established. Current work aims to improve constrained-residency performance, move the right expert data earlier when useful, qualify additional environments, and turn the research runtime into a reproducible private-alpha package.
MER remains pre-release systems research. AWS inference and performance qualification remain ahead. There is no general availability or production SLA.