Run models beyond VRAM
Keep the active expert working set in GPU memory while bounded RAM and local NVMe back a larger expert pool.
See the memory modelEnter the expert gate.
Scroll to enterOperator-Controlled AI Infrastructure
MER is Amalgafy's experimental GPU-native inference runtime for sparse Mixture-of-Experts models. It keeps the active expert working set in GPU memory while bounded RAM and local NVMe back a larger expert pool, extending useful model capacity beyond available VRAM.
Conceptual expert-memory flow. NVMe backs the larger expert pool, RAM stages missing experts, and selected experts execute in VRAM. Experts already resident in VRAM can run immediately without a storage transfer. The animation illustrates routing, not measured timing.
Partnerships & Programs
Keep the active expert working set in GPU memory while bounded RAM and local NVMe back a larger expert pool.
See the memory modelThe qualified GPU-native path keeps the token loop and active expert execution on the GPU.
Review the qualified pathUse a resident expert immediately; stage a missing expert from RAM or NVMe toward VRAM as demand changes.
Explore hit and miss flowA 15.21 GiB routed expert pool continued inference with a 2 GiB expert-VRAM budget and a 12 GiB host-memory ceiling.
Inspect the qualification evidenceHow MER works
When the model needs an expert, MER checks the GPU first. A resident expert can run immediately; a missing expert moves from bounded RAM or local NVMe into the managed GPU residency. MER is also testing route-aware expert movement using observed routing behavior. The mechanism is under active measurement; it is not yet a public speedup claim.
A sparse MoE router chooses the small subset of expert networks needed for each token.
Use an expert immediately if it is in VRAM; otherwise move it from RAM or load it from NVMe through RAM.
Run the selected expert on the GPU. Route-aware movement is being tested to place likely-needed expert data earlier within the available resource budget.
Why MER
VRAM remains the fastest place to work. MER fills its managed expert budget with the active working set while operator-controlled RAM and local storage back a larger routed expert namespace. Fully resident execution remains faster when everything comfortably fits.
Reserve safe space for the model core, context, active requests, temporary GPU work, and transfers.
Use nearly all remaining safe VRAM for active experts while RAM and NVMe hold the larger pool.
More VRAM should reduce miss frequency and improve performance; storage latency remains visible and workload-dependent.
Measured evidence
MER has moved beyond architecture bring-up. Its real Qwen3-Coder qualification workload runs through a GPU-owned token path in repeated qualification on physical NVIDIA L4 hardware, including constrained expert residency and miss recovery. Current work focuses on performance, portability, and private-alpha readiness.
Current Qwen3-Coder Q4_0 qualification footprint across 6,144 routed experts.
Review the current workloadOut-of-core qualification used a managed expert-residency budget far below the routed expert footprint.
See the out-of-core proofGPU-native autoregressive execution and NVMe-to-RAM-to-GPU expert staging were exercised on physical hardware.
Review the GPU execution scopeFrozen qualification workloads preserve expected generated token and text output as residency architecture changes.
Review the correctness scopeQwen3-Coder-30B-A3B-Instruct, pure Q4_0, 6,144 routed experts.
2 GiB managed expert residency and a 12 GiB host-memory ceiling in the existing constrained qualification.
Pre-release systems research with real-model, real-hardware qualification.
Constrained-residency performance, route-aware expert movement, portability, and private-alpha readiness.
Private alpha preparation
MER is being prepared for a controlled private alpha. We are reviewing suitable sparse-MoE workloads, hardware profiles, and memory constraints while portability and performance qualification continue.