AMALGAFYMER / Micro-Expert-Router

Enter the expert gate.

Scroll to enter

Operator-Controlled AI Infrastructure

Use all your VRAM. Run beyond it.

MER is Amalgafy's experimental GPU-native inference runtime for sparse Mixture-of-Experts models. It keeps the active expert working set in GPU memory while bounded RAM and local NVMe back a larger expert pool, extending useful model capacity beyond available VRAM.

Conceptual expert-memory flow. NVMe backs the larger expert pool, RAM stages missing experts, and selected experts execute in VRAM. Experts already resident in VRAM can run immediately without a storage transfer. The animation illustrates routing, not measured timing.

Partnerships & Programs

MER engine features

Capacity

Run models beyond VRAM

Keep the active expert working set in GPU memory while bounded RAM and local NVMe back a larger expert pool.

See the memory model
GPU execution

Keep active work on the GPU

The qualified GPU-native path keeps the token loop and active expert execution on the GPU.

Review the qualified path
Memory hierarchy

Move only what is needed

Use a resident expert immediately; stage a missing expert from RAM or NVMe toward VRAM as demand changes.

Explore hit and miss flow
Validation

Prove constrained residency

A 15.21 GiB routed expert pool continued inference with a 2 GiB expert-VRAM budget and a 12 GiB host-memory ceiling.

Inspect the qualification evidence

How MER works

Route precisely. Move less.

When the model needs an expert, MER checks the GPU first. A resident expert can run immediately; a missing expert moves from bounded RAM or local NVMe into the managed GPU residency. MER is also testing route-aware expert movement using observed routing behavior. The mechanism is under active measurement; it is not yet a public speedup claim.

01

Select

A sparse MoE router chooses the small subset of expert networks needed for each token.

02

Use or move

Use an expert immediately if it is in VRAM; otherwise move it from RAM or load it from NVMe through RAM.

03

Execute

Run the selected expert on the GPU. Route-aware movement is being tested to place likely-needed expert data earlier within the available resource budget.

Why MER

Use the VRAM you have. Virtualize the rest.

VRAM remains the fastest place to work. MER fills its managed expert budget with the active working set while operator-controlled RAM and local storage back a larger routed expert namespace. Fully resident execution remains faster when everything comfortably fits.

01

Use the VRAM you have

Reserve safe space for the model core, context, active requests, temporary GPU work, and transfers.

02

Virtualize the rest

Use nearly all remaining safe VRAM for active experts while RAM and NVMe hold the larger pool.

03

Measure every miss

More VRAM should reduce miss frequency and improve performance; storage latency remains visible and workload-dependent.

Measured evidence

A real engine, not a concept deck.

MER has moved beyond architecture bring-up. Its real Qwen3-Coder qualification workload runs through a GPU-owned token path in repeated qualification on physical NVIDIA L4 hardware, including constrained expert residency and miss recovery. Current work focuses on performance, portability, and private-alpha readiness.

2 GiB

Constrained expert VRAM

Out-of-core qualification used a managed expert-residency budget far below the routed expert footprint.

See the out-of-core proof
NVIDIA L4

Real hardware

GPU-native autoregressive execution and NVMe-to-RAM-to-GPU expert staging were exercised on physical hardware.

Review the GPU execution scope
Exact

Deterministic output

Frozen qualification workloads preserve expected generated token and text output as residency architecture changes.

Review the correctness scope
Current workload

Qwen3-Coder-30B-A3B-Instruct, pure Q4_0, 6,144 routed experts.

Constrained memory

2 GiB managed expert residency and a 12 GiB host-memory ceiling in the existing constrained qualification.

Qualified stage

Pre-release systems research with real-model, real-hardware qualification.

Active frontier

Constrained-residency performance, route-aware expert movement, portability, and private-alpha readiness.

Private alpha preparation

Bring us the constraint MER is built to test.

MER is being prepared for a controlled private alpha. We are reviewing suitable sparse-MoE workloads, hardware profiles, and memory constraints while portability and performance qualification continue.

Private alpha preparation. Workload review requests go to sales@amalgafy.com.