Why MER

Model capacity should not stop at VRAM.

MER is an experimental GPU-native inference runtime written in Rust for sparse Mixture-of-Experts models. The active expert working set fits a managed GPU-residency budget while bounded host memory and local NVMe back the larger expert pool.

  • Active expert working set in VRAM
  • Larger namespace in RAM / NVMe
  • Operator-controlled inference

The model-fit equation changes.

Fully resident accelerator inference requires every frequently needed weight to fit in accelerator memory. It remains the faster choice when the complete model comfortably fits and maximum latency or throughput is the priority.

Sparse MoE models expose a different operating point: keep the transformer state and active expert working set in GPU memory, while RAM and NVMe back the larger routed expert pool. That can make useful inference possible when the expert footprint exceeds the practical accelerator-memory envelope.

Deployment tradeoff

Fully resident when it fits. Virtualized when capacity matters.

Fully resident accelerator deployment and MER virtualized expert residency
FactorFully resident acceleratorMER virtualized residency
Model-fit requirementComplete model fits accelerator memoryActive expert working set fits the managed VRAM budget
Expert residencyAll routed expert weights residentDemand-managed VRAM hot set backed by RAM and NVMe
Storage participationPrimarily startup and loadingLocal NVMe supplies misses through RAM staging
When expert VRAM increasesAlready-resident execution remains availableLarger hot set, fewer misses, and potentially higher throughput
Miss costNo expert-residency miss in steady stateWorkload-dependent storage, staging, and upload latency
Best fitMaximum performance when the model fitsUseful, local inference under memory constraint

More VRAM changes the hot set, not the thesis.

Different GPUs provide different safe expert-residency budgets and therefore different miss rates and performance profiles. MER does not claim one universal model-to-VRAM multiplier or that every supported workload will behave like the current qualification.

Operator control

Keep execution inside infrastructure you control.

Local model execution

Run the engine and model weights on operator-controlled systems rather than depending on an external hosted inference API for every request.

Deployment flexibility

Evaluation targets include workstations, on-premises systems, private cloud, disconnected research networks, edge sites, and sovereign infrastructure.

MER does not claim security, privacy, or compliance certifications that have not been completed.

Amalgafy

Infrastructure beyond hyperscale assumptions.

Amalgafy builds infrastructure for making advanced AI inference practical on hardware outside hyperscale data-center assumptions. MER is its current core systems project, focused on local inference, operator control, and better use of existing hardware.

Environmental positioning

Hardware utilization is an engineering objective.