Local model execution
Run the engine and model weights on operator-controlled systems rather than depending on an external hosted inference API for every request.
Why MER
MER is an experimental GPU-native inference runtime written in Rust for sparse Mixture-of-Experts models. The active expert working set fits a managed GPU-residency budget while bounded host memory and local NVMe back the larger expert pool.
Fully resident accelerator inference requires every frequently needed weight to fit in accelerator memory. It remains the faster choice when the complete model comfortably fits and maximum latency or throughput is the priority.
Sparse MoE models expose a different operating point: keep the transformer state and active expert working set in GPU memory, while RAM and NVMe back the larger routed expert pool. That can make useful inference possible when the expert footprint exceeds the practical accelerator-memory envelope.
Deployment tradeoff
| Factor | Fully resident accelerator | MER virtualized residency |
|---|---|---|
| Model-fit requirement | Complete model fits accelerator memory | Active expert working set fits the managed VRAM budget |
| Expert residency | All routed expert weights resident | Demand-managed VRAM hot set backed by RAM and NVMe |
| Storage participation | Primarily startup and loading | Local NVMe supplies misses through RAM staging |
| When expert VRAM increases | Already-resident execution remains available | Larger hot set, fewer misses, and potentially higher throughput |
| Miss cost | No expert-residency miss in steady state | Workload-dependent storage, staging, and upload latency |
| Best fit | Maximum performance when the model fits | Useful, local inference under memory constraint |
Different GPUs provide different safe expert-residency budgets and therefore different miss rates and performance profiles. MER does not claim one universal model-to-VRAM multiplier or that every supported workload will behave like the current qualification.
Operator control
Run the engine and model weights on operator-controlled systems rather than depending on an external hosted inference API for every request.
Evaluation targets include workstations, on-premises systems, private cloud, disconnected research networks, edge sites, and sovereign infrastructure.
MER does not claim security, privacy, or compliance certifications that have not been completed.
Amalgafy
Amalgafy builds infrastructure for making advanced AI inference practical on hardware outside hyperscale data-center assumptions. MER is its current core systems project, focused on local inference, operator control, and better use of existing hardware.
Environmental positioning