Falsifiable next step
Now test the reason to build it.
The next experiment must use a model that cannot fit in the test GPU’s VRAM and compare against the practical fallback: CPU/RAM offload.
Claim: old hardware can run an oversized MoE correctly
Decision rule written before the run
A model that fails resident GPU allocation completes end-to-end with same-quant correctness and bounded GPU memory.
The hybrid cannot generate correctly or exceeds the small GPU memory budget.
Capacity works correctly but has severe speed, concurrency, memory-tier, energy, or hardware-generation limits. Report them without erasing the capacity result.