Optimizing Transformer Model Serving Parameters – An Apple Silicon GPU Case Study
July 11, 2026
One MacBook Pro M4 Max, an open-source MoE, and a hands-on exercise in optimizing it for inference: why the thing that collapses LLM throughput by 100× is prefill, not decode, and the two prefill-side fixes that got it back.