A 128B parameter Mixture-of-Experts model balancing capability and efficiency, with strong performance on reasoning, coding, and multilingual tasks. Served via NVIDIA NIM with fast inference and optimized latency.