Mirantis is building a new enterprise AI infrastructure product that lets organizations run and govern large language models on their own Kubernetes clusters. You will join a small senior team early, with broad ownership of the model-serving layer and its path to production.What you'll doDesign and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management.Package the platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks.Integrate the serving layer with the platform's API gateway, identity, and metering services.Build the observability for operating GPU inference in production (serving metrics, GPU telemetry).Contribute across a multi-service codebase and help set engineering direction through design docs and reviews.