Senior AI Infrastructure Engineer, Kubernetes
Firmus
Sydney, Australia · Onsite · Full Time
Posted
Job description
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. Firmus AI Cloud Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. Role Summary The Senior Kubernetes Engineer, AI Infrastructure owns the technical design and delivery of the backend infrastructure that powers the Firmus Kubernetes platform. This is a hands-on principal-level individual contributor role, responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities across GPU-accelerated bare-metal environments. They solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign-off for platform designs. They work across AI Platforms, Solutions Architecture & Delivery, networking, security, and operations to create a secure, resilient, multi-tenant platform that can be deployed and operated consistently at AI-factory scale. Key Responsibilities Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design. Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure…