About the role
About the Role
Join an early-stage team building infrastructure that makes production Kubernetes easier to deploy and manage for demanding AI workloads. You will help shape the platform layer and its multi-cluster capabilities, working across systems engineering, automation, and cloud-native infrastructure.
What You'll Do
* Build platform features in Go, including custom Kubernetes operators and controllers.
* Design GitOps workflows with ArgoCD to streamline continuous deployment.
* Develop infrastructure-as-code patterns with Terraform and Helm to provision and manage clusters.
* Work with distributed storage technologies such as Ceph and WEKA for scalable cluster storage.
* Create observability systems with Prometheus and Grafana to surface cluster health and performance.
* Build and optimize container networking with Cilium for security and observability.
* Design federated Kubernetes architectures for multi-cluster management.
* Develop automation that reduces operational work for teams running production workloads.
What We're Looking For
* At least 2 years of software development experience, with approximately 5 years preferred.
* Hands-on experience writing Go for Kubernetes environments and managing production-scale clusters.
* Experience with distributed storage such as Ceph or WEKA, federated Kubernetes, and multi-cluster management.
* Experience with GitOps and ArgoCD, plus infrastructure as code using Terraform, Helm, or comparable tools.
* Familiarity with cloud platforms, managed Kubernetes, container networking, or service mesh technologies.
* Experience with Prometheus and Grafana is useful, as are contributions to open-source infrastructure projects.
Compensation & Benefits
Annual salary of $180,000 to $210,000 USD. Visa sponsorship is available.
Location
This is a full-time, on-site role in San Francisco, California, United States.