About the role
About the Role
This is an early-team platform engineering role at a seed-stage AI infrastructure startup building an open-source, GitOps-native distributed operating system on top of Kubernetes. You will work closely with the founding team to design and evolve core platform features, making multi-cluster management intuitive for developers running demanding AI workloads. The work spans the full infrastructure stack, from bare metal to production-grade Kubernetes.
What You'll Do
* Build and extend core platform features in Go, including custom Kubernetes operators and controllers.
* Design and implement GitOps workflows with ArgoCD to make continuous deployment seamless.
* Develop infrastructure-as-code patterns using Terraform and Helm to provision and manage clusters.
* Work on distributed storage solutions using Ceph and WEKA for high-performance, scalable cluster storage.
* Create observability and monitoring systems with Prometheus and Grafana to surface cluster health and performance.
* Build and optimize container networking with Cilium for network security and observability.
* Design and implement federated Kubernetes architectures for multi-cluster management.
* Build automation tooling that reduces operational overhead for developers running production workloads.
What We're Looking For
* 2 to 5+ years of software development experience, with a strong systems or infrastructure focus.
* Hands-on proficiency in Go, including writing Go in a Kubernetes environment.
* Production Kubernetes experience: managing clusters at meaningful scale, writing operators and controllers, working with CRDs.
* Experience with distributed storage solutions such as Ceph or WEKA.
* Experience designing federated Kubernetes architectures for multi-cluster management.
* Familiarity with GitOps workflows and ArgoCD.
* Experience with infrastructure-as-code tools such as Terraform and Helm; Ansible or Kubespray is a plus.
* Background with container networking, CNI plugins, or service mesh technologies such as Istio or Linkerd.
* Experience with cloud platforms (AWS, GCP, or Azure) and their managed Kubernetes offerings.
* Familiarity with observability tooling such as Prometheus and Grafana.
* Experience with GPU infrastructure or bare metal environments is a strong plus.
* Contributions to Kubernetes ecosystem tooling or similar open-source infrastructure projects are a bonus.
Compensation & Benefits
Salary range: $180,000 to $210,000 USD annually. Visa sponsorship is available.
Location
On-site in San Francisco, California. Candidates must be able to work in person full-time.