About the role
About the Role
Join an early-stage AI infrastructure team as its first QA engineer, owning quality and test automation across a GPU fleet management platform. Your work will help ensure that conversational and web interfaces, backend services, and Kubernetes operations behave reliably and safely in production.
What You'll Do
* Test conversational workflows end to end, including command accuracy and actions performed on live clusters.
* Build evaluations for AI agent behavior, checking response quality, safe actions, and handling of ambiguous or invalid requests.
* Create end-to-end tests for the web interface and verify consistent behavior across product interfaces.
* Write Go integration tests for backend workflows such as installation, upgrades, and failover.
* Provision and tear down Kubernetes test clusters to validate controllers, operators, and custom resources across versions.
* Verify alert delivery across product interfaces and chat, and ensure alerts fire appropriately.
* Gate releases in CI and turn bugs and incidents into regression tests.
What We're Looking For
* At least 3 years of experience in QA, SDET, or test automation, including ownership of product quality and test suites.
* Strong Go skills for integration testing or backend test automation, plus experience building end-to-end web UI tests.
* Hands-on experience with Kubernetes, React, RPC systems such as gRPC, CI/CD, and Linux fundamentals.
* Familiarity with monitoring and alerting tools such as Prometheus, Alertmanager, or Grafana.
* Experience testing conversational or LLM-powered products, building agent evaluations, or working with GPU infrastructure is valuable.
* Exposure to multicluster tooling, chaos engineering, Kubernetes testing tools, and UI automation frameworks is a plus.
Compensation & Benefits
Annual salary: $180,000 to $200,000 USD. Visa sponsorship is not available.
Location
On-site in San Francisco, California.