About the role
About the Role
Join a small, senior team building autonomous AI agents for executive productivity. You will develop new agent capabilities and create reliable ways to evaluate them, helping ensure that important real-world tasks are completed accurately before new capabilities reach users.
What You'll Do
* Build agents that complete delegated work end to end without requiring the user to check each step.
* Develop trustworthy evaluations that measure how accurately agents perform different tasks.
* Ship new agent capabilities in the tools people already use, supported by evidence that they are ready.
* Investigate production failures and deliver durable, validated fixes.
* Build large-scale failure detection to identify known patterns and unexpected issues.
* Improve agent recovery when a tool or website fails.
* Create repeatable workflows so engineers can measure whether agent changes improve performance.
What We're Looking For
* Approximately 2 to 11 years of relevant experience, including a few years in one role with demonstrated progression.
* Experience building or evaluating production AI or machine learning systems, especially agents or LLM-powered products.
* Strong Python skills and familiarity with Django, React, and TypeScript.
* Experience with evaluation and monitoring tools such as Braintrust or Raindrop is relevant.
* Comfort working across implementation, evaluation, and iteration in a fast-moving environment.
Compensation & Benefits
Salary range: $200,000 to $300,000 USD annually. Visa sponsorship is not available.
Location
Hybrid, with four days per week in the office in Palo Alto, California, United States.