About the role
Help build a benchmark of real, expert-level problems that today’s best AI agents still can’t solve. Experts take a genuinely hard problem from their own field and turn it into a task an AI must solve from scratch in a Linux terminal.
The goal: stump the model. A task counts only if, when three frontier AI models each try it three times, they mostly fail. It’s not annotation — contributors design the problems that expose where AI still breaks.
Who we need:
* Masters degree or higher * 10+ years of professional experience in a technical domain * At least one publication (academic or professional)
Domains in demand:
* Coding * Machine Learning * Systems * Security * Hardware
* This is an ongoing project — at least 7 weeks * Onboarding process is approximately 90 minutes total
This enterprise client helps the world’s most innovative companies improve their AI models by providing human feedback.
Originally posted on Himalayas