About the role
Job Description
• Build and validate predictive models including censored bid-landscape modeling, contextual over-indexing, conversion propensity prediction with delayed labels, and positive-unlabelled learning • Design and implement offline evaluation frameworks using inverse propensity scoring and doubly-robust estimators over logged decisions • Define exploration strategies and propensity logging approaches to support reliable model evaluation and optimization • Calibrate and optimize models for individual advertisers while independently monitoring ranking and calibration quality • Develop and operate scalable training orchestration pipelines across hourly, daily, and weekly execution schedules • Build and maintain model registry workflows including lineage tracking, evaluation gates, and auditable promotion processes • Implement isolated per-advertiser model instances with dedicated configuration and namespace separation • Own model publishing pipelines with freshness SLO compliance and documented fallback procedures • Run shadow deployments and champion/challenger experiments with production-grade measurement logging • Monitor feature drift, prediction drift, train/serve skew, calibration decay, and label latency in production environments • Ensure reproducibility through pinned environments, containerized builds, and reproducible data snapshots • Participate in post-launch optimization cycles and evaluate business impact using statistically grounded lift measurements • Prepare technical documentation and support knowledge transfer to the Customer’s engineering and data teams
Qualifications
• 6+ years of combined commercial experience in Data Science and ML Engineering, including at least 2 years in each area • Strong production experience with machine learning systems delivering measurable business impact • Deep expertise in Data Science/ML Engineering with solid hands-on competence in the complementary domain • Strong practical experience with gradient-boosted trees such as XGBoost, LightGBM, or CatBoost • Advanced knowledge in at least one of the following areas: delayed labels, PU learning, off-policy evaluation, hierarchical estimation, constrained optimization • Production-level Python and strong SQL skills • Hands-on experience with ML orchestration, CI/CD pipelines, and model registry management • Practical experience with Kubernetes and Docker in production environments • Strong experimentation and evaluation skills, including statistical interpretation of results • Upper-Intermediate or higher English level
WILL BE A PLUS
• Experience in AdTech, RTB, ranking, pricing, or real-time marketplace systems • Knowledge of contextual bandits and off-policy evaluation techniques • Experience with multi-tenant ML systems and data isolation approaches • Background in batch scoring systems with freshness SLA requirements • Hands-on experience with MLflow, Kubeflow, Airflow, or Argo • Experience with GCP services including Vertex AI and BigQuery • Familiarity with Terraform and on-prem Linux infrastructure