About the role
Job Description
* Design and build scalable cloud-native data platforms from greenfield to production * Design and implement near-real-time ingestion pipelines using event-driven patterns * Define and enforce data platform standards including Data Lake and Lakehouse principles, medallion architecture, and data contracts * Refactor, optimize, and modernize Spark and PySpark scripts for performance and maintainability * Introduce best practices for code quality, testing, and CI/CD across data pipelines * Drive adoption of AI tooling and agentic workflows within the Data Engineering team * Ensure data quality, observability, scalability, and reliability across all platforms and pipelines * Design and deliver self-service tooling and microservices that simplify platform usage * Collaborate with cross-functional stakeholders including Product, Machine Learning, and Data Science teams * Contribute to architecture decisions and technical R&D initiatives
Qualifications
* 5+ years of professional experience in Data Engineering * Strong Python and SQL development skills * Hands-on experience with Apache Spark and PySpark including query optimization and performance tuning * Experience working with Databricks or Snowflake * Practical experience with at least one major cloud provider such as Azure, AWS, or GCP * Experience with stream processing technologies including Kafka or Spark Structured Streaming * Strong understanding of ETL/ELT patterns, data modelling, and data warehousing concepts * Experience with orchestration tools such as Apache Airflow or Azure Data Factory * Knowledge of Infrastructure as Code tools including Terraform * Understanding of production-grade systems including observability, scalability, reliability, and performance * Ability to independently lead technical initiatives from concept to delivery * Strong communication and collaboration skills * Upper-Intermediate or higher English level
WILL BE A PLUS
* Familiarity with RAG pipeline design and LLM integration patterns * Knowledge of data governance frameworks and tools such as Unity Catalog or Apache Atlas * Experience with dbt for data transformation and modelling * Familiarity with MLflow, Feature Stores, or ML platform integrations
Additional Information
PERSONAL PROFILE
* Proactive and self-driven mindset * Strong analytical and architectural thinking * Ability to work independently and take ownership of technical decisions * Passion for innovation and modern engineering practices * Knowledge-sharing and team-oriented approach * Strong problem-solving skills