About the role
Summary: We are seeking an experienced Senior DevOps Engineer with over 8 years in the field to lead the design, implementation, and management of infrastructure and deployment pipelines across hybrid environments, including on-premises data centers and the cloud. This role requires deep expertise in AWS, Azure, Kubernetes, Infrastructure as Code (IaC), CI/CD, and monitoring tools. The engineer will collaborate closely with development, QA, and operations teams to ensure the delivery of reliable, scalable, and secure enterprise-grade applications. Additionally, the role involves mentoring junior team members and leading automation initiatives.
Responsibilities:
* Manage and support hybrid infrastructure, including physical data centers, AWS, and Azure services. * Design and implement Infrastructure-as-Code using tools like Terraform, Pulumi, and CloudFormation. * Configure and maintain Kubernetes clusters, ensuring scalability and high availability. * Build, maintain, and optimize CI/CD pipelines using Jenkins, GitLab CI, Bitbucket, and GitHub. * Automate deployment, scaling, and monitoring processes. * Set up and maintain observability using Grafana, Prometheus, and other monitoring tools. * Perform root cause analysis of performance issues and participate in on-call support. * Manage SSL/TLS and code signing certificates and implement infrastructure security best practices. * Support compliance and audit requirements through documentation and monitoring. * Collaborate with cross-functional teams to improve infrastructure reliability and delivery processes. * Document processes in Confluence, track work in Jira, and contribute to knowledge sharing. * Write root cause analysis and incident post-mortems to prevent future issues.
Required Skills & Experience:
* 8+ years of experience in DevOps, Site Reliability Engineering, or related roles. * Strong expertise with AWS and Azure services. * Experience with Infrastructure-as-Code tools and GitOps. * Proven experience with Kubernetes/EKS setup and operations. * Proficiency in CI/CD pipelines and GitOps tools. * Strong knowledge of monitoring and observability tools and incident management. * Scripting and automation skills in Python, TypeScript, and Bash/shell. * Understanding of networking, Linux/Windows Server administration, and troubleshooting. * Experience with data streaming, CDC pipelines, and administering relational and in-memory data stores. * Configuration management and server automation experience. * Knowledge of security best practices and familiarity with compliance/audit frameworks.
Soft Skills:
* Strong problem-solving and root cause analysis capabilities. * Excellent communication and documentation skills. * Ability to collaborate in team meetings and contribute to on-call rotations. * Mentoring and knowledge-sharing mindset.