About the role
Job Description
We are looking for a hands-on Site Reliability Engineer (SRE) to help ensure the reliability, availability, scalability, and performance of our production eFX/Crypto platforms.
The role has a strong focus on Kubernetes, production operations, automation, observability, and infrastructure reliability. You will work closely with Infrastructure and Development engineers to operate and continuously improve our production environment across on-premise and private cloud infrastructure.
Key Responsibilities
* Monitor production systems, respond to incidents, and ensure optimal uptime and performance of applications and infrastructure. * Operate, maintain, and troubleshoot Kubernetes clusters and containerized applications in production, and on-premise infrastructure. * Contribute to the design and implementation of scalable and reliable infrastructure solutions. * Automate IT and operational processes to reduce manual work and improve efficiency. * Implement and maintain CI/CD pipelines and deployment automation. * Monitor and improve system observability using tools such as Prometheus, Grafana, and Elastic/Kibana. * Analyse faults, perform root-cause analysis, and implement corrective and preventive actions. * Perform stress, resilience, disaster recovery, and BCP testing to validate platform stability. * Work closely with Infrastructure and Development teams to improve production readiness and system reliability. * Participate in the on-call rotation and provide operational support for production systems. * Continuously improve operational processes, automation, monitoring, and reliability practices.
Qualifications
* 3+ years of experience in SRE, DevOps, Systems Engineering, Platform Engineering, or a similar production engineering role. * Strong hands-on Kubernetes experience in production environments — mandatory. * Solid understanding of Kubernetes architecture and day-to-day operations, including: * Services mesh and Istio * Users management * Health checks and probes * Kubernetes networking and DNS * Troubleshooting and performance issues * Linux administration and troubleshooting skills. * Experience with Docker and container technologies. * Hands-on experience with on-premise infrastructure and/or private cloud environments. * Experience with Infrastructure as Code and automation tools such as Terraform and Ansible. * Experience with CI/CD pipelines. * Experience with observability and monitoring tools such as Prometheus, Grafana, Elastic/Kibana. * Good programming and scripting skills in Python and Shell. * Strong understanding of networking fundamentals, including TCP/IP, DNS, HTTPS, HTTP, and load balancing. * Good understanding of SRE principles, including SLIs, SLOs, SLAs, availability, reliability, and service performance. * Good understanding of Java web applications and web servers such as Spring Boot and Apache Tomcat. * Fluent English, both written and spoken, with willingness to learn French.
Nice to Have
* CKA or CKAD certification. * Experience with Helm, ArgoCD, Github Actions , or Kubernetes operators. * Knowledge of Apache Kafka or RabbitMQ. * Knowledge of MT4/MT5. * Experience with Windows Server 2016/2022. * Experience in eFX, trading, cryptocurrency, banking, or other financial services. * Experience working with high-availability or low-latency production systems.
Who You Are:
* A team player who loves helping colleagues and bridging gaps between departments, * Detail-oriented and meticulous in your work, * Fluency in English (spoken and written) and willing to learn french. * Passionate about technology and eager to stay up-to-date with AI innovations.
Additional Information
Please note that Swissquote never requests sensitive personal information or payment of any kind during the recruitment process. Any such request is fraudulent.
SQ2