Site Reliability Engineer - Cloud Operations

Swissquote

Gland, VD, Switzerland Full-time 1 mo ago
Workplace
On-site
Location
Gland, VD, Switzerland
Who can apply
Based in Switzerland: you usually need the right to work there

Check your CV against this job

Free · No signup · A 0–100 ATS match score and the keywords you're missing.

Check my CV free

$14.99/month, cancel anytime. Already have an account? Log in

About the role

Job Description

In this role, you will:

* Migrate and modernize production applications on Kubernetes, * Integrate third-party software into our production platforms and make it fit our operational standards, * Work alongside Software and IT Engineers to improve reliability, performance and operational readiness, * Design and operate applications on our service mesh platform, * Integrate safe deployment patterns such as canary releases and progressive rollouts, * Define SLOs, SLIs and useful operational KPIs, then use them to drive improvements, * Improve observability across metrics, logs and traces so problems are easier to spot and understand, * Explore and integrate AI tools that can help with troubleshooting, incident analysis and remediation, * Test how systems behave under load, during failures and when dependencies disappear, * Automate repetitive operational work whenever it makes sense, * Provide Level-3 support and participate in the on-call rotation.

Qualifications

* At least 3 years of experience in SRE, DevOps, Platform Engineering or a similar production-focused role, * Solid hands-on experience running production workloads on Kubernetes, OpenShift, EKS or a similar Kubernetes platform, * Good knowledge of Helm and how to package, configure and maintain applications with it, * Experience working with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience, * A good understanding of service-to-service networking, traffic routing, mTLS and TLS, * Experience with GitOps and modern deployment strategies such as canary or progressive delivery, * A practical understanding of SRE concepts such as SLIs, SLOs and error budgets, * Experience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry, * Strong Linux and networking fundamentals, including TCP/IP, DNS and load balancing, * Comfortable troubleshooting JVM-based applications in production and able to investigate issues related to heap usage, garbage collection or JVM configuration, * Comfortable automating things with Python, Go, Bash or another programming language, * Experience or strong interest in applying AI to observability, incident response or operational automation, * Experience or a strong interest in operating applications that depend on GPU resources or other AI infrastructure, * Familiarity with Infrastructure as Code tools such as Terraform, Ansible or Puppet.

Nice-to-Haves

* Experience with Argo CD, Argo Rollouts or Argo Workflows, * Deeper experience with Istio, Linkerd or Envoy-based service mesh platforms, * Experience designing or operating Kubernetes platforms at scale, * Experience running Java or Spring Boot applications in production, * Hands-on experience tuning JVM applications for performance or low-latency workloads, * Experience integrating applications with self-hosted AI platforms such as vLLM, * Experience troubleshooting AI infrastructure integrations, including model access, GPU availability and NVIDIA MIG configurations. * Knowledge of Cilium, eBPF or other modern Kubernetes networking technologies, * Experience with public cloud or large private cloud environments, * CKAD, CKA, CKS or equivalent hands-on Kubernetes experience, * A homelab, self-hosted services or side projects where you get to experiment, break things and build them again.

Who You Are

* You like understanding why systems behave the way they do, especially when something goes wrong, * You automate repetitive work instead of accepting it as part of the job, * You’re comfortable working across development, infrastructure and operations teams, * You don’t mind getting deep into software you didn’t build yourself, * You’re curious about AI and where it can genuinely improve day-to-day operations, * You are fluent in English and have good conversational French, * You enjoy keeping up with cloud-native technologies and trying new approaches when they solve a real problem.

Additional Information

Please note that Swissquote never requests sensitive personal information or payment of any kind during the recruitment process. Any such request is fraudulent.

SQ2

Get matched & apply with FindAJobAI

Upload your resume once. We score every job against your profile, tailor your resume and cover letter, and autofill the application.