Site Reliability Engineer - Cloud Operations
Swissquote · Gland
Description du poste
About the role
Are you passionate about Kubernetes, distributed systems and keeping production platforms reliable at scale? Join our Cloud Operations team and help us migrate and modernize applications running at the core of Swissquote. We’re looking for a Site Reliability Engineer who enjoys working close to production, solving reliability problems and improving how applications are deployed and operated.
Key responsibilities
- Migrate and modernize production applications on Kubernetes.
- Integrate third‑party software into our production platforms and make it fit our operational standards.
- Work alongside Software and IT Engineers to improve reliability, performance and operational readiness.
- Design and operate applications on our service mesh platform.
- Integrate safe deployment patterns such as canary releases and progressive rollouts.
- Define SLOs, SLIs and useful operational KPIs, then use them to drive improvements.
- Improve observability across metrics, logs and traces so problems are easier to spot and understand.
- Explore and integrate AI tools that can help with troubleshooting, incident analysis and remediation.
- Test how systems behave under load, during failures and when dependencies disappear.
- Automate repetitive operational work whenever it makes sense.
- Provide Level‑3 support and participate in the on‑call rotation.
Required profile
- You like understanding why systems behave the way they do, especially when something goes wrong.
- You automate repetitive work instead of accepting it as part of the job.
- You’re comfortable working across development, infrastructure and operations teams.
- You don’t mind getting deep into software you didn’t build yourself.
- You’re curious about AI and where it can genuinely improve day‑to‑day operations.
- You are fluent in English and have good conversational French.
- You enjoy keeping up with cloud‑native technologies and trying new approaches when they solve a real problem.
Required skills
- At least 3 years of experience in SRE, DevOps, Platform Engineering or similar.
- Hands‑on experience running production workloads on Kubernetes (OpenShift, EKS, etc.).
- Good knowledge of Helm.
- Experience with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience.
- Understanding of service‑to‑service networking, traffic routing, mTLS and TLS.
- Experience with GitOps and modern deployment strategies such as canary or progressive delivery.
- Practical understanding of SRE concepts such as SLIs, SLOs and error budgets.
- Experience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry.
- Strong Linux and networking fundamentals, including TCP/IP, DNS and load balancing.
- Comfortable troubleshooting JVM‑based applications, heap usage, garbage collection.
- Comfortable automating with Python, Go, Bash or another language.
- Familiarity with IaC tools such as Terraform, Ansible or Puppet.
Questions fréquentes
Pourquoi signalez-vous cette offre ?
Aller plus loin
Salaires, guides et recherches en Suisse.
Postulez en 30 secondes
Entrez votre email pour postuler. Un compte sera cree automatiquement.
En continuant, vous acceptez nos conditions d'utilisation.
Deja un compte ? Connexion
Une question sur cette offre ?
Posez-la ici : vous recevrez le récapitulatif de l'offre par e-mail, tout de suite.
Publie il y a 6 heures
Expire dans 1 mois
3 vues · 0 interesses
Boostez vos chances
Importez votre CV : nous vous proposons les offres qui matchent votre profil.
Analyse de votre CV en cours...
Swissquote
Gland