Site Reliability Engineer - Cloud Operations
Swissquote · Gland
Job description
About the role
Are you passionate about Kubernetes, distributed systems and keeping production platforms reliable at scale? Join our Cloud Operations team and help us migrate and modernize applications running at the core of Swissquote. We’re looking for a Site Reliability Engineer who enjoys working close to production, solving reliability problems and improving how applications are deployed and operated.
Key responsibilities
- Migrate and modernize production applications on Kubernetes.
- Integrate third‑party software into our production platforms and make it fit our operational standards.
- Work alongside Software and IT Engineers to improve reliability, performance and operational readiness.
- Design and operate applications on our service mesh platform.
- Integrate safe deployment patterns such as canary releases and progressive rollouts.
- Define SLOs, SLIs and useful operational KPIs, then use them to drive improvements.
- Improve observability across metrics, logs and traces so problems are easier to spot and understand.
- Explore and integrate AI tools that can help with troubleshooting, incident analysis and remediation.
- Test how systems behave under load, during failures and when dependencies disappear.
- Automate repetitive operational work whenever it makes sense.
- Provide Level‑3 support and participate in the on‑call rotation.
Required profile
- You like understanding why systems behave the way they do, especially when something goes wrong.
- You automate repetitive work instead of accepting it as part of the job.
- You’re comfortable working across development, infrastructure and operations teams.
- You don’t mind getting deep into software you didn’t build yourself.
- You’re curious about AI and where it can genuinely improve day‑to‑day operations.
- You are fluent in English and have good conversational French.
- You enjoy keeping up with cloud‑native technologies and trying new approaches when they solve a real problem.
Required skills
- At least 3 years of experience in SRE, DevOps, Platform Engineering or similar.
- Hands‑on experience running production workloads on Kubernetes (OpenShift, EKS, etc.).
- Good knowledge of Helm.
- Experience with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience.
- Understanding of service‑to‑service networking, traffic routing, mTLS and TLS.
- Experience with GitOps and modern deployment strategies such as canary or progressive delivery.
- Practical understanding of SRE concepts such as SLIs, SLOs and error budgets.
- Experience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry.
- Strong Linux and networking fundamentals, including TCP/IP, DNS and load balancing.
- Comfortable troubleshooting JVM‑based applications, heap usage, garbage collection.
- Comfortable automating with Python, Go, Bash or another language.
- Familiarity with IaC tools such as Terraform, Ansible or Puppet.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Switzerland.
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 13 hours ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Swissquote
Gland