Talent.com
swissquote
Site Reliability Engineer - Cloud Operationsswissquote • gland, waadt, Switzerland
Cerca altre offerte di lavoro
Site Reliability Engineer - Cloud Operations

Site Reliability Engineer - Cloud Operations

swissquote • gland, waadt, Switzerland
22 giorni fa
Descrizione dell’offerta di lavoro

ph3Company Description /h3 /brpAt Swissquote, we’re all in. All in to shake things up. All in to build the bank people actually want to use. All in to make finance less boring — and a lot more powerful. We’re Switzerland’s leading digital bank — 1,400+ people across Europe, the Middle East and Asia, building real financial solutions for over a million clients worldwide. From trading and investing to everyday banking, we cover the full picture. We move fast, but we build things we’re genuinely proud of. Have a look behind the scenes by checking Humans of Swissquote on Instagram. Growing fast creates room. Room to try things, own things, and grow at a pace most places can’t offer. Whether you like the spotlight or prefer to just put your head down and do great work, there’s space for both here. We’ve ditched the dress code — but never the chance to celebrate. Big win or small, we make it count. The kind of place where the atmosphere takes care of itself. As an equal opportunity employer, we welcome candidates from all backgrounds, experiences and perspectives to join our team and contribute to our shared success. Feeling it? It’s a good start. /p /br /brh3Team Mission and Stakeholders /h3 /brpAre you passionate about Kubernetes, distributed systems and keeping production platforms reliable at scale? Join our Cloud Operations team and help us migrate and modernize applications running at the core of Swissquote. We’re looking for a Site Reliability Engineer who enjoys working close to production, solving reliability problems and improving how applications are deployed and operated. /p /br /brh3Job Description /h3 /brul /brliMigrate and modernize production applications on Kubernetes, /li /brliIntegrate third-party software into our production platforms and make it fit our operational standards, /li /brliWork alongside Software and IT Engineers to improve reliability, performance and operational readiness, /li /brliDesign and operate applications on our service mesh platform, /li /brliIntegrate safe deployment patterns such as canary releases and progressive rollouts, /li /brliDefine SLOs, SLIs and useful operational KPIs, then use them to drive improvements, /li /brliImprove observability across metrics, logs and traces so problems are easier to spot and understand, /li /brliExplore and integrate AI tools that can help with troubleshooting, incident analysis and remediation, /li /brliTest how systems behave under load, during failures and when dependencies disappear, /li /brliAutomate repetitive operational work whenever it makes sense, /li /brliProvide Level-3 support and participate in the on-call rotation. /li /br /ul /br /brh3Qualifications /h3 /brul /brliAt least 3 years of experience in SRE, DevOps, Platform Engineering or a similar production-focused role, /li /brliSolid hands-on experience running production workloads on Kubernetes, OpenShift, EKS or a similar Kubernetes platform, /li /brliGood knowledge of Helm and how to package, configure and maintain applications with it, /li /brliExperience working with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience, /li /brliA good understanding of service-to-service networking, traffic routing, mTLS and TLS, /li /brliExperience with GitOps and modern deployment strategies such as canary or progressive delivery, /li /brliA practical understanding of SRE concepts such as SLIs, SLOs and error budgets, /li /brliExperience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry, /li /brliStrong Linux and networking fundamentals, including TCP/IP, DNS and load balancing, /li /brliComfortable troubleshooting JVM-based applications in production and able to investigate issues related to heap usage, garbage collection or JVM configuration, /li /brliComfortable automating things with Python, Go, Bash or another programming language, /li /brliExperience or strong interest in applying AI to observability, incident response or operational automation, /li /brliExperience or a strong interest in operating applications that depend on GPU resources or other AI infrastructure, /li /brliFamiliarity with Infrastructure as Code tools such as Terraform, Ansible or Puppet. /li /br /ul /br /brh3Nice-to-Haves /h3 /brul /brliExperience with Argo CD, Argo Rollouts or Argo Workflows, /li /brliDeeper experience with Istio, Linkerd or Envoy-based service mesh platforms, /li /brliExperience designing or operating Kubernetes platforms at scale, /li /brliExperience running Java or Spring Boot applications in production, /li /brliHands-on experience tuning JVM applications for performance or low-latency workloads, /li /brliExperience integrating applications with self-hosted AI platforms such as vLLM, /li /brliExperience troubleshooting AI infrastructure integrations, including model access, GPU availability and NVIDIA MIG configurations. /li /brliKnowledge of Cilium, eBPF or other modern Kubernetes networking technologies, /li /brliExperience with public cloud or large private cloud environments, /li /brliCKAD, CKA, CKS or equivalent hands‑on Kubernetes experience, /li /brliA homelab, self-hosted services or side projects where you get to experiment, break things and build them again. /li /br /ul /br /brh3Who You Are /h3 /brul /brliYou like understanding why systems behave the way they do, especially when something goes wrong, /li /brliYou automate repetitive work instead of accepting it as part of the job, /li /brliYou’re comfortable working across development, infrastructure and operations teams, /li /brliYou don’t mind getting deep into software you didn’t build yourself, /li /brliYou’re curious about AI and where it can genuinely improve day‑to‑day operations, /li /brliYou are fluent in English and have good conversational French, /li /brliYou enjoy keeping up with cloud-native technologies and trying new approaches when they solve a real problem. /li /br /ul /br /brh3Additional Information /h3 /brpPlease note that Swissquote never requests sensitive personal information or payment of any kind during the recruitment process. Any such request is fraudulent. /p /brpSQ2 /p /p #J-18808-Ljbffr

Creare un avviso di lavoro per questa ricerca

Site Reliability Engineer - Cloud Operations • gland, waadt, Switzerland