Migrate and modernize production applications on Kubernetes
Integrate third-party software into our production platforms and make it fit our operational standards
Work alongside Software and IT Engineers to improve reliability performance and operational readiness
Design and operate applications on our service mesh platform
Integrate safe deployment patterns such as canary releases and progressive rollouts
Define SLOs SLIs and useful operational KPIs then use them to drive improvements
Improve observability across metrics logs and traces so problems are easier to spot and understand
Explore and integrate AI tools that can help with troubleshooting incident analysis and remediation
Test how systems behave under load during failures and when dependencies disappear
Automate repetitive operational work whenever it makes sense
Provide Level-3 support and participate in the on-call rotation.
Qualifications :
At least 3 years of experience in SRE DevOps Platform Engineering or a similar production-focused role
Solid hands-on experience running production workloads on Kubernetes OpenShift EKS or a similar Kubernetes platform
Good knowledge of Helm and how to package configure and maintain applications with it
Experience working with service mesh technologies such as Istio or Linkerd or strong Kubernetes networking experience
A good understanding of service-to-service networking traffic routing mTLS and TLS
Experience with GitOps and modern deployment strategies such as canary or progressive delivery
A practical understanding of SRE concepts such as SLIs SLOs and error budgets
Experience with observability and tracing tooling such as Prometheus Grafana Elastic Stack or OpenTelemetry
Strong Linux and networking fundamentals including TCP/IP DNS and load balancing
Comfortable troubleshooting JVM-based applications in production and able to investigate issues related to heap usage garbage collection or JVM configuration
Comfortable automating things with Python Go Bash or another programming language
Experience or strong interest in applying AI to observability incident response or operational automation
Experience or a strong interest in operating applications that depend on GPU resources or other AI infrastructure
Familiarity with Infrastructure as Code tools such as Terraform Ansible or Puppet.
Nice-to-Haves
Experience with Argo CD Argo Rollouts or Argo Workflows
Deeper experience with Istio Linkerd or Envoy-based service mesh platforms
Experience designing or operating Kubernetes platforms at scale
Experience running Java or Spring Boot applications in production
Hands-on experience tuning JVM applications for performance or low-latency workloads
Experience integrating applications with self-hosted AI platforms such as vLLM
Experience troubleshooting AI infrastructure integrations including model access GPU availability and NVIDIA MIG configurations.
Knowledge of Cilium eBPF or other modern Kubernetes networking technologies
Experience with public cloud or large private cloud environments
CKAD CKA CKS or equivalent hands-on Kubernetes experience
A homelab self-hosted services or side projects where you get to experiment break things and build them again.
Who You Are
You like understanding why systems behave the way they do especially when something goes wrong
You automate repetitive work instead of accepting it as part of the job
Youre comfortable working across development infrastructure and operations teams
You dont mind getting deep into software you didnt build yourself
Youre curious about AI and where it can genuinely improve day-to-day operations
You are fluent in English and have good conversational French
You enjoy keeping up with cloud-native technologies and trying new approaches when they solve a real problem.
Additional Information :
Please note that Swissquote never requests sensitive personal information or payment of any kind during the recruitment process. Any such request is fraudulent.
SQ2
Remote Work :
No
Employment Type :
Full-time
Experience: years Vacancy: 1
Creare un avviso di lavoro per questa ricerca
Site Reliability Engineer Cloud Operations • Nyon, Vaud, Switzerland