StellenbeschreibungpThe bQuant DevOps Engineer /b owns and operates the application engineering platform, cloud infrastructure, and data execution environment that application and modeling workloads depend on. /ppThe role is accountable for the end‑to‑end reliability, scalability, security, performance, and cost efficiency of the platform, supporting large‑scale, data‑intensive and compute‑heavy workloads such as parallel Databricks clusters and high‑memory systems. /ppThis position requires broad and deep technical expertise, strong operational judgment, and the ability to act as a central technical interface between engineering teams, IT operations, security/SecOps, and data users. It is a senior ownership role with significant impact on delivery speed, platform stability, and infrastructure cost. /pbr/pbPlatform Infrastructure Ownership /b /pliEnd-to-end ownership and delivery accountability for the application execution platform and production model runs (Databricks, Servers). /liliDay-to-day operational responsibility: availability, incident handling, runtime management, and delivery continuity for business-critical runs (incl. peak periods like YE). /liliManage Azure cloud infrastructure, including: Virtual machines, storage, identity and access management (RBAC)Networking components such as firewalls, peering, and cross‑subscription connectivity /liliLead standardization with central teams across observability, security controls, platform services, and “golden paths,” while keeping delivery running /liliEnsure platform reliability, scalability, and long‑term sustainability /lipbCI/CD Automation /b /pliDesign, maintain, and continuously improve CI/CD pipelines using Azure DevOps and related tooling /liliBuild and evolve automation using scripting and build tools /liliOptimize pipeline performance, reliability, and parallel execution to support large‑scale workloads /lipbContainer Runtime Management /b /pliOwn Docker image creation, lifecycle management, and governance /liliOptimize build processes, caching strategies, and container security /liliSupport containerized execution environments for compute‑heavy workloads and services /lipbInfrastructure as Code Configuration Management /b /pliMaintain and evolve Infrastructure as Code using Terraform /liliOperate and improve configuration management systems (e.g. SaltStack) /liliReduce configuration drift and improve reproducibility across environments /lipbObservability, Security Secrets /b /pliOwn and operate observability platforms (e.g. ELK, Prometheus, Grafana) /liliEnsure meaningful metrics, logs, dashboards, and alerting are in place /liliManage secrets and platform security tooling (e.g. Wiz, Snyk) /liliCollaborate closely with Security and SecOps teams on controls, findings, and improvements /lipbData Platform Capacity Planning /b /pliConfigure/support Databricks usage for application workloads; manage workspace-level configuration/permissions as delegated; partner with central Databricks lead for global administration and optimization. /liliSupport large‑scale data and compute workloads /liliLead capacity planning for: Highly parallel Databricks clusters (e.g. up to ~10 × 80‑node clusters)Memory‑intensive systems (multi‑terabyte RAM)Data pipelines producing terabytes of data /liliBalance performance, reliability, and cost across platform decisions /lipbOperations Incident Response /b /pliAct as senior escalation point for platform and infrastructure incidents /liliParticipate in a limited on‑call rotation /liliInvestigate incidents and execute or coordinate remediation /liliPerform manual interventions when automation is insufficient /liliDrive post‑incident reviews and platform improvements /lipbCross‑Team Organizational Coordination /b /pliServe as primary technical contact for: IT OperationsSecurity and SecOpsArchitecture and governance bodiesService management processes /liliCoordinate platform‑related work across teams /liliSupport customer‑facing technical discussions related to platform capabilities and constraints /libr/pbExperience /b /pliBackground in bInsurance, Finance, or Scientific / High‑Performance Computing environments /b /liliStrong experience in bplatform or DevOps engineering /b within production environments /liliSolid expertise in bcloud infrastructure /b, preferably Microsoft Azure /liliHands‑on experience with bCI/CD, Infrastructure as Code, and container platforms /b /liliProven experience operating bdata‑intensive and compute‑heavy systems /b /lipbTechnical Competencies /b /pliStrong troubleshooting and operational mindset /liliAbility to manage and balance competing constraints: CostPerformanceSecurityReliability /liliDeep understanding of platform stability, scalability, and automation /lipbProfessional Competencies /b /pliSenior‑level autonomy and decision‑making capability /liliOwnership mindset; accountable for outcomes rather than tasks /liliAbility to operate as a trusted senior technical interface across teams /li