Sovereign Cloud Engineer (m/w/d)

TMT Prüfservice GmbH & Co KG

Walldorf · Onsite · Full Time

Posted

Job description

For our growing team, we are looking for an experienced Sovereign Cloud / Site Reliability Engineer (SRE) to support the secure, reliable, and continuous operation of modern cloud-native and Kubernetes platforms. You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement. The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment. Mandatory Requirements – Please Read Before Applying Citizenship You must hold valid citizenship in a country that is a full member of both the European Union (EU) and NATO . If you hold multiple citizenships, all citizenships must be from countries that are members of both the EU and NATO. Employment in Germany You must: Be employed directly by a legal entity registered in Germany. Hold a German employment contract. Be subject exclusively to German labor law. Comply with all applicable German tax, social-security, and employment regulations. Reside in Germany. **Employment through non-German entities, including foreign subcontractors or affiliates, does not meet this requirement. ** Security Clearance – Ü2 You must hold a valid and verifiable Ü2 security clearance in accordance with the German Security Clearance Act ( Sicherheitsüberprüfungsgesetz – SÜG ) and applicable preventive personnel sabotage-protection requirements. Tasks Kubernetes & Platform Operations Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components. Troubleshoot availability, performance and deployment issues. Ensure secure, scalable, resilient and highly available platforms. CI/CD & Automation Operate and optimize Jenkins and ArgoCD pipelines for automated deployments. Implement Infrastructure as Code (IaC) and Git-based deployment workflows. Develop automation and operational tools using Python, Go and/or Bash. Automate provisioning, health and compliance checks, alerting and reporting. Monitoring & Observability Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL. Develop and maintain Grafana dashboards. Continuously i…

Apply for this job