Senior DevOps Engineer / SRE (m/f/d)
Voize
Berlin · Onsite · Full Time
Posted
Job description
🎤 Why voize? Because we’re more than just a job! At voize, we believe the greatest gift to frontline workers is time - time to care, connect, and be present. Today, that time is lost to busywork and complex systems that pull them away from what matters most: people . Our vision is to change that by building AI companions that seamlessly take over digital workflows. We don't replace humans with technology - we amplify their impact. Our mission is backed with a $50M Series A funding led by Balderton Capital, with support from HV Capital, Y Combinator and other leading VCs. Today, 2,000+ facilities trust voize, and over 200,000 users rely on our AI companion to ease their daily workload. As a dynamic team, we combine first-in-class technology with meaningful social impact. And now, we’re looking for you to join us on this mission! 💡 Your Mission: Delight customers, optimize processes! As DevOps Engineer , you will build and own voize's platform and infrastructure — from cloud and edge systems to ML data pipelines; ensuring our systems are secure, reliable, and compliant. You will enable product and ML teams to ship state-of-the-art healthcare AI with confidence to hospitals, care facilities, and mobile devices across Europe. 🚀 Your Daily Business - No two days are alike Own and operate our Kubernetes clusters (AWS EKS production, bare-metal K3s for ML training, on-premises appliances running K3s) using GitOps (FluxCD) and Infrastructure as Code (CloudFormation, Ansible) Manage the lifecycle of on-premises gateway appliances deployed at customer sites — VM image builds, TLS certificate automation, staged rollouts via GitOps, remote monitoring and troubleshooting Build and maintain monitoring, alerting, and observability infrastructure (Prometheus, Grafana, Loki, Tempo, OpenTelemetry) across cloud and edge environments Drive compliance automation — security hardening, access controls, audit logging, encrypted secrets management (SOPS/KMS), and evidence collection for C5, HIPAA, and HDS certifications Support and scale ML training and data processing infrastructure — GPU cluster management, training job orchestration, and data pipeline reliability, working closely with the ML team to power state-of-the-art speech and language models Incident response — detection, triage, resolution, and post-incident reviews for infrastructure issues. 🤝 Your Skillset - What y…