Senior Site Reliability Engineer (all genders)
Fact FinderRole brief
IntroductionFACT-Finder develops product discovery technology for eCommerce and is deployed with its Next Generation and Infinity products at leading online shops across Europe. Both products are currently moving toward a modern, hybrid platform based on Kubernetes and Harvester – with the option to scale fully into the cloud in the medium term. As a Senior Site Reliability Engineer (SRE), you ensure that our systems remain fast, available, and scalable throughout this transformation. You work alongside the hosting team and experienced engineers, actively shaping the path toward a modern SaaS company. Your Responsibilities You define and own SLOs, SLIs, and Error Budgets across both products and make data-driven decisions on reliability and performance. You drive incident response forward: fast detection, clear communication, blameless postmortems, and sustainable follow-ups. You consistently reduce manual work through automation and GitOps (e.g., Argo CD / Flux) and expand self-healing and self-service capabilities. You support the development of an NG Search Operator (Custom Kubernetes Operator / CRDs) and the introduction of Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler). You advance our observability – metrics, logs, traces, alerting, and runbooks that truly help on-call. You plan capacity and costs across on-premise (Frankfurt, Stockholm) and cloud – including burst scenarios into the public cloud. You leverage AI tools to significantly improve diagnosis, alerting, and operational workflows. Your ProfileExperience as an SRE, Infrastructure, or Production Engineer in a SaaS or platform environment – or a strong software/operations background with a clear commitment to growing into the SRE role. Solid understanding of SLOs, Error Budgets, Incident Management, and Observability. Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns. Experience or strong interest in Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack). Familiarity with GitOps (Argo CD / Flux), container storage (Longhorn, Ceph), and Kubernetes networking (load balancing, ingress). Knowledge of auto-scaling primitives (HPA, VPA, Cluster Autoscaler, KEDA) and capacity planning on-premise and in the cloud. Understanding of networks in production-grade data centers (including VLANs). Strong automation instinct and a mindset to structurally eliminate toil. Practical experience using AI tools in operational environments. Excellent English skills; German is a plus. THE JOY OF WORKING WITH US Impact from day one: Your work directly affects the revenue of leading eCommerce brands across Europe. Modern tech stack: Kubernetes, Harvester, GitOps, Auto-Scaling, and an exciting path toward the cloud – with room to build things new and right. AI-first mindset: We use AI not as a buzzword, but as an integral part of our daily work. Ownership & Growth: Clear responsibility, short decision-making paths, and the opportunity to actively shape your role. Flexible work: Hybrid work model with a focus on results. Strong team: Experienced engineers, open feedback culture, and an environment where reliability is taken seriously as an engineering discipline. LocationBerlin, Munich, Pforzheim, or Stockholm (Hybrid) [Auto-translated from German]
