Role brief
YOUR TASKS You work in the IT landscapes of our customers with a clear focus on middleware, integration platforms, and private cloud environments. In the spirit of a Site Reliability Engineer, you take responsibility for stable, scalable, and high-performing systems. You support, operate, and develop middleware components such as application servers, integration and interface solutions, as well as their integration into private cloud infrastructures and connection to backend and frontend systems. You ensure highly available and resilient operations by actively applying SRE principles such as automation, monitoring, incident and problem management, and continuous improvement. You assume technical responsibility with the customer, coordinate tasks, and work closely with internal project, cloud, and operations teams. You support transformation and migration projects, particularly in the context of cloud strategies, private cloud architectures, hybrid scenarios, and system modernization. You analyze and optimize systems with regard to availability, performance, and scalability, and establish appropriate observability and monitoring solutions. You document architectures, interfaces, systems, and configurations in a comprehensible and customer-oriented manner. You define and collect relevant metrics (e.g., SLIs/SLOs) and prepare these as a basis for decision-making for customers and stakeholders. You are the point of contact for our customers in operational management; this includes, if necessary, participation in an on-call rotation organized internally by the team. YOUR PROFILE You have completed training or a degree in IT or comparable qualifications and bring several years of experience in operating complex IT systems – ideally in a private cloud environment or as a Site Reliability Engineer. You have sound knowledge in the area of middleware (e.g., application servers, integration and interface platforms) as well as a good understanding of modern cloud and private cloud architectures. You bring experience with SRE methods and tools, particularly in the areas of monitoring, logging, alerting, automation, and incident management. You understand the interaction of applications, operating systems, networks, virtualization, and security components in complex, distributed system landscapes. You analyze problems in a structured manner, develop sustainable solutions, and communicate confidently with technical and business contacts. You configure and optimize systems independently, develop individual solutions if needed, and use automation and scripts to make operations efficient and reliable. You have a pronounced understanding of metrics in IT operations (e.g., availability, latency, error rates) and can prepare and present these appropriately for your audience. Project work is second nature to you; you structure tasks, maintain an overview, and actively promote teamwork. Become part of our team! We look forward to your application. [Auto-translated from German]
