Platform & SRE Engineer
A leading IT company is looking for a Platform & SRE Engineer to join a team responsible for a large-scale Event-Driven streaming platform. This platform ensures real-time data synchronization between legacy Mainframe environments and a cloud-native ecosystem. The mission is hands-on: you will be responsible for the operation and stabilization of the OpenShift platform, the evolution of the GitOps foundation, troubleshooting data pipelines, and supporting development teams.
The objective is to solve concrete infrastructure problems and improve the reliability and efficiency of the platform.
What you do
- Operate and stabilize the multi-cluster OpenShift infrastructure and shared services.
- Administer and secure OpenShift clusters: workloads, CRDs, secrets, NetworkPolicies, ServiceAccounts, SCCs, RBAC, quotas, and LimitRanges.
- Manage storage integration (PVC/NFS, S3) and enterprise PKI.
- Maintain and evolve the GitOps foundation based on ArgoCD and Helm, especially in a multi-environment context.
- Develop and debug complex Helm charts.
- Support CI/CD pipelines, including Jenkins, SonarQube, and Flyway migrations.
- Operate Kafka / IBM EventStreams clusters: broker and topic configuration, partitions, retention, rolling restarts, consumer groups, and lag monitoring.
- Provide operational support for Apache Flink, including savepoints/checkpoints, memory sizing, backpressure, and job restarts.
- Diagnose and resolve incidents related to CDC pipelines, Flink jobs, and Kafka brokers.
- Configure and maintain OpenTelemetry Collector pipelines.
- Create and maintain Grafana dashboards, alert rules, and operational runbooks.
- Administer Grafana Loki and utilize PromQL / LogQL.
- Support application teams with their network, OAuth2/OIDC, build, integration, and deployment issues.
- Participate in IAM management with Entra ID / Active Directory, app registrations, MSAL, and Azure Key Vault CSI.
Profile
Infrastructure & Containerization.
- Strong command of Linux.
- Solid experience with Kubernetes / Red Hat OpenShift.
- Concrete experience in administering and troubleshooting containerized environments.
CI/CD & GitOps.
- Strong command of Git.
- Experience with ArgoCD or strong willingness to master it.
- Good knowledge of Helm.
- Concrete experience with CI/CD pipelines: Jenkins, GitHub Actions, or equivalent.
- Experience with ArgoCD ApplicationSets, App-of-Apps, Helm chart authoring, or Kustomize is a plus.
Streaming & Data Pipelines.
- Experience with Kafka / EventStreams / Strimzi.
- Knowledge of Kafka Connect and Schema Registry (Apicurio/Avro).
- Good understanding of CDC concepts.
- Experience with Apache Flink, including Flink Kubernetes Operator, savepoints, and state backends, is a plus.
Observability.
- Experience with Prometheus and Grafana.
- Knowledge of centralized logging systems such as Loki, ELK, or equivalent.
- Experience with OpenTelemetry Collector and the LGTM stack (Loki, Grafana, Tempo, Mimir) is a plus.
- Knowledge of PromQL / LogQL appreciated.
Development.
- Good understanding of the Java / Spring Boot ecosystem, mainly for diagnosing containerized applications.
- Pure developer experience is not required.
Databases.
- Knowledge of PostgreSQL.
- Familiarity with Elasticsearch / Kibana.
Other appreciated skills.
- Knowledge of Angular.
- Familiarity with Mainframe / DB2 / IBM MQ environments.
- Knowledge of Infrastructure as Code with Ansible / Terraform.
- Automation via Bash / Python.
- Use and strong interest in artificial intelligence tools in daily work: code assistants, script and runbook generation, log analysis, and automation.
Practical
Languages.
- English required.
- French or Dutch.
How to apply
View the full assignment text and application details once your tailored application is ready.
Order a tailored application to view the full assignment and application details.
More context, less searching.
You get enough context to judge whether this job is relevant. The full brief, client details and next steps stay available inside the app.