Site Reliability Engineer
A technology company is rapidly rolling out new AI-driven features and agents. To make this reliable, scalable, and measurable, we are looking for a Site Reliability Engineer. You will be the architect behind operational visibility and the driving force behind our metrics. You will provide the fundamental data layer that proves our systems perform at the highest quality level. You will translate complex telemetry data into clear insights for all stakeholders.
What you will do.
- Design and implement the core monitoring structure for the SaaS application. Define clear Service Level Indicators (SLIs) and Service Level Objectives (SLOs) in Datadog for critical endpoints. Identify and eliminate performance bottlenecks using error budget alerting and real-user monitoring (RUM).
- Build the operational visibility layer. Aggregate existing OTel signals in Datadog and link this to PR-per-developer tracking, AI costs, and consumption figures to provide direct insight into adoption and throughput before and after the widespread rollout of AI tooling.
- Take observability to the next level by deeply focusing on the AI agents themselves. Make skill invocations, success rates, drift, and output quality measurable to ensure everything runs at the appropriate quality level.
Technologies.
- Observability & Reliability: Datadog (APM & RUM), OpenTelemetry (OTel), StatusPage, DORA metrics frameworks.
- Feature Management & Deployment: LaunchDarkly, GitHub Actions, Docker, Terraform.
- AWS Services: SNS, SQS, S3, Lambda, API Gateway.
- Core Tech & Collaboration: MySQL, OpenSearch, PHP/Laravel, Vue.js, Linear, Confluence, Slack.
Profile
- An organized, methodical, and flexible team player with a passion for data-driven platform reliability and application performance.
- Strong communication skills to coach developers and clearly visualize performance data for the entire team.
- Extensive experience with modern SRE and observability standards, specifically with Datadog APM, synthetics, and OpenTelemetry.
- Experience with defining and monitoring SLIs, SLOs, and setting up proactive alerting based on error budgets.
- Knowledge of progressive delivery tools such as LaunchDarkly and incident communication via StatusPage.
- Analytically minded, proactive, and always looking for ways to optimize application latency and infrastructure costs.
- Experience in software engineering or DevOps/SRE, ideally within a scale-up or SaaS environment.
Practical
- An exciting role in a growing, international SaaS scale-up.
- High autonomy, flat hierarchies, and short decision-making lines.
- Flexible work options.
- Team events and opportunities for personal development.
- Choice of your own hardware (Windows, Mac, Linux).
How to apply
View the full assignment text and application details once your tailored application is ready.
Order a tailored application to view the full assignment and application details.
More context, less searching.
You get enough context to judge whether this job is relevant. The full brief, client details and next steps stay available inside the app.