About the role
Job Description
Your responsibilities
* Own the operational stability of the Integration and Delivery Tools platform across all squads, and act as the escalation point for platform-level incidents. * Design, document and run the oncall and live-support model for the integration side, and take part in the oncall rotation yourself. * Lead major incident calls and post-incident reviews; drive RCA, problem management and the closure of systemic remediation. * Audit and improve monitoring and observability (Datadog), with clear ownership, meaningful alerts and reduced alert noise. * Define and track SLAs, SLOs and platform KPIs, and use dashboards and trends to anticipate risks and steer improvements. * Plan and run disaster recovery and resilience exercises, and close the gaps they reveal. * Drive performance and resilience testing together with the squads, and apply quality gates for critical journeys and releases. * Plan cloud capacity and seasonal readiness, and contribute to FinOps and cost optimisation of the platform. * Coordinate squads, vendors and business stakeholders during escalations and outages, and keep incident hygiene and ticket quality high. * Contribute to the platform operations strategy together with the Platform Lead and architects, and coach squads on an operations assurance mindset.
Qualifications
Required key competencies and qualifications
* At least 5 years of hands-on experience in operations, SRE or platform engineering, running production systems of large enterprise applications. * At least 3 years of experience leading or coordinating teams, incident response or operational workstreams (line management is not required). * Strong hands-on experience with monitoring and observability tooling, ideally Datadog, including alerting design and dashboards. * Solid understanding of SRE principles: SLOs, SLAs, error budgets, incident, problem and change management. * Experience with cloud environments and planning for capacity, resilience, disaster recovery and cost (FinOps). * Practical experience in performance and resilience testing, or a background in QA or test leadership with a strong operations focus. * Proven ability to define and implement operational processes (oncall, runbooks, RCA, post-incident reviews). * Strong analytical and problem-solving skills, able to investigate incidents across multiple data sources. * Excellent communication and stakeholder management skills for technical and non-technical audiences. * Fluent English, written and verbal; based in Romania.
Nice to have
* Familiarity with integration platforms, ideally MuleSoft, and API-based (Apigee) architectures. * Experience with release governance and CI/CD pipelines. * Experience in chaos or failover drills and seasonal peak readiness programs. * Scripting and automation skills to reduce operational toil.
Additional Information
What we offer at METRO.digital?
* Hybrid and agile work: thrive in a flexible, multicultural environment.
At METRO.digital, we promote work-life balance through a hybrid working model. You’ll be part of self-organizing, multicultural teams that collaborate in an agile setup.
* People development: when you grow so do we!
We want you to become the best version of yourself with individual and company-wide programs and trainings for people development. Focused among other on development, leadership, appreciation ... it´s time to upskill your career.
* Support with individual solutions:
We are people-caring! We offer support whenever you need it - at every stage of your professional journey.
We offer support whenever you need it - at every stage of your professional journey.
Want to know more about all our benefits? Discover more here.
Let´s connect soon. Apply for the role now!
Position grade within our career framework: Operations Manager (Md8).