Description
A technology-focused financial services organisation in Estonia is seeking a Fintech Infrastructure Reliability Manager to strengthen the reliability, scalability, and operational resilience of its core technology environment.
This role is suited to an experienced infrastructure professional who understands that financial platforms require exceptional availability, controlled change management, rapid incident response, and measurable operational performance.
Your Mission
You will lead initiatives that improve platform reliability while helping engineering and infrastructure teams operate increasingly complex fintech systems with greater automation, visibility, and resilience.
Key Responsibilities
- Lead infrastructure reliability and operational-resilience initiatives across critical fintech platforms.
- Establish service-level objectives, reliability metrics, and operational performance standards.
- Improve monitoring, alerting, logging, tracing, and system observability.
- Lead the development of incident-management and escalation practices.
- Analyse recurring incidents and coordinate root-cause remediation.
- Partner with engineering teams to improve application and infrastructure reliability.
- Strengthen deployment, rollback, disaster-recovery, and capacity-management processes.
- Automate infrastructure and operational workflows wherever practical.
- Support cloud architecture, container orchestration, networking, and infrastructure optimisation.
- Develop resilience testing and operational-readiness practices.
- Maintain clear documentation covering critical services, dependencies, recovery procedures, and operational controls.
- Mentor infrastructure and reliability engineers and promote strong engineering practices.
Candidate Profile
- 6+ years of experience in infrastructure engineering, SRE, DevOps, cloud engineering, or platform operations.
- Strong experience with AWS, Azure, GCP, or comparable cloud environments.
- Excellent knowledge of Kubernetes, containers, Linux, networking, CI/CD, and infrastructure-as-code.
- Experience operating distributed systems with demanding availability requirements.
- Strong understanding of observability, incident response, disaster recovery, and capacity planning.
- Experience in fintech, banking, payments, or another high-availability industry is highly valuable.
- Strong leadership and problem-solving capabilities.
Technical Environment
Kubernetes, cloud infrastructure, infrastructure-as-code, container platforms, CI/CD pipelines, distributed systems, databases, observability tooling, automated deployment, security controls, and resilience testing.
Why This Opportunity
You will have direct responsibility for improving the reliability of technology that supports financial services where downtime, data integrity, and operational disruption can have significant consequences.
Ideal Candidate: A technically strong reliability leader who combines hands-on infrastructure knowledge with disciplined operational management and a continuous-improvement mindset.
Are you interested in this position?
Apply by clicking on the “Apply Now” Button below!
#JobsHubEstonia #GlobalRecrument
#CareerOpportunities #HiringNow
#JobSeekersNetwork #EstoniaJobs
#RecruitmentServices #EmploymentPortal.