Description
Engineering Mission
A technology-intensive fintech organisation in Estonia is seeking a Senior Fintech Reliability & Resilience Engineer to strengthen the reliability of systems responsible for critical financial transactions and customer-facing services.
This role goes beyond conventional infrastructure operations. You will engineer systems that remain dependable under traffic spikes, infrastructure failures, deployment changes, third-party outages, and other conditions that can affect financial services.
Your Responsibilities
- Design and implement reliability strategies for critical financial platforms.
- Establish service-level objectives, reliability indicators, and error budgets for core services.
- Analyse production incidents and lead systematic root-cause investigations.
- Build resilient architectures capable of handling infrastructure and service failures.
- Develop automated monitoring, alerting, tracing, and observability across distributed systems.
- Improve disaster-recovery and business-continuity capabilities.
- Conduct resilience testing and controlled failure experiments.
- Work with engineering teams to eliminate recurring reliability problems.
- Improve deployment safety through progressive delivery, automated rollback, and release controls.
- Monitor capacity, performance, and system dependencies as transaction volumes grow.
- Develop operational runbooks and incident-response procedures.
- Partner with security and infrastructure teams to strengthen production environments.
- Mentor engineers in reliability engineering, production readiness, and operational excellence.
Technical Environment
The role involves modern cloud-native technologies such as Kubernetes, Docker, Terraform, Linux, distributed systems, observability platforms, CI/CD, cloud infrastructure, databases, messaging systems, and automated operational tooling.
Candidate Profile
- 6+ years of experience in SRE, platform engineering, DevOps, infrastructure, or distributed-systems engineering.
- Strong experience operating high-availability production systems.
- Excellent understanding of cloud infrastructure, Kubernetes, networking, databases, and distributed architectures.
- Strong programming or scripting ability in Python, Go, Java, or a comparable language.
- Practical experience with incident management, observability, performance engineering, and disaster recovery.
- Strong understanding of secure production environments.
- Experience in fintech, banking, payments, or other mission-critical technology environments is highly advantageous.
Why This Opportunity
Financial platforms cannot simply be designed to work—they must be engineered to continue working when conditions are imperfect.
Ideal Candidate: A technically rigorous reliability engineer who enjoys solving difficult production problems and understands that resilience is a fundamental part of financial-product quality, not an afterthought.
Are you interested in this position?
Apply by clicking on the “Apply Now” Button below!
#JobsHubEstonia #GlobalRecrument
#CareerOpportunities #HiringNow
#JobSeekersNetwork #EstoniaJobs
#RecruitmentServices #EmploymentPortal.