Description
About the role
Our Site Reliability Engineers in the Platform team are working towards making sure we build up an ever more resilient and scalable platform for engineering teams to build on top of.
Main tasks and responsibilities:
Developing automation and implementing internal systems.
Participating in architectural discussions to select the best long-term solution and collaborating with other SRE Engineers and Software Engineers.
Implementation of new systems and tuning of existing ones to enable better reliability, availability, scalability, or reduce costs.
Troubleshooting production system issues.
Making sure all main services are measured, monitored, and covered by alerting systems.
Participating in on-call activities for the most critical parts of the infrastructure, helping to solve incidents in a no-blaming culture.
About you:
You have in-depth knowledge of Linux troubleshooting, including networking, filesystems, security, and the kernel.
You have a solid understanding of large-scale distributed systems.
You are experienced with AWS services, such as EC2 and EKS.
You have experience with infrastructure automation tools, such as Terraform and Ansible.
You are a team player, with good communication skills, that works well with others in the group and the rest of the organization.
Are you interested in this position?
Apply by clicking on the “Apply Now” Button below!
#JobsHubEstonia #GlobalRecrument
#CareerOpportunities #HiringNow
#JobSeekersNetwork #EstoniaJobs
#RecruitmentServices #EmploymentPortal
Related Jobs
Description About the role Our Site Reliability teams solve a broad range of challenges, from keeping the infrastructure behind Bolt’s backend services reliable to continuously improving how we operate at scale. The Site Reliability Cloud...
Remote
Description ABOUT THE JOB In this role, you’ll: Own and evolve Webshare’s production infrastructure – lead the migration from Docker Swarm to Kubernetes (or hybrid K8s + Ansible). Maintain high availability across hundreds of servers...
Remote
Description About the role As a Site Reliability Engineer, you will help design, automate, and operate our infrastructure across GCP and AWS. You will apply cloud governance and reliability best practices, participate in an on-call...
Remote
×