Description
ABOUT THE JOB
You run and improve our cloud-native platform day to day. You work independently with minimal guidance, you’re a trusted resource for less-experienced colleagues, and your impact lands within Engineering and the teams you support. This is a hands-on role for an engineer who already runs production Kubernetes with confidence and treats AI agents as part of their toolchain.
Our whole business, from a loan application to a funding decision, runs on a cloud-native platform on AWS: EKS, Istio, Flux GitOps, Terraform, Helm, and our own in-house-built cloud infrastructure operator that provisions cloud resources (queues, buckets and more) declaratively, straight from Kubernetes. On that same platform runs our agentic AI operating system and the AI infrastructure behind it. And yes — you’ll build with agentic coding tools every day. Most of our platform team already does.
Your responsibilities
-
Operate and extend our AWS/EKS platform, Istio service mesh, Flux GitOps, KEDA autoscaling, OPA Gatekeeper, cert-manager, Velero and ALB ingress, across dev, acceptance and production.
-
Run and extend our in-house cloud infrastructure operator that lets teams provision cloud resources (queues, buckets and more) declaratively, straight from Kubernetes.
-
Build and maintain shared CI/CD templates and self-service tooling so teams ship with less friction.
-
Own reliability and observability (Prometheus, Grafana, Loki, Tempo) for your area, and turn incidents into runbooks so they don’t happen twice.
-
Solve complex – but largely bounded or recurring – platform problems, balancing short- and long-term trade-offs.
-
Automate the toil: if something is manual and repetitive, your instinct is to make the platform (or an agent) do it.
What you bring
-
Minimum 4 years of platform / DevOps / SRE experience, and you can run Kubernetes in production with minimal guidance (EKS a plus).
-
Strong infrastructure-as-code and GitOps discipline — Terraform, Helm, Flux or Argo CD — and solid CI/CD (GitLab a plus).
-
You’re genuinely AI-native: agentic coding tools (Claude Code, Cursor, and friends) are already part of how you build, and you’re excited to push that frontier rather than sceptical of it. This is the mindset we care about most.
-
A security- and reliability-first instinct: you think about blast radius, least privilege and observability before things break.
-
A first-principles thinker and an owner: you finish things, you write the runbook, and you warn affected teams before they feel the change.
-
Bonus: hands-on experience with AI/LLM infrastructure — vector stores, RAG, LLM observability, MCP, or Bedrock/OpenAI.
Don’t tick every box? Tell us anyway — we hire for trajectory and judgement, not a checklist. That’s usually how the best engineers start here.
Are you interested in this position?
Apply by clicking on the “Apply Now” Button below!
#JobsHubEstonia #GlobalRecrument
#CareerOpportunities #HiringNow
#JobSeekersNetwork #EstoniaJobs
#RecruitmentServices #EmploymentPortal