Description

ABOUT THE JOB

Maintaining our start-up spirit, we prioritize thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and, of course, our business.

You’ll join the AI Team — the group driving all AI products and technology. We build and ship AI across the company: AI financial co-pilot, voice agent, and internal AI-powered processes. Our belief: your AI agent is only as good as your eval loop — we can build AI as good as the evals we run on it. Your mission: own that eval loop across every AI product we ship — pre-launch quality gates, post-launch monitoring, continuous improvement. Close collaboration with AI engineers, Product, and domain experts across the company. Core stack: Databricks, DeepEval, Claude Code

What You Will Be Doing
  • Own and extend our offline eval suite across products — datasets (capability + regression), judges, metrics

  • Build and maintain online quality dashboards: resolution rate, CSAT, thumbs up/down, LLM-as-judge signals, error rate, latency

  • Close the production feedback loop: mine failure patterns from real traffic → turn them into regression cases → propose fixes to Product and domain experts

  • Harden methodology: judge stability, non-determinism handling

  • Translate numbers into decisions – weekly syncs, clear trade-offs, no dashboards for their own sake

,

Must-Haves
  • Python and SQL — you can build an analysis end-to-end

  • Solid foundation in statistics — sampling, hypothesis testing, variance, understanding what a noisy metric is

  • Analytical mindset — you start from the business question, not from the tool

  • 3+ years in analyst / data scientist roles, at least one in a product context

,

Nice-to-Haves
  • Experience in quality analytics for ML systems — ranking, recommendations, classification, etc.

  • Hands-on experience evaluating LLM applications (RAG, agents, tool use, judges)

  • Experience building LLM agents — side projects, toy builds, personal experiments all count

,

How we work — one thing we mean seriously
  • AI-assisted coding is our default authoring environment, not a bonus

  • Claude Code is our main tool — you’ll reach for it for SQL, Python, analyses, dashboards, and internal scripts

  • We’re looking for analysts who are already curious and fluent with AI coding — or genuinely excited to become fluent fast

  • We care about what you ship and how clearly you think

  • If this idea excites you rather than worries you, you’ll feel at home here

Are you interested in this position?

Apply by clicking on the “Apply Now” Button below!

#JobsHubEstonia #GlobalRecrument
#CareerOpportunities #HiringNow
#JobSeekersNetwork #EstoniaJobs
#RecruitmentServices #EmploymentPortal.