Senior AI Backend Engineer - Agent Evaluation & Quality

Sallaยท Technology
Apply Now โ†—

About this role

About the role

We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that.

You'll build the evaluation systems behind our agents - the judges, test harnesses, and simulators that tell us whether an agent is working and where it's failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us.

Evaluation is the focus, but it won't be the boundary. Because you'll understand the agents' failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them.

Responsibilities

  • Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode.
  • Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes.
  • Build user simulators to generate test coverage and adversarial cases before real users hit them.
  • Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time.
  • Partner with product to turn "what good looks like" into concrete, measurable criteria.
  • Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.
  • Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering.
  • Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break.
  • A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.
  • Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability.
  • 5+ years software engineering, with recent hands-on LLM/agent work.

Nice to have

  • Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing.
  • Observability tooling (Arize, LangSmith, or similar).
  • Arabic language / NLP experience.
  • E-commerce or merchant-facing product experience.

Frequently Asked Questions

Is the salary disclosed for the Senior AI Backend Engineer - Agent Evaluation & Quality position at Salla?
The salary for this Senior AI Backend Engineer - Agent Evaluation & Quality role at Salla is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Is the Senior AI Backend Engineer - Agent Evaluation & Quality job at Salla remote?
Yes, this Senior AI Backend Engineer - Agent Evaluation & Quality position at Salla is remote, with team members based in Makkah, Makkah Province, Saudi Arabia, TELECOMMUTE. You can work from home or anywhere in the supported regions.
Which team or department does the Senior AI Backend Engineer - Agent Evaluation & Quality at Salla belong to?
This Senior AI Backend Engineer - Agent Evaluation & Quality position is part of the Technology department at Salla. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Senior AI Backend Engineer - Agent Evaluation & Quality position at Salla?
Click the "Apply Now" button on this page. You will be redirected to Salla's official application portal hosted on workable where you can submit your application directly.
When was the Senior AI Backend Engineer - Agent Evaluation & Quality job at Salla posted?
This Senior AI Backend Engineer - Agent Evaluation & Quality position at Salla was posted on Jul 19, 2026. Apply as soon as possible โ€” early applications are often reviewed first.
Senior AI Backend Engineer - Agent Evaluation & Quality
Salla
Apply for this role โ†—

You'll be redirected to Salla's official application page on workable.