About this role
We're BlueCat—a Great Place to Work for good reason. Our team solves critical network challenges for some of the world's largest organizations. In simple terms, we manage the systems that keep networks running smoothly, securely, and reliably—the backbone infrastructure that powers digital transformation for enterprises globally. Our Intelligent Network Operations platform delivers and enables AI-driven agentic ops at scale, automating and simplifying how companies manage, secure, and optimize their networks.
But what makes us different is how we work: we believe great work happens in an environment where you're trusted, heard, and supported. With teams globally, we're building a workplace culture that values collaboration and integrity as much as innovation. If you're looking to advance your career with a company that invests in its people, this is it.
About the Team
The BlueCat Horizon team powers all BlueCat SaaS products. Our mission is to deliver enterprise-grade software on a reliable, fast, globally distributed, and cost-efficient cloud platform. We are building an AI-first platform, where agentic systems run as core production workloads on Kubernetes (EKS) and AWS AgentCore.
About the Role
We are looking for a Site Reliability Engineer to help make our AI agent platform production-ready, secure, observable, and scalable. This role focuses on operating, automating, and hardening systems built by development teams, ensuring experimental AI capabilities are reliably delivered as enterprise-grade production services. You will work across platform engineering, infrastructure, and development teams to improve reliability, automation, and delivery speed.
What You’ll Do
Operate and scale Kubernetes (EKS) clusters running AI and cloud-native workloads
Support productionization of AWS AgentCore-based systems
Design, build, and maintain CI/CD pipelines using GitLab CI/CD for applications and platform services
Build and improve automated deployment workflows across environments
Implement and maintain observability systems (metrics, logs, traces)
Define and enforce SRE practices: SLIs, SLOs, alerting, incident response, postmortems
Partner with development teams to ensure services are production-ready and operable
Participate in on-call rotation and support incident resolution and root cause analysis
Build infrastructure using Terraform for repeatable deployments
Improve cost efficiency, performance, and reliability of the platform
What You Bring
5+ years of experience in SRE, DevOps, or Platform Engineering
Strong hands-on experience with AWS and Kubernetes (EKS)
Experience operating or supporting AWS AgentCore or similar AI/agent platforms
Strong proficiency in Python for automation
Solid understanding of AWS identity and access management (IAM), including SSO/Identity Center, roles, policies, trust relationships, SCPs, and cross-account access patterns
Experience with Infrastructure as Code (Terraform preferred)
Strong understanding of CI/CD using GitLab CI/CD
Ability to quickly onboard into existing AWS environments and operate independently
Experience working with on-call and incident management in production systems
Nice to Have
Experience with AI/LLM or agent-based systems
Running stateful workloads on Kubernetes (PostgreSQL, Redis, etc.)
Knowledge of RAG, vector databases, or knowledge systems
Multi-region AWS architecture experience
Cost optimization and capacity planning experience
Frequently Asked Questions
Is the salary disclosed for the Site Reliability Engineer – AI-first Platform position at bluecatnetworks?
The salary for this Site Reliability Engineer – AI-first Platform role at bluecatnetworks is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Site Reliability Engineer – AI-first Platform position at bluecatnetworks located?
This Site Reliability Engineer – AI-first Platform role at bluecatnetworks is based in Belgrade. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Site Reliability Engineer – AI-first Platform role at bluecatnetworks full-time or part-time?
This is listed as a Full Time position. It is posted as a Site Reliability Engineer – AI-first Platform role in the DNS Edge COGS department at bluecatnetworks.
Which team or department does the Site Reliability Engineer – AI-first Platform at bluecatnetworks belong to?
This Site Reliability Engineer – AI-first Platform position is part of the DNS Edge COGS department at bluecatnetworks. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Site Reliability Engineer – AI-first Platform position at bluecatnetworks?
Click the "Apply Now" button on this page. You will be redirected to bluecatnetworks's official application portal hosted on lever where you can submit your application directly.
When was the Site Reliability Engineer – AI-first Platform job at bluecatnetworks posted?
This Site Reliability Engineer – AI-first Platform position at bluecatnetworks was posted on Jul 13, 2026. Apply as soon as possible — early applications are often reviewed first.