Senior Site Reliability Engineer

Jolera· Professional Services
Apply Now ↗

About this role

Job Purpose

Lead for a team of site reliability engineers delivering who deliver incident detection, triage, and runbook-based remediation for production cloud-native environments, to support our North American customers. Set the operational standard for triage and recovery, act as the senior escalation point, and serve as the primary technical liaison to the Service Delivery Manager.

Key Responsibilities

• Lead incident detection, triage, and first response across production cloud and Kubernetes environments, to support our North American customers.

• Execute and oversee approved runbooks for service restoration — workload and node restarts, scaling, rollbacks, and database stabilization — within agreed operational boundaries.

• Act as the senior escalation authority; prepare clear escalation summaries covering impact, actions taken, current state, and recommended next steps.

• Author, review, and maintain operational runbooks; continuously improve detection, alerting, and automation.

• Engage cloud-provider support (AWS, GCP) for platform-level failures and vendor escalations.

• Technically supervise and mentor the SRE team; review handoffs and assure consistency across shifts.

• Own daily shift handoffs and contribute to monthly service reporting and reviews.

People Management

• Provides technical leadership and day-to-day supervision

• Contributes to coaching, performance input, and skills development; formal line management sits with the Service Delivery Manager.

Financial Responsibility

• Accountable for protecting service levels and cost-to-serve through efficient, automation-first operations.

• Key Performance Indicators (KPIs)

• Service-level (SLO/SLA) attainment

• Mean time to acknowledge / mean time to resolve

• Runbook coverage and quality

• Escalation accuracy and completeness

• Shift-handoff quality and reporting timeliness

• Repeat-incident reduction and automation adoption

Education & Certifications

• Bachelor’s degree in Computer Science, Engineering, or equivalent experience.

• Certified Kubernetes Administrator (CKA) required; CKAD, AWS, and Google Cloud certifications strongly preferred.

Experience

• 7+ years in SRE, DevOps, or production infrastructure operations, including 3+ years operating Kubernetes in production.

• Proven track record leading incident response for production cloud workloads.

• Managed-services / MSP or 24×7 operations experience preferred.

Skills & Competencies

Technical Skills

• Kubernetes operations across AWS EKS and GCP GKE

• AWS and GCP core services (compute, storage, networking, scaling, IAM)

• Relational database operational recovery (e.g., PostgreSQL)

• Observability platforms (e.g., Datadog)

• Scripting and automation (Bash, Python, Go or equivalent); read-level Terraform/IaC

• Incident command and structured troubleshooting

Soft Skills

• Calm, decisive incident leadership under pressure

• Clear written and verbal English

• Mentoring and team collaboration

• Time management

Tools / Software

• Datadog

• Jira / ServiceNow

• Confluence / GitHub Wiki

• AWS & GCP consoles

• Slack / Microsoft Teams

What We Offer

  • Competitive compensation package
  • Competitive benefits package
  • Company Perks, Good Life gym, and various brand discounts
  • Company events, recognitions, and celebrations
  • Career development and growth opportunities

Frequently Asked Questions

Is the salary disclosed for the Senior Site Reliability Engineer position at Jolera?
The salary for this Senior Site Reliability Engineer role at Jolera is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Senior Site Reliability Engineer position at Jolera located?
This Senior Site Reliability Engineer role at Jolera is based in Colombo, Western Province, Sri Lanka. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Which team or department does the Senior Site Reliability Engineer at Jolera belong to?
This Senior Site Reliability Engineer position is part of the Professional Services department at Jolera. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Senior Site Reliability Engineer position at Jolera?
Click the "Apply Now" button on this page. You will be redirected to Jolera's official application portal hosted on workable where you can submit your application directly.
When was the Senior Site Reliability Engineer job at Jolera posted?
This Senior Site Reliability Engineer position at Jolera was posted on Jun 16, 2026. Apply as soon as possible — early applications are often reviewed first.
Senior Site Reliability Engineer
Jolera
Apply for this role ↗

You'll be redirected to Jolera's official application page on workable.