Site Reliability Engineer — Multi-Cloud Infrastructure

skit· Technology
Apply Now ↗
📍 Bangalore, KA, IndiaFULL TIME

About this role

About the Role
Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/

Job Title: Site Reliability Engineer — Multi-Cloud Infrastructure
Type: Full-time
Location: Bangalore

Why this role exists:
We run a voice AI platform for regulated enterprises in banking, telecom, and collections, spread across AWS, GCP, and Azure — for resilience, for cost, and because client data-residency rules leave us no choice. That's a lot of surface area: compute, networking, storage, identity, clusters, pipelines, and supporting services, all needing to stay healthy across three providers.
This role owns the day-to-day reliability and operations of that estate. It's the generalist counterpart to our real-time-platform SRE: where they go deep on the latency-critical call path, you go broad — keeping the whole infrastructure dependable, well-automated, and cost-sane, and sharing the on-call load. If you like knowing how everything fits together and making the boring parts reliable and self-serve, this is a good seat.

What you'll own:
  • Multi-cloud operations. Provision, operate, and keep healthy compute, networking, storage, and identity across AWS, GCP, and Azure — with sensible consistency instead of three snowflakes.
  • Infrastructure as code. Manage the estate through Terraform (or equivalent) and version control — reproducible environments, reviewed changes, no undocumented hand-tweaks.
  • CI/CD and delivery. Keep build and deploy pipelines fast and reliable so engineers ship safely and often.
  • Clusters and workloads. Run Kubernetes/container platforms and the supporting services (databases, queues, caches, internal tooling) that everything depends on.
  • Monitoring and on-call. Maintain monitoring and alerting for infrastructure health, take a turn in the rotation, and respond to and mitigate incidents with clear communication and blameless follow-up.
  • Cost and hygiene. Keep an eye on cloud spend, rightsizing, and waste; own the unglamorous but essential hygiene — patching, backups, secrets, and access.
  • Automation and toil reduction. Replace manual, repetitive operations with automation and self-service so the team scales without headcount scaling with it.

What the first year looks like:
  • First 90 days. Learn the estate across all three clouds. Take a turn on call. Close the most obvious gaps in monitoring, backups, and access hygiene.
  • By 6 months. More of the estate under consistent infrastructure-as-code. Reliable, reviewed CI/CD. A clearer, quieter alerting setup and documented runbooks for the common incidents.
  • By 12 months. Measurably less manual toil through automation and self-service. Sensible cost controls in place. Provisioning and environment setup that's repeatable rather than tribal knowledge.

What we're looking for:
Must-have
  • A few years in SRE, DevOps, or infrastructure operations for production systems, including on-call.
  • Hands-on experience across at least two of AWS, GCP, and Azure (all three is a strong plus).
  • Kubernetes and containers in production.
  • Infrastructure-as-code (Terraform or similar) and CI/CD pipelines.
  • Monitoring and alerting practice (e.g. Prometheus/Grafana) and structured incident handling.
  • A scripting/programming language for automation (Python, Go, or Bash beyond one-liners).
  • Solid Linux systems and networking fundamentals.
Nice-to-have
  • All three clouds run in production, and comfort designing for consistency across them.
  • Cost optimization / FinOps.
  • Secrets management, security hardening, and compliance/data-residency contexts.
  • PostgreSQL and other stateful-service operations at scale.
  • Some exposure to real-time or voice infrastructure — enough to back up the platform SRE on call.

Our Stack:
Representative — you'll help shape it. Multi-cloud across AWS, GCP, and Azure; Kubernetes/containers; Terraform and GitHub Actions CI/CD; PostgreSQL; Grafana/Tempo for monitoring; Modal for ML deployment; LiveKit/SIP telephony on the platform side.

How you'll know you're succeeding:
The infrastructure just works, across all three clouds, and when it doesn't it's caught early and fixed cleanly. Engineers provision what they need without filing tickets. Cloud spend is understood, not surprising. And the on-call rotation trends calmer because the estate is increasingly automated and self-healing.

We're an equal-opportunity employer and evaluate every candidate on merit. [Add benefits, compensation band, and application instructions before posting.]


Frequently Asked Questions

Is the salary disclosed for the Site Reliability Engineer — Multi-Cloud Infrastructure position at skit?
The salary for this Site Reliability Engineer — Multi-Cloud Infrastructure role at skit is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Site Reliability Engineer — Multi-Cloud Infrastructure position at skit located?
This Site Reliability Engineer — Multi-Cloud Infrastructure role at skit is based in Bangalore, KA, India. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Site Reliability Engineer — Multi-Cloud Infrastructure role at skit full-time or part-time?
This is listed as a FULL TIME position. It is posted as a Site Reliability Engineer — Multi-Cloud Infrastructure role in the Technology department at skit.
Which team or department does the Site Reliability Engineer — Multi-Cloud Infrastructure at skit belong to?
This Site Reliability Engineer — Multi-Cloud Infrastructure position is part of the Technology department at skit. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Site Reliability Engineer — Multi-Cloud Infrastructure position at skit?
Click the "Apply Now" button on this page. You will be redirected to skit's official application portal hosted on keka where you can submit your application directly.
When was the Site Reliability Engineer — Multi-Cloud Infrastructure job at skit posted?
This Site Reliability Engineer — Multi-Cloud Infrastructure position at skit was posted on Jul 5, 2026. Apply as soon as possible — early applications are often reviewed first.
Site Reliability Engineer — Multi-Cloud Infrastructure
skit
Apply for this role ↗

You'll be redirected to skit's official application page on keka.