Inference Engineer

Weekday AIยท Weekday's Client via platform
Apply Now โ†—
๐Ÿ“ Bengaluru, Karnataka, IndiaFull time

About this role

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿฐ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฏ๐Ÿฒ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿฎ๐Ÿฐ-๐Ÿฏ๐Ÿฒ ๐—Ÿ๐—ฃ๐—”)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are seeking a highly skilledย Inference Engineerย with strong expertise inย Large Language Models (LLMs)ย andย vLLMย to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.

As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.

Key Responsibilities

  • Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
  • Build and optimize high-performance model serving pipelines usingย vLLMย and other modern inference frameworks.
  • Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
  • Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
  • Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
  • Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
  • Develop APIs, microservices, and deployment workflows for AI-powered applications.
  • Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
  • Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
  • Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.

What Makes You a Great Fit

  • 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
  • Strong hands-on expertise withย Large Language Models (LLMs)ย andย vLLMย for production-scale inference.
  • Experience deploying and optimizing transformer-based models using modern inference frameworks.
  • Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
  • Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
  • Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
  • Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
  • Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
  • Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
  • Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
  • Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
  • Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.

Frequently Asked Questions

Is the salary disclosed for the Inference Engineer position at Weekday AI?
The salary for this Inference Engineer role at Weekday AI is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Inference Engineer position at Weekday AI located?
This Inference Engineer role at Weekday AI is based in Bengaluru, Karnataka, India. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Inference Engineer role at Weekday AI full-time or part-time?
This is listed as a Full time position. It is posted as a Inference Engineer role in the Weekday's Client via platform department at Weekday AI.
Which team or department does the Inference Engineer at Weekday AI belong to?
This Inference Engineer position is part of the Weekday's Client via platform department at Weekday AI. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Inference Engineer position at Weekday AI?
Click the "Apply Now" button on this page. You will be redirected to Weekday AI's official application portal hosted on workable where you can submit your application directly.
When was the Inference Engineer job at Weekday AI posted?
This Inference Engineer position at Weekday AI was posted on Jul 29, 2026. Apply as soon as possible โ€” early applications are often reviewed first.
Inference Engineer
Weekday AI
Apply for this role โ†—

You'll be redirected to Weekday AI's official application page on workable.