Inference Engineer
About this role
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ฎ๐ฐ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฏ๐ฒ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ฎ๐ฐ-๐ฏ๐ฒ ๐๐ฃ๐)
Experience: 5+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are seeking a highly skilledย Inference Engineerย with strong expertise inย Large Language Models (LLMs)ย andย vLLMย to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.
As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.
Key Responsibilities
- Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
- Build and optimize high-performance model serving pipelines usingย vLLMย and other modern inference frameworks.
- Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
- Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
- Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
- Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
- Develop APIs, microservices, and deployment workflows for AI-powered applications.
- Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
- Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
- Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.
What Makes You a Great Fit
- 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
- Strong hands-on expertise withย Large Language Models (LLMs)ย andย vLLMย for production-scale inference.
- Experience deploying and optimizing transformer-based models using modern inference frameworks.
- Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
- Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
- Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
- Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
- Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
- Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
- Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
- Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
- Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.
Frequently Asked Questions
Is the salary disclosed for the Inference Engineer position at Weekday AI?
Where is the Inference Engineer position at Weekday AI located?
Is the Inference Engineer role at Weekday AI full-time or part-time?
Which team or department does the Inference Engineer at Weekday AI belong to?
How do I apply for the Inference Engineer position at Weekday AI?
When was the Inference Engineer job at Weekday AI posted?
You'll be redirected to Weekday AI's official application page on workable.