Member of Technical Staff, ML Inference Engineering

Sanas· Engineering
Apply Now ↗
📍 Palo Alto, CAFULL TIME

About this role

About the Role

Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do

Performance Optimization

  • Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference
  • Analyze and improve latency, throughput, memory usage, and compute efficiency
  • Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
  • Implement low-level optimizations using CUDA, Triton, and other performance tooling
  • Improve support for mixed precision, quantization, and model graph optimization
  • Build and maintain performance benchmarking and monitoring infrastructure
  • Scale inference and training systems across multi-GPU, multi-node environments

Inference Systems & Reliability

  • Own and evolve our inference engine, enabling reliability and performance at scale
  • Develop and optimize runtime inference services for large-scale AI applications
  • Implement robust, fault-tolerant systems for data ingestion and processing

Requirements

Must-have:

  • 5+ years of experience writing high-quality, high-performance code
  • Familiarity with NVIDIA GPU architecture and CUDA
  • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
  • A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to
  • A record of shipping research or systems that other people build on, whether in a lab or in industry

Nice-to-have:

  • Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
  • Experience developing large-scale, high-load production systems
  • Experience maintaining or contributing to open-source ML projects
  • Experience managing machine learning workloads on Kubernetes clusters
  • Experience with InfiniBand or RoCE networking
  • Experience with bare-metal provisioning and lifecycle management
  • Experience operating large-scale AI training or inference clusters
  • Experience with hardware health monitoring and predictive failure detection
  • Experience with distributed storage systems

Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more.

Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language.

Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house.

Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication.

If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you.

Frequently Asked Questions

Is the salary disclosed for the Member of Technical Staff, ML Inference Engineering position at Sanas?
The salary for this Member of Technical Staff, ML Inference Engineering role at Sanas is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Member of Technical Staff, ML Inference Engineering position at Sanas located?
This Member of Technical Staff, ML Inference Engineering role at Sanas is based in Palo Alto, CA. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Member of Technical Staff, ML Inference Engineering role at Sanas full-time or part-time?
This is listed as a FULL TIME position. It is posted as a Member of Technical Staff, ML Inference Engineering role in the Engineering department at Sanas.
Which team or department does the Member of Technical Staff, ML Inference Engineering at Sanas belong to?
This Member of Technical Staff, ML Inference Engineering position is part of the Engineering department at Sanas. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Member of Technical Staff, ML Inference Engineering position at Sanas?
Click the "Apply Now" button on this page. You will be redirected to Sanas's official application portal hosted on rippling where you can submit your application directly.
When was the Member of Technical Staff, ML Inference Engineering job at Sanas posted?
This Member of Technical Staff, ML Inference Engineering position at Sanas was posted on Aug 10, 2026. Apply as soon as possible — early applications are often reviewed first.
Member of Technical Staff, ML Inference Engineering
Sanas
Apply for this role ↗

You'll be redirected to Sanas's official application page on rippling.