HPC/ML Infrastructure Engineer

spellbrush· Research Team
Apply Now ↗
📍 San Francisco or TokyoFullTime

About this role

We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.

You may be a good fit if:

You love anime and the anime aesthetic.

This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.

You’re familiar with the modern HPC software landscape

Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)

As well as the traditional sysadmin landscape

Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools.

You're not afraid of physical computers

We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.

And you're comfortable working on small, fast-paced teams.

We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world.

We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.

Frequently Asked Questions

Is the salary disclosed for the HPC/ML Infrastructure Engineer position at spellbrush?
The salary for this HPC/ML Infrastructure Engineer role at spellbrush is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the HPC/ML Infrastructure Engineer position at spellbrush located?
This HPC/ML Infrastructure Engineer role at spellbrush is based in San Francisco or Tokyo. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the HPC/ML Infrastructure Engineer role at spellbrush full-time or part-time?
This is listed as a FullTime position. It is posted as a HPC/ML Infrastructure Engineer role in the Research Team department at spellbrush.
Which team or department does the HPC/ML Infrastructure Engineer at spellbrush belong to?
This HPC/ML Infrastructure Engineer position is part of the Research Team department at spellbrush. See the full job description for more information about the team structure and responsibilities.
How do I apply for the HPC/ML Infrastructure Engineer position at spellbrush?
Click the "Apply Now" button on this page. You will be redirected to spellbrush's official application portal hosted on ashby where you can submit your application directly.
When was the HPC/ML Infrastructure Engineer job at spellbrush posted?
This HPC/ML Infrastructure Engineer position at spellbrush was posted on Jun 15, 2026. Apply as soon as possible — early applications are often reviewed first.
HPC/ML Infrastructure Engineer
spellbrush
Apply for this role ↗

You'll be redirected to spellbrush's official application page on Ashby ATS.