Synced from Ashby · Jun 15

HPC Infrastructure Engineer

SpellbrushSan Francisco or TokyoPosted Jun 15, 2026
Infrastructure EngineerOn-siteSenior
Apply now - freeSave & get alerts

Mirrored from Spellbrush's own Ashby careers system · refreshed hourly

10
Other open Spellbrush roles
Jun 15
Posted
Ashby
Applicant system
Job description

We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.

You may be a good fit if:

You love anime and the anime aesthetic.

This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.

You’re familiar with the modern HPC software landscape

Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)

As well as the traditional sysadmin landscape

Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools.

You're not afraid of physical computers

We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.

And you're comfortable working on small, fast-paced teams.

We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world.

We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.

View original posting on Ashby

What applying to Spellbrush usually looks like

Based on publicly available information, candidates applying through ashby can generally expect an online application form covering resume submission, work history, and sometimes short screening questions tailored to the role, whether that be ai-engineer, frontend-engineer, game-designer, or another open position. The process may include multiple stages, such as an initial recruiter or hiring manager conversation, followed by technical or role-specific assessments, and possibly a take-home exercise or portfolio review depending on the discipline. Interviews conducted via Ashby often involve structured scorecards, so candidates may notice consistent question themes across stages. Communication is typically handled through automated status updates within the platform, and response times vary by company and role. Applicants to Spellbrush should prepare materials relevant to the specific position and expect the process to reflect general practices common among companies using this ATS.

Based on publicly available information. LandEarly does not verify interview process details.

Land this one early - before the req fills.

LandEarly tailors your resume and screening answers to each posting, submits within minutes of a role going live, and tracks every application in one place.

Free to start · No credit card · Cancel anytime

Keep exploring

What this role pays, where else it is open, and how to write the application.

HPC Infrastructure Engineer
Spellbrush · San Francisco or Tokyo
Apply