Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.
In this role, you'll own token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new request scheduling and parallelism strategies at the engine level, and help us use a heterogenous mix of hardware at max efficiency.
What you’ll do
- Modify and extend state-of-the-art inference engines like vLLM and SGLang.
- Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile.
- Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
- Write custom GPU kernels to excel in specific regimes, such as cascade attention https://flashinfer.ai/2024/02/02/cascade-inference.html
What we’re looking for
- Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
- Interest in MLSys research - great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.
- Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
- Great interpersonal communication - please don't use LLMs to write prose. We desk-reject most LLM-generated cover letters and resumes.
Interview process
1. Meet the CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear.
2. Meet the CEO. This is an early step because we respect your time. Ask any question and get a definitive answer immediately.
3. Come in to Sail's SF office for an interview day. Meet the whole team, then you'll have 3-4 hours to work on a problem that closely simulates the work we do daily. It's an objectively scored task, so you'll have immediate feedback on how well your code is working - just like we do in production! AI assistance is highly encouraged, and we'll provide a laptop with all the best tools set up. Finish with a short presentation describing your process, learnings, and results.
4. Offer. Once the team decides we want to work with you, we make a strong offer quickly and will be quite persistent over email/text/calls :)
LIFE AT SAIL
We work out of a beautiful, sunny office in downtown San Francisco. All meals are on us (and actually great; SF is a food paradise and it would be a shame to eat only bowl slop). Everyone gets a Studio Display at their desk. We are serious about investing in anything that saves us time or energy. There are six different ways to make coffee or tea in the office. A friendly (hypoallergenic) black cat named Coco visits occasionally.
Get roles like this in your inbox
New openings at recently funded startups, sent as they appear. Free.