Build what comes next.
Help build the systems that bring personalized models into everyday use.
Infrastructure Engineer —
Inference & GPU Systems
The work
Run the models we serve. You’ll own inference end to end: how models are deployed and served, how requests are batched and scheduled across GPUs, how new hardware is benchmarked before it carries real traffic, and how the whole system is observed and kept reliable. The steady work is making every model faster and cheaper to run, without changing what it says.
Relevant experience
Linux, containers, GPU systems, distributed services, and diagnosing performance problems. Experience with inference engines such as vLLM, SGLang, or TensorRT-LLM, and with batching, KV-cache management, or quantization, is a plus.
Send a short introduction and links to your work.
Start a conversation
Member of Technical Staff —
Post-training & Personalization
The work
Shape how models adapt to the people who use them. Fine-tuning, training data, evaluation, and bringing adapted models into inference.
Relevant experience
Hands-on machine learning, PyTorch or equivalent, experiment design, and research code others can rely on.
Send a short introduction and links to your work.
Start a conversation