Research Engineer, Infrastructure
Get alerts for roles like this
More Research Engineer, Infrastructure roles in San Francisco — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
Who We Are We are an applied AI lab building end-to-end software agents. We're the team behind Devin, the first AI software engineer, and Windsurf, an AI-native IDE. These products represent our vision for AI that doesn't just assist engineers, but works alongside them as a genuine teammate.
Our team is small and talent-dense: world-class competitive programmers, former founders, and researchers from the frontier of AI, including Scale AI, Palantir, Cursor, Google DeepMind, and others. Role Mission Research moves at the speed of the infrastructure underneath it.
Every training run, evaluation loop, and experimental iteration depends on systems that are fast, reliable, and built to scale. This role exists to make sure nothing in the stack becomes the bottleneck that slows down the frontier.
You will own the core systems that researchers depend on daily: distributed training infrastructure, experiment orchestration, data pipelines, and the tooling that turns raw compute into usable research velocity. This is not a support role.
You will work directly alongside researchers, understand the science deeply enough to anticipate what they need next, and build systems that hold up under the pressure of training jobs running across thousands of GPUs.
We don't distinguish between research and engineering; the best infrastructure engineers here are also the ones who understand why the research works. What You'll Accomplish Distributed Training Infrastructure: Build and own the systems that run large-scale training jobs reliably across GPU clusters.
This includes job launchers, checkpointing and recovery, fault tolerance, and the monitoring that keeps researchers informed and unblocked.
Scaling Agent Rollouts: Own the infrastructure that runs hundreds of thousands of concurrent coding agent rollouts in VM sandboxes, from high-fidelity environment design to the distributed systems that hold up at our largest RL training scales. Performance Optimization: Profile and improve traini
Similar Jobs
Other open roles matched to this job's skills and location.