Member of Technical Staff, AI Compute & Data Infrastructure

Vinci Palo Alto HQ Publicerat 15 september 2026
full_timeonsitesenior
ABOUT VINCI Every physical thing you touch exists because somebody successfully navigated the laws of physics: the chips in your phone, the vehicles on the road, the data centers powering AI. Physics determines what can be built, how well it performs, and where it breaks. Yet the tools engineers use to understand physical behavior are too slow and too specialized to use continuously while designing, so critical decisions get made with only a partial view of how a system will behave. Our mission is to make physical reasoning as accessible to engineers as language became through modern AI. This is not an attempt to build slightly better engineering software. It is an attempt to change how physical products are designed. Our technology is used today by many of the world's most advanced semiconductor and electronics organizations, including nearly half of the twenty largest companies in the industry. We are backed by Khosla Ventures and Eclipse Ventures. THE ROLE Training a foundation model for physics means holding petabytes of simulation data and keeping GPU clusters saturated with it. We are hiring the engineer who will own that layer: the storage where our training data lives, and the compute the AI team trains and validates models on. This is a founding role for the discipline, with a wide surface and a small team. You will set the compute and data architecture, and you will also be the person who finds out why a job has been sitting in pending for two hours. The AI team decides what to train and how to judge the result. Your job is to make sure the compute and the data are there when they need them, that the runs finish, and that we can afford it. The measure of the role is what the AI team can do with what you build: how many experiments they can run, how quickly results come back, how reliably a long run finishes without an infrastructure failure, and how much training we get for what we spend. Nothing in this layer stays fixed. Our training data will grow, and it will take on new sources and new types alongside what we have today. GPU hardware and the market for it move faster than most infrastructure, and our appetite for compute keeps increasing. We are hiring the person who leads us through that on the compute and data side: who identifies the next step before we are forced into it, makes the case for what it costs, and then builds it. WHAT YOU'LL DO - Compute architecture. Define how our GPU capacity is organized, provisioned, and grown, and how that architecture holds up as both the fleet and the team get larger. - Workload scheduling. Own how competing work claims that capacity: queueing, priority, preemption, gang scheduling, and quota between training runs, validation jobs, and data preparation. This starts as a judgment call among a handful of stakeholders and becomes an allocation problem worth solving algorithmically. Recognizing when that transition arrives, and building for it, is part of the role. - Data infrastructure at petabyte scale. Design the storage, ingestion, and access paths that keep training jobs fed at full throughput, and keep datasets versioned and reproducible as they change. Build for new sources and data types arriving at similar or larger scale, not only for what we hold today. - Training and validation platform. Own the systems the AI team uses to launch, checkpoint, resume, and monitor runs, and make sure validation workloads get the compute and data access they need without competing with training for it. - Reliability, utilization, and cost. Set and meet targets for cluster utilization, job success rate, and time from submission to first batch. Own the cost of a training run and be able to account for where it goes. - Hands-on engineering. Write and review production code, lead architecture reviews, and take the hardest debugging problems yourself. - Mentorship and teaching. Mentor the engineers who join this team as it grows, and help hire them. Review designs outside your own work. Make the AI team more capable with the infrastructure than they were before, and write things down so the answer to a recurring question lives somewhere other than in your head. - Technical direction. Work with the AI team on what the next model will require from infrastructure before it is required, and with leadership on capacity planning and compute investment. Say what our compute and data footprint needs to look like a year out and what it will cost. The leadership in this role is technical, not managerial. WHAT WE'RE LOOKING FOR - 10+ years building large-scale distributed systems, including several years running GPU infrastructure for large model training - Direct experience operating multi-node distributed training on modern accelerators: runs that held dozens or hundreds of GPUs for days at a time, with the scheduling, interconnect behavior, checkpointing, and failure recovery that requires. You have found out why a run was slower than the hardware allowed and fixed it. - Direct experience serving training data at petabyte scale, where throughput and storage cost were both constraints you had to answer for - Experience setting scheduling and quota policy on shared GPU capacity, with a view on the tradeoff between fleet utilization and how long people wait in the queue - A track record of building infrastructure whose users are researchers and engineers, and of being measured on what those users were able to do with it, not on the system itself - Infrastructure you built that survived a substantial change in scale or in the character of the workload, with a clear account of what held up and what you had to replace - A history of mentoring engineers and of teaching people outside your specialty enough to work on their own - The ability to set technical direction, make the case for it to people outside infrastructure, and then implement it. At this stage the role is one engineer and a small team, not one engineer directing several. - Ad

Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.

Visa jobbet och ansök

Ursprunglig annons: jobs.ashbyhq.com