Member of Technical Staff, Software
Substrate Bio
London · Onsite · Full Time
Posted
Job description
The opportunity Substrate is building a laboratory that runs itself. Something has to turn a scientist's intent into work the instruments actually execute, schedule it across the lab, and capture everything that happens as structured data. That software does not fully exist yet. It is being written now, from the first line, by a small, elite engineering team -and you would build it with them. We call this our infrastructure software layer: customer intent in, executed experiments and clean, agent-ready data out, with full provenance captured as the lab runs. Provenance is one half of the bar, with scientific quality, that makes Substrate's data worth training on. About Substrate Substrate is building the critical infrastructure layer between AI and biology: an AI-native automated lab that produces biological data at scale. AI for biology has a data problem, not a compute problem. Biological foundation models can predict, but they cannot run experiments, and the high-quality, large-scale data they need does not exist. Substrate generates it, with quality and provenance built in. We are venture-backed, building our first lab at 20 Triton Street in London, with US expansion to follow. What started as four co-founders is now a rapidly expanding team across science, intelligence, software, operations and partnerships, with people who have come from Automata, Palantir, Owkin, Illumina and Exscientia. We expect to be over 30 people within the year. We are not a cloud lab and we are not a CRO. We are the infrastructure that turns scientific intent into executed experiments and structured, AI-ready data, and over time into proprietary datasets and our own infrastructural intelligence. What you’ll do You will build the infrastructure software that runs the lab, working across the stack with the founding software engineer and the team. There are two products. The execution product turns a customer's intent into executed lab work: a translation layer converts an experiment into versioned, runnable workflows, an orchestration layer schedules and runs them across the lab on top of Automata's LINQ, and the output lands as structured, AI-ready data under a shared ontology. The observation product captures metadata everywhere it is generated and maps it into a knowledge graph, so every run carries full provenance. Where you land depends on you and on what the lab needs next…