Scientist II / Senior ML Scientist, Cofolding and Structure-Aware ML
Lilasciences
Cambridge, MA USA; London; San Francisco, CA USA, United States · Onsite · Full Time
Posted
Job description
Your Impact at LILA Lila Sciences is seeking a Machine Learning Scientist, Cofolding and Structure-Aware ML to train next-generation cofolding models for drug discovery. This role is focused on improving models that reason over proteins, ligands, binding context, and experimental data, potentially using contrastive learning and related representation-learning approaches. This person should have direct experience training modern scientific ML models, not only using pretrained systems. You will work with ML researchers, computational chemists, computational biophysicists, data engineers, and drug discovery teams to develop models that learn from DEL and related datasets, connect molecular and protein context, and improve AI-driven discovery decisions. The models developed in this role should produce outputs that medicinal and computational chemists as well as biophysicists can interrogate, validate, and use in downstream agent-driven discovery decisions. What You'll Be Building Train and evaluate cofolding models for protein-ligand and related molecular discovery applications. Use contrastive learning, representation learning, self-supervised learning, or related methods where they help improve cofolding models trained on molecules, proteins, structures, and experimental readouts. Develop modeling approaches that make DEL data more useful for learning binding, enrichment, selectivity, and structure-activity signals. Build and evaluate models informed by Boltz, AlphaFold-style cofolding, equivariant GNNs, and related structure-aware ML methods. Design training objectives, including contrastive, self-supervised, or multimodal objectives, that connect ligands, proteins, structures, assays, simulations, and experimental data. Build rigorous evaluation frameworks that distinguish meaningful molecular learning from dataset artifacts, leakage, or spurious correlations. Collaborate with data and platform teams to define datasets, labels, negatives, controls, and metadata needed for model training. Partner with computational chemistry and biophysics teams to connect model outputs to physically and chemically meaningful hypotheses. Work with low-data learning scientists to identify which DEL, assay, simulation, or structural data would most improve model performance in focused chemical spaces. Work with research engineers to scale training, inference, and evaluation wo…