Scientist II / Senior ML Scientist, Data-Efficient Learning for Drug Discovery
Lilasciences
Cambridge, MA USA; London; San Francisco, CA USA, United States · Onsite · Full Time
Posted
Job description
Your Impact at LILA Lila Sciences is seeking a Machine Learning Scientist, Data-Efficient Learning for Drug Discovery to build models and learning strategies for settings where data is scarce, expensive, and intentionally generated. This role is focused on training useful models from low-quantity but high-quality datasets ranging from as few as tens to low thousands of examples, often in tightly focused areas of chemical space, and deciding what data should be acquired next. This is an applied scientific ML role in a frontier research area. The work is not a matter of applying standard models out of the box. You will use and develop approaches across active learning, meta-learning, fine-tuning, uncertainty estimation, experimental design, and multimodal modeling to help Lila build closed-loop systems that learn efficiently from targeted data acquisition. This role connects model training with scientific decision-making: data acquisition plans should be useful to computational chemists evaluating compound priorities, computational biophysicists deciding when simulation is warranted, and cofolding modelers deciding which protein-ligand data would improve structure-aware models. What You'll Be Building Build ML models that perform well in low-data regimes for drug discovery and molecular optimization. Design data acquisition strategies that identify which compounds, assays, DEL selections, simulations, structural predictions, or experiments should be run next to maximize learning. Develop active learning, meta-learning, fine-tuning, transfer learning, and uncertainty-aware modeling approaches for focused chemical spaces. Train models on low-quantity, high-quality datasets generated by Lila's experimental, computational, and agentic discovery systems. Build multimodal models that can integrate DEL data, simulation outputs, assay data, protein and structural information, chemical features, literature or text-derived signals, images, and experimental metadata. Partner with experimental, computational, and drug discovery teams to ensure data acquisition plans are scientifically meaningful and operationally feasible. Evaluate models through learning curves, prospective validation, retrospective benchmarks, uncertainty calibration, and decision-focused metrics. Develop closed-loop learning workflows that continuously update models as new data arrives from experiment…