We're combining a technology called DNA encoded libraries (DEL) with machine learning. DELs let you make very large numbers of compounds without losing track of each in a large mixture (1 million to 1 billion compounds in one tube). This means you can run a whole lot of experiments in parallel (basically 1e6-1e9 per tube).
We believe this is a very promising way to generate large volumes of labelled data, which we know is what ML loves. The trick is to structure both the biochemistry and the ML properly so as to avoid artifacts and generate useful molecules reliably.
Happy to talk about it in more detail, give me a shout at ntilmans at anagenex.com!