Skip to main content

How do you leverage federated learning models in drug discovery?

Data modelling

Depending on what modelling workflows you use, you can also use more noisy data to inform certain modelling setups – one of the opportunities offered by the capabilities of data management in drug discovery. Find out more in our exclusive White Paper (Image: Shutterstock)

There are a few questions that drug discovery teams need to address before considering the adoption of federated learning techniques. Artificial Intelligence (AI) has become critical to modern drug discovery, but as models become more sophisticated, the quality and diversity of the data used to train them has become an equally important competitive advantage. Pharmaceutical firms, biotechs and research organisations hold swathes of valuable data, yet concerns around intellectual property, patient privacy and commercial sensitivity make sharing it difficult. Federated learning is widely considered to be the solution.

 

Instead of moving sensitive datasets between organisations, federated learning enables AI models to be trained collaboratively while the underlying data remains within each organisation. It results in an opportunity to build more accurate AI models without compromising ownership (and adjacent confidentiality concerns).

 

“Where companies need to look is how do you apply these models now to make better decisions or, in the case of large pharma, how do you build a better AI drug discovery capability?” suggests Robin Roehm, CEO of Apheris.

 

In this exclusive roundtable debate with Scientific Computing World, he said: “You also need to look at how quickly you generate new data and enrich these respective models, how you might customise and combine them with other models in your ecosystem, in order to move molecules forward. Companies might have very different needs here. A small company might want to have models at their fingertips to be able to query them, whereas a large pharma company might have a model catalogue and various groups that want to play with these models.”

 

He adds: “Then there is the question of how much superiority you ultimately get out of this. That is a function of what data is being contributed, what is the quality of the data and what’s your training paradigm as a machine learning project. We have released results with the structural biology network and have shown that federated models outperform all alternatives significantly, which is a huge achievement for co-folding models where so much investment is going into that. There are lots of providers building on the same publicly available data set and we have delivered to the partners the best model, significantly outperforming every other model that they could build internally or any public model.

 

“That is a function of not just federated learning, but how much data has been contributed. Depending on what modelling workflows you use, you can also use more noisy data to inform certain modelling setups. For example, you might generate synthetic data from more noisy data and feed that back into the training. If you get two orders of magnitude more data with that, that goes some way to addressing the data harmonisation issue compared to using raw data.
 

Measures of success
 

“We’ve also talked about benchmarking – how do you measure success? Ultimately, pharma companies are looking only to improve their own drug programmes, so that has to be the measure for most.

 

“What you want to compare is the federated model against a model that you could generate on your own data or against a public model only, and then apply it to your use case.”

 

John Androsavich, General Manager, Ginkgo Datapoints, Ginkgo BioWorks, picks up on the measures of success. “One metric is whether or not scientists use it,” he observes. “You can’t just have a good model on paper. It needs to be practical and scientists need to adopt it to increase the efficiency of their workflows on the day-to-day.”

 

Niña Cortina, Co-Founder of LiVeritas Biosciences, voices a view on behalf of smaller pharma players. “A concern for all small companies is that these federations feel like exclusive clubs for large pharma companies,” she says. “We are democratising access to analytical technology for small and mid-sized biopharma, and it would be a missed opportunity if federated learning reproduced the same access imbalances that have historically disadvantaged smaller players in the industry.”

 

These issues – and many more – are explored in depth in The Path to AI Federated Learning for Drug Discovery, our exclusive white paper, produced in partnership with Revvity Signals, featuring insights from leaders at companies such as LiVeritas Biosciences, Eli Lilly and Company, and AstraZeneca. Whether you're working in pharmaceutical R&D, biotechnology, AI development or scientific data management, the discussion offers valuable perspectives on where federated learning is today – and where it is heading next.

Download the free white paper here to discover how federated learning is transforming collaborative AI and accelerating the future of drug discovery.

Topics

Media Partners