Affiliation:
1. Oxford Protein Informatics Group, Department of Statistics, University of Oxford , 24-29 St Giles’, Oxford OX1 3LB, UK
2. Large Molecule Research, Roche Pharma Research and Early Development, Roche Innovation Center Munich , DE-82377 Penzberg, Germany
Abstract
Abstract
Antibodies are key proteins of the adaptive immune system, and there exists a large body of academic literature and patents dedicated to their study and concomitant conversion into therapeutics, diagnostics, or reagents. These documents often contain extensive functional characterisations of the sets of antibodies they describe. However, leveraging these heterogeneous reports, for example to offer insights into the properties of query antibodies of interest, is currently challenging as there is no central repository through which this wide corpus can be mined by sequence or structure. Here, we present PLAbDab (the Patent and Literature Antibody Database), a self-updating repository containing over 150,000 paired antibody sequences and 3D structural models, of which over 65 000 are unique. We describe the methods used to extract, filter, pair, and model the antibodies in PLAbDab, and showcase how PLAbDab can be searched by sequence, structure, or keyword. PLAbDab uses include annotating query antibodies with potential antigen information from similar entries, analysing structural models of existing antibodies to identify modifications that could improve their properties, and facilitating the compilation of bespoke datasets of antibody sequences/structures that bind to a specific antigen. PLAbDab is freely available via Github (https://github.com/oxpig/PLAbDab) and as a searchable webserver (https://opig.stats.ox.ac.uk/webapps/plabdab/).
Funder
F. Hoffmann-La Roche
GlaxoSmithKline
Engineering and Physical Sciences Research Council
University of Oxford
Publisher
Oxford University Press (OUP)