Cambridge Healthtech Institute's 5th Annual

Machine Learning for Protein Engineering

Streamlining Biologic Development

May 14 - 15, 2026 ALL TIMES EDT

The use of ML/AI tools and modeling is enabling a vastly different process of drug development that is optimized for efficiency and success at every step of the biologics pipeline. Cambridge Healthtech Institute's 5th annual Machine Learning for Protein Engineering conference at PEGS Boston will provide an overview of tools being used in the industry with use cases and real examples of large-scale validation and implementation successes, as well as tips for overcoming challenges. Newly evolved developability metrics will employ the best proxies of success and correlate their use with real results. This conference will cover key advancements in algorithm development, generative language models and evaluation, and data-driven decision-making, providing researchers with tools to enable prediction with greater accuracy and efficiency. These techniques promise to transform discovery, prediction, developability, simulation, and optimization of biologics.

Sunday, May 10

2:00 pmRecommended Pre-Conference Short Course

SC1: In silico and Machine Learning Tools for Antibody Design and Developability Predictions

*Separate registration required. See short course page for details.

Thursday, May 14

7:30 amRegistration Open

7:30 am

From Scientist to Start-Up: An Interactive Entrepreneurship Breakfast

PANEL MODERATOR:

Catharine Smith, Executive Director, Termeer Foundation

Join us for an interactive breakfast conversation on the journey from scientist to entrepreneur, featuring founder, CSO, CEO, and investor perspectives. Panelists will share how they navigated the leap from postdoc to scientist to startup leadership, from securing initial funding and building teams to cultivating networks of mentors and advisors. Please see Networking Events Page for details https://www.pegsummit.com/networking-events. Free to attend - sign up in advance on the registration page.

PANELISTS:

Natalie Galant, PhD, CEO, Paradox Immunotherapeutics

Luca Giani, Senior Principal, AV; CoFounder & CEO, Ilios Therapeutics

Noor Jailkhani, PhD, CEO & Co-Founder, Matrisome Bio

8:30 amTransition to Sessions

8:40 amOrganizer's Remarks

USE OF AI IN COMPLEX MODALITIES: MULTISPECIFICS AND NOVEL SCAFFOLDS

8:45 am

Chairperson's Remarks

Maria Wendt, PhD, Global Head (Vice President) of Digital and Biologics Strategy and Innovation, Large Molecule Research, Novel Modalities, Synthetic Biology and AI, Sanofi

8:50 am

Accurate Protein-Binder Design Using BindCraft

Lennart Nickel, Graduate Student, Biotechnology & Bioengineering, École Polytechnique Fédérale de Lausanne

Protein-protein interactions are fundamental to most biological processes but remain difficult to design due to their complex structural determinants. We present BindCraft, an open-source and automated platform for de novo protein binder design that achieves high-affinity binding without experimental optimization or prior structural information. BindCraft enables the generation of functional binders against a broad spectrum of targets, including receptors, allergens, and nucleases. By integrating state-of-the-art structure prediction and design modules, BindCraft advances the field toward a “one design one binder” paradigm, establishing a generalizable framework with broad implications for therapeutic, diagnostic, and biotechnological applications.

9:20 am

Designing Biochemical Function with Generative AI

Rohith Krishna, PhD, Postdoctoral Fellow, Computational Biology & Machine Learning, University of Washington

Deep learning has accelerated protein design, but most existing methods are restricted to generating protein backbone coordinates and often neglect interactions with other biomolecules. I will present the next generation of protein design methods that include side-chain coordinates for design of more complex biomolecular function. Finally, I will show a series of applications of these algorithms to design of experimentally characterized functional proteins.

9:50 am

Miniprotein Optimization through Efficient Models

James Bowman, PhD, CTO, AI Proteins

De novo designed miniproteins are a powerful therapeutic modality that can be engineered for high-affinity target binding, site-specific conjugation, stability, and rapid tissue distribution and clearance. Using generative AI, synthetic biology, and laboratory automation, we design, produce, and characterize thousands of 45–60 amino acid miniproteins. Machine learning optimization improves expression, stability, and developability, enabling discovery of potent therapeutic candidates, including an in vivo active TNFR1 antagonist.

10:20 am Closing the Loop: A Step-by-Step Look at Integrated Wet Lab—AI Antibody Discovery and Development

Shuji Sato, Vice President, Innovative Solutions, Business Development, MindWalk

This approach enables antibody discovery from sequence only, with traceable outputs across discovery stages. LensAI, powered by HYFT, applies multi-parametric analysis by integrating AI models with large-scale lab and in silico data. It supports triaging hits from in vivo, in vitro, and repertoire sequencing, performs clone clustering, and evaluates epitope and paratope features with developability, manufacturability, and immunogenicity predictors for sequence-level mapping toward engineered preclinical leads.

10:35 am Multimodal Fusion of Empirical Structural Data and Deep Learning for Improved Modeling of Antibody–Antigen Complexes

Dan Benjamin, CTO & CoFounder, R&D, Immuto Scientific

We present a multimodal AI framework that fuses radical footprinting data with deep learning to more accurately model antibody–antigen complexes, including disordered and hypervariable interfaces. Using Immuto’s high-throughput radical footprinting platform, large-scale structural datasets, comprising hundreds of protein complexes, are generated to train AI systems, providing solution-state, dynamic data across all protein classes, including multipass transmembrane proteins. These experimental constraints enable advanced, actionable models for antibody engineering, affinity maturation, and specificity design.

10:50 amCoffee Break in the Exhibit Hall with Poster Viewing

ENTREPRENEUR MEET-UP

Fostering Entrepreneurship and Models for Start-Ups

Natalie Galant, PhD, CEO, Paradox Immunotherapeutics

Catharine Smith, Executive Director, Termeer Foundation

Are you a founder or aspiring founder? Are you an academic entrepreneur? Join Natalie and Catharine and PEGS attendee founders and entrepreneurs for networking and discussion. We will discuss existing resources for academic entrepreneurs, founders, and start-up leaders, and areas where the ecosystem can better support you.

PLENARY FIRESIDE CHAT

11:35 am

Plenary Fireside Chat Introduction

Eric Smith, PhD, Vice President, Bispecific Antibodies, Regeneron Pharmaceuticals, Inc.

11:40 am PANEL DISCUSSION:

How to Think about Designing Smart Biologics in the Age of GenAI: Integrating Biology, Technology, and Experience

PANEL MODERATOR:

Christopher J. Langmead, PhD, AI-Driven Molecular Design, Danaher Corporation

Artificial intelligence and machine learning are reshaping how we design, optimize, and understand biologics,  from sequence generation and developability prediction to in silico screening and automated lab validation. Yet, turning AI’s promise into real-world discovery impact requires new ways of thinking about data, infrastructure, and collaboration across disciplines and organizations. In this fireside chat, leaders from across industry and academia will discuss how AI is changing the landscape of biologics discovery, what challenges still slow adoption, and how teams are reimagining the interface between computation and experiment. 

The conversation will explore:

  • How AI is accelerating early discovery and molecular design for biologics
  • Emerging strategies for integrating experimental data and large language models
  • The challenges of data quality, interoperability, and interpretability
  • The evolving roles of scientists, data, and automation in the next generation of discovery labs​
PANELISTS:

Surge Biswas, PhD, Founder & CEO, Nabla Bio, Inc.

Rebecca Croasdale-Wood, PhD, Senior Director, Augmented Biologics Discovery & Design, Biologics Engineering, Oncology, AstraZeneca

Joshua Meier, Co-Founder & CEO, Chai Discovery

Maria Wendt, PhD, Global Head (Vice President) of Digital and Biologics Strategy and Innovation, Large Molecule Research, Novel Modalities, Synthetic Biology and AI, Sanofi

12:35 pmNetworking Luncheon in the Exhibit Hall and Last Chance for Poster Viewing

DEVELOPABILITY AT-SCALE

2:05 pm

Chairperson's Remarks

M. Frank Erasmus, PhD, Head, Bioinformatics, Specifica, an IQVIA business

2:10 pm

Predicting sdAb Biophysical and Developability Properties

Andrei Kamenski, PhD, Senior Data Scientist, InSilico Biologics Discovery, Novo Nordisk R&D UK

Single-domain antibodies (sdAbs), such as camelid VHH, are a strong focus area in biologics discovery due to their small size and modularity. However, the link between sdAb sequence, structure, and developability is poorly understood. To bridge this gap, we trained generalisable machine learning models on custom early-stage developability datasets, predicting key properties such as thermostability, hydrophobicity, and polyreactivity. In this talk, I will share our latest modeling insights and provide a glimpse into how the models are applied in our drug discovery workflows.

2:40 pm

Application of AI to Developability Screening, a Skeptic's View

Andrew C.R. Martin, DPhil, Emeritus Professor of Bioinformatics and Computational Biology, University College London

While AI has been used in bioinformatics since the early 1990s for problems such as protein secondary structure prediction, advances in AI over the last 5 years have revolutionized many areas of life from animation to bioinformatics. These changes have been driven by approaches such as protein language models and generative models used in AlphaFold for protein structure prediction. There have been several publications that use such approaches for 'ab initio' antibody design, but I for one remain skeptical. Nonetheless, there are clear applications for modern AI techniques around antibody developability and improving candidate antibody-based drugs.

3:10 pm

TherAbDesign: Bridging AI and Biophysics for Antibody Developability Optimization

Amy Wang, PhD, Senior ML Scientist, Prescient Design, Genentech

Antibodies are promising protein therapeutics, but successful development requires meeting strict developability criteria. We present TherAbDesign, a machine learning method that evaluates and optimizes antibodies based on sequence alone, proposing modifications that mimic the biophysical properties of successful therapeutics. This approach circumvents computationally expensive structure prediction and physics-based calculations. We show that this method improves known developability liabilities, such as viscosity, without explicitly modeling their mechanism of action.

3:40 pm Scalable, Robust and Easy to Use Biologics Discovery & Development with Generative AI

Eli Bixby, Co Founder & Head of ML, ML Research & Engineering, Cradle

Generative AI is redefining what’s possible in biologics discovery — enabling scientists to explore protein sequence space more efficiently, design higher-performing molecules, and accelerate the path from concept to candidate. At Cradle, we are developing intuitive software tools that allow protein engineers to directly leverage generative AI models within their existing workflows. This presentation will showcase how Cradle’s technology enables scalable and reproducible protein design, combining advanced AI architectures with user-centered design to reduce cycle times and ultimately create better biologics.

4:10 pmNetworking Refreshment Break

INNOVATION SHOWCASE

4:40 pm

Elise Pepermans, CEO and Co-Founder, ImmuneSpec

4:45 pm

Sunny Sharma, Founding Team, Fovus

DEVELOPABILITY AT-SCALE (CONT.)

4:50 pm

Designing Optimal Proteins at Scale

Jeliazko R Jeliazkov, PhD, Lead Protein Design Scientist, Machine Learning, Profluent Bio

ML-based protein design routinely achieves functional success, but whether models can generate variants that are simultaneously optimal across many properties remains an open question. We present alignment of foundational pLMs as a solution for multi-parameter protein optimization, with applications spanning gene editors to antibodies.

5:10 pm

Benchmarking Language Models for Antibody and Nanobody Tasks

Koji Tsuda, PhD, Professor, Computational Biology & Medical Sciences, University of Tokyo

Recent advances in protein language models (PLMs) have demonstrated strong performance on structure and function prediction. To evaluate their performance in nanobody-related tasks, we developed a comprehensive benchmark suite, NbBench. Benchmarking of eleven models revealed that antibody language models excel in antigen-related tasks, while thermostability and affinity-related talks remain challenging across all models. We further discuss how PLMs and their benchmarks could impact on antibody research.

5:40 pm PANEL DISCUSSION:

Are In Silico Tools Truly Reducing Clinical Failure and Accelerating Development?

PANEL MODERATOR:

M. Frank Erasmus, PhD, Head, Bioinformatics, Specifica, an IQVIA business

  • Validity of Proxies: Are we over-relying on proxies (e.g., AC-SINS, HIC) that lack standardization and correlate poorly at high concentrations? 
  • Manufacturing vs. Efficacy: By optimizing for algorithmic "safety," are we discarding potent binders to prioritize ease of manufacturing over patient benefit?
  • Generative Bias: As we shift to generative design, are we simply baking historical biases into new models?
  • The Negative Data Gap: Can we build robust "early warning" models if we fail to characterize and train systems on negative data (clinical failures)?
  • False Positives: Are we discarding safe, viable molecules based on theoretical binding risks that never trigger an actual immune response? 
  • The MHC Limitation: Current standards predict MHC binding, but binding does not equal T-cell activation
  • Predicting Tolerance: Can current tools evolve to reliably predict functional tolerance and activation, rather than just binding affinity?
PANELISTS:

Hunter Elliott, PhD, Vice President, AI/ML, BigHat Biosciences

Sandeep Kumar, PhD, Senior Vice President, Digital Biologics, Natural Antibody

Andrei Kamenski, PhD, Senior Data Scientist, InSilico Biologics Discovery, Novo Nordisk R&D UK

6:10 pmClose of Day

Friday, May 15

7:15 amRegistration Open

INTERACTIVE ROUNDTABLE DISCUSSIONS

7:30 amInteractive Roundtable Discussions with Continental Breakfast

Interactive Roundtable Discussions are informal, moderated discussions, allowing participants to exchange ideas and experiences and develop future collaborations around a focused topic. Each discussion will be led by a facilitator who keeps the discussion on track and the group engaged. To get the most out of this format, please come prepared to share examples from your work, be a part of a collective, problem-solving session, and participate in active idea sharing. Please visit the Interactive Roundtable Discussions page on the conference website for a complete listing of topics and descriptions.

TABLE 1:

The AIntibody Challenge: Inaugural Results and What's New in Challenge 2

Andrew R.M. Bradbury, MD, PhD, CSO, Specifica, an IQVIA business

M. Frank Erasmus, PhD, Head, Bioinformatics, Specifica, an IQVIA business

  • Review final results from the initial AIntibody Challenge, which included >30 teams across pharma, biotech, academia and AI companies. The final manuscript will be submitted in February and published later in Nature Biotech 
  • Evaluate target affinity, developability (minimum score), and submission time
  • Discuss plans for the second challenge​
TABLE 2:

From Binding to Application: What Will It Take for De Novo Binders to Succeed?

Lennart Nickel, Graduate Student, Biotechnology & Bioengineering, École Polytechnique Fédérale de Lausanne

  • Are de novo–designed miniproteins moving beyond academic proof-of-concept toward robust, reproducible therapeutic platforms?
  • What technical, biological, and manufacturing hurdles must still be addressed for de novo formats to truly rival antibody-derived scaffolds?
  • In which applications do antibodies and miniproteins naturally coexist, and where do miniproteins provide superior performance or design freedom?
  • What do we know and how much can we predict the immunogenicity of de novo protein binders?​

LAB-IN-THE-LOOP

8:25 am

Chairperson's Remarks

Victor Greiff, PhD, Professor, University of Oslo; CTO, IMPRINT

8:30 am

Better Antibodies Engineered with a GLIMPSE of Human Data

Lance Hepler, PhD, Co-Founder, R&D, Infinimmune Inc.

Infinimmune presents GLIMPSE, an antibody language model trained on proprietary human data that achieves state-of-the-art performance. We used GLIMPSE within our lab-in-a-loop platform to engineer an anti-IL-13 antibody, enhancing its drug-like properties including potency, extended half-life, affinity, stability, and manufacturability via liability removal. This work demonstrates the practical application of language models for optimizing therapeutics while maintaining their humanness, moving beyond typical proof-of-concept studies.

9:00 am

Training Data Composition Determines Machine-Learning Generalization and Biological Rule Discovery

Victor Greiff, PhD, Professor, University of Oslo; CTO, IMPRINT

Supervised machine learning in antibody discovery relies on positive and negative examples, making dataset composition crucial for performance and bias. We evaluated how different negative-class definitions affect generalization and rule discovery in antibody-antigen binding using synthetic structure-based data. Models trained with negatives more similar to positives had reduced in-distribution performance but markedly better out-of-distribution generalization. Ground-truth analyses revealed that inferred binding rules shift with negative set choice, and experimental validation confirmed these findings, emphasizing dataset design for robust, biologically meaningful models.

9:30 am

Computational Design of Antibody Repertoires

Ariel Tennenhouse, Graduate Student, Biomolecular Sciences, Weizmann Institute of Science

We are developing a new strategy for designing repertoires of billions of structurally diverse and stable human antibodies. I will first describe two methods we developed for atomistic antibody design that enable this strategy and show that each method can optimize antibodies across a variety of criteria without prior mutational data. This shows that optimizing native-state energy is an excellent first approach for antibody optimisation. I will then describe a proof-of-concept universal repertoire of 500 million variants we designed and show we can reliably select highly developable and reasonably high-affinity antibodies against diverse targets.

10:00 am Lab-in-the-Loop AI Drug Discovery: De novo Antibody Design and Benchmarking with Amazon Bio Discovery

Speaker to be Announced, Amazon Web Svcs

Jiwon Kim, Amazon Web Services

AI is compressing preclinical drug discovery timelines, but fragmented tools and limited benchmarking data remain barriers. We present Amazon Bio Discovery, a unified application for lab-in-the-loop AI drug design, and a new antibody developability benchmark dataset spanning 5,000 sequences across 50 targets and 7 assays. Demoing a case study of agent-guided de novo nanobody design against a novel cancer target with Memorial Sloan Kettering Cancer Center, we show how combining AI-driven candidate generation, in silico ranking, and wet-lab validation produces sub-nanomolar binders without prior antibody data.

10:30 amNetworking Coffee Break

DE NOVO BIOLOGICS DESIGN: USING AI TO CREATE BRAND-NEW ANTIBODIES AND PROTEINS FROM SCRATCH

10:44 am

Chairperson's Remarks

Surge Biswas, PhD, Founder & CEO, Nabla Bio, Inc.

10:45 am

From Proof-of-Concept to Proof-of-Productivity and Scale

Prashanth Vishwanath (PV), PhD, Director, AI & ML, Computer Science & Data Strategy, Takeda Pharmaceutical Co Ltd

Proof-of-concept has been demonstrated showing how AI methods can be used to design and optimize large molecules. We must shift our focus to scaling to maximize the productivity and innovation gains. This talk will cover a selection of PoCs and then how we are scaling digital biologics at Takeda across our portfolio.

10:55 am

Massively Multiplexed in vivo Screening of AI-Designed Proteins Enables Programmable Tissue Targeting

Pierce J. Ogden, PhD, Co-Founder & CSO, Manifold Biotechnologies Inc.

At Manifold Bio, we’ve built a direct-to-vivo platform that connects AI-driven protein design to functional data from living systems. Using this approach, we generate and evaluate thousands of designed binders to novel targets simultaneously in vivo. This massively multiplexed framework has yielded functional brain shuttles capable of crossing the blood–brain barrier, and we are now extending it to other tissues to enable selective delivery of diverse therapeutics. By integrating AI design and in vivo multiplex readouts, we aim to generalize tissue targeting across the entire body.

11:05 am

Push-Button Biologics Design

Surge Biswas, PhD, Founder & CEO, Nabla Bio, Inc.

We recently announced JAM-2, which can design antibodies with drug quality properties with high success rates. We'll discuss these results, and also share examples of what successful deployment on real drug discovery programs partnered with large pharma looks like. We'll discuss roadblocks and share practical lessons/advice for how to build teams and infrastructure to ensure AI driven biologics discovery delivers real drugs not just headlines.

11:15 am PANEL DISCUSSION:

De novo Biologics Design: Using AI to Create Brand-New Antibodies and Proteins from Scratch

PANEL MODERATOR:

Surge Biswas, PhD, Founder & CEO, Nabla Bio, Inc.

PANELISTS:

Prashanth Vishwanath (PV), PhD, Director, AI & ML, Computer Science & Data Strategy, Takeda Pharmaceutical Co Ltd

Pierce J. Ogden, PhD, Co-Founder & CSO, Manifold Biotechnologies Inc.

Maria Wendt, PhD, Global Head (Vice President) of Digital and Biologics Strategy and Innovation, Large Molecule Research, Novel Modalities, Synthetic Biology and AI, Sanofi

12:15 pmClose of Summit





No Agenda API URL configured.

Register

View By:


Premier Sponsors

       Integral-Molecular_NEW  Kactus     ThermoFisher_Red