Delivering Insights to Accelerate Protein Engineering
Good afternoon, and thank you for being here today. There's been a number of good talks already this week that have discussed expanding the use of antibodies beyond therapeutics and thinking about how they can be used as delivery vehicles. Today, I will continue that narrative and discuss how DOPMADX LUMA can be a partner in those efforts. Just as a way of a quick agenda, we're gonna go through four parts. The first is the discovery journey, what goes into it. We're also gonna talk a little bit about designing for the practical and what we can actually do to support that. Talk a little more about letting data drive timely decisions and what that actually means and future looking, what enhancements do we see coming to help research? This is just a quick safe harbor statement. So when I started in the industry about thirty years ago, we were just beginning to understand the potential of antibody therapeutics. In that time, the advances in molecular biology and biochemistry have really pushed our understanding of antibodies, but also immunoglobulins in general. Think of all the talks we've heard about CAR Ts. Additionally, we've learned how to produce antibodies and antibody products at industrial scale, so that now we can ask: What is the best way? And not: What is possible? Right now, our understanding of antibodies, multi specific antibodies, and antibody delivery, we can look at treating diseases in completely novel ways. Take for instance, the scenario I'm proposing today. We know that the formation of amyloid plaque correlates with tau protein tangles and which leads to disease progression in Alzheimer's. What if we found that modifications in CYP40, a molecular chaperone, or HTRA, a serine protease, could not only slow but reduce the formation of tau protein tangles. How would we deliver them? What we have first is a challenge with the blood brain barrier. We have to get across that in order to deliver our enzymes. We need a strategy you could be to use anti transparent antibody or antibody fragment to mediate that delivery. But is an antibody the right format? We can actually start to answer these questions today. As someone whose family has been impacted by Alzheimer's, more potent treatments are needed. Right now, when we look at our benchmark for success, it's not whether you're in academia or an industry, it's how do we move from idea to innovation as rapidly as possible. Success. So success as a former scientist and as a business analyst, I'm gonna ask the question, what are our blockers? And it's an old story, a universal one. Whether you're an industry or academia, and whether or not you're looking at this diagram and seeing yourself spread across multiple sites, across multiple continents, or thinking about how you're collaborating within your university system, it is pulling all of our data together has become extremely difficult. And now in the age of AI, it's becoming a bigger and bigger challenge because we are using AI as a part of our daily work. This situation that we're looking at here not only slows down our us as human scientists, but it it's curbing how we use one of our most powerful tools. And while this is an old story, it's one that I think we are now at a point where we can really address and need to address. And what we really need is a way to traverse the nodes of our data and then look at it in multiple dimensions. This would power not only us as a human scientist, but it would it would help, our AI and ML tools interrogate the data. And that's exactly what Luma does. And as an AI native platform, we can begin to realize the full value of all of our connected data. To do this, LUMA brings together design, process, and results. So the goal here is to enable more rapid iteration cycles powered by experimental datas and all aspects of our research. Some of you may have heard that this referred to as lab in a loop, and that's exactly what LUMA empowers. But what does that really mean, for the scientists? It means better collaboration, planning tools, optimizing the execution of our experiments. And when we start to think in terms of AI, what this can help us with is explore more rapidly the insights that we are developing and also begin to think in terms of predictive modeling. So why is this important? Let's discuss the scenario I proposed. When we when we are doing our work, we want to work in this connected ecosystem. And what that really means is that data from each step is leveraged across the process so that we can maintain and build on the knowledge that we're generating. When we start looking at this, it begins with our design panel. We wanna make sure that when we're building this, that our scientific objectives are clear, and we wanna make sure that we can achieve them. When we're starting to build these more and more complex formats, the question becomes, can we actually build them and make them at scale? If we can grow them up, the question then becomes, can we purify them? Can we isolate them? So that we can use them to test potency and make sure that they are still as potent as we change and move through different formats. And ultimately, we wanna use this information to produce decisions that lead to us either to we are on the right track or we need to continue to iterate. And to make this all possible, it starts with a very accurate scientific data model that represents the molecules that we're working with. And the underlying foundation of this concept is building blocks. What this allows us to do is to break down immunoglobulins into composable units. This enables a scalable way to iterate on design, but in a way that structures the data for both ease of experimental workflow development and exploration in the future. It also enables, as a scientist, an intuitive way to look at the molecule. And I think anyone who's been working in the field long enough, when we look at a molecule like this, we can very quickly realize that we are looking at a bispecific with an SCFV off the FC with with with an FV off the FC. And this is important part of the design of these molecules. And most important here is that this is the data model that underlies everything I'm gonna talk about. So moving ahead into the discovery part. We know let's talk a little bit about how we iterate through this. We know we have a molecule that we want, an anti transferrin antibody, that we wanted to use to deliver these two potential enzymes. What could we do? We could design a single antibody with both enzymes that might potentially be very potent because you have both enzymes present. However, we know we have a scientific challenge with the blood brain barrier. So in looking at that challenge, we realized that having a molecule of this size may not be effective. We may want to make it smaller. So removing one enzyme might be will will achieve that goal. But is that enough? Well, what if we also have a challenge with clustering the transferrin receptor? So as a result, we may want to remove one of the fabs. But if we do that, we can actually gain even more reduction in size by moving to an SCFV. So as we start to walk through this, we start to build a plan. And that plan says, we begin with this multi specific antibody molecule, but we're gonna iterate down to reduce the size reduce the molecular size. But in addition to to optimizing for that one feature, we also start to ask ourselves, can we leverage sequence and structure to improve the developability and remove liabilities during manufacturing? And that's where the strength of Bioglyph really starts to shine. And so let's take a look at Bioglyph. When we're looking at Bioglyph, we noticed when we enter the palette is that we have this build these building blocks over here on the left hand side. And what this does, again, it is the underlying data model. So we're not only talking about drawing molecules and pictures, but we're actually talking about a data model that establishes the connection between the different composable units of an antibody fragment. We can also, over here, start beginning to assign targets to our to our molecule. What this allows us to do very rapidly is see which parts of the molecule are binding to which and which are the parts of the molecule that are adding enzymatic functionality. Moreover, what we can also do on the same design palette is begin to iterate through our scientific process that we had discussed earlier. So now in real time, whether I'm collaborating in my office or whether we're we're on on a Teams call or whether in a conference room, you can very quickly, very rapidly iterate through all the potential scientific options for the molecule you want to develop. This is extremely powerful because what it allows us to do as scientists is have conversations about are we headed in the right direction? Are we using the right formats? Have we tried anything like this in the past? Do we wanna go even smaller? And as much as we we might wanna think about being successful the first time out, what is our next step if none of these work? Because this is a visual tool, we can very quickly and rapidly start to ask not only questions about the molecules themselves, but scientifically, are we hitting in the right direction? And I think on top of the data model that the building blocks provide, this visual tool provides a very easy, extensible collaboration tool to scientists in the lab. Moving forward, let's say, for argument's sake, we found our panel. We're very happy with where we're at. We can start to actually iterate on this and build out the specific pieces. We know the models. Here they are. Here are our designs. Now I wanna make this real. And to do that, I'm gonna start assigning fabs. I'm gonna start signing my SCFVs. I'm gonna start looking at what FCs do I wanna use and how many iterations of the enzymes. Right? Because I may have multiple iterations there. When you start to look at this panel, you can start to see how rapidly we are increasing the number of molecules that we wanna produce. This brings us to another really interesting point of that in addition to the in addition to the data model that we can do with Bioglyph. And that is we begin to start looking at from a real practical way, what are we designing? What might not be easy to see is that you have one hundred and forty four molecules with just what we discussed before just about twenty four of the of these fabs a few SCFVs and a few combinations. We are getting to a hundred and forty four molecules. And what this allows us to do before we even begin making anything is to have a very honest conversation about capacity. So again, thinking about what Bioglyph does, and I'm gonna show this later in the presentation, it is really a scientific tool that allows us to iterate through looking at two, at three-dimensional designs of proteins, and I'm gonna show that. But it also is a collaboration tool that allows us to start talking about things like capacity. So as a as a molecular biologist, as someone whose lab might be requested these hundred and forty four molecules, If I'm sitting there at the table when we're having this conversation, we can actually look and say, can we make that? And we can have those hard conversations before we start committing to doing the work so that we can make sure that we are meeting the goals not only of this project, but all of the projects that we have. And so what I wanna make sure I'm highlighting here because as in this part of the presentation, I'm presenting how we can leverage Bioglyph as a collaboration tool. There are other parts of the tool that we'll go through that actually do molecular design, and I'm gonna hit on those later in the in the presentation. But I don't wanna lose sight of the fact that this is still a scientific tool, and we will we'll talk more about the data models going forward. But now what I wanna do is move to we've created our panels. We have this underlying data model. We're moving this ahead into the actual part of process of building the molecules. And that's where we really begin to see how LUMA can move from design into execution. Now I'm not gonna be able to show everything today, but one of the key things I wanna highlight here is that when we track Moomah, we are actually creating deep scientific relationships and lineages so that we're not only mapping out the data and the metadata, but we're also tracking the process and the results with high degrees of precision. And how do we do that? When we look at how we integrate Luma with LabConnect and adaptive workflows, which again, I won't be going through the workflow steps today in great detail, what we are gonna bring together are the experimental and process data with our design data to build an operational picture. And so what this provides is a necessary tool to move very rapidly from that, as I said before, from the idea through knowledge and to being able to actually make decisions very rapidly. And so what does that look like? Again, looking at the connective ecosystem and, again, every experience is customizable, but one way to look at it would be the following. When we land on our page, get some very high level information about our project and a and a quick project summary. We see things like the project name. We see the number of constructs that we've been building and we can see some of the quick metadata around the whole project as a whole. Something to think about very quickly here is that when we look at the number of constructs that we've already made, we would have made about a hundred and twenty. I remember we were talking about building about a hundred and forty four constructs. What we see already is that we have about three percent that aren't expressing. And so what this gives us the ability to do wrap for the is let's drill into that data and see what is there any correlation? Is this just random, or or is this something that we might wanna use before we continue with the remainder of the of the constructs? So we can drill down into our expression tighter tighter data. And what we see here is that we're actually doing not one, but three three pilot runs. And the data across the pilot runs are pretty consistent with run three maybe showing the most variability among the other two. If we wanna take this a step further, we can look at our data and compare it with with properties like aggregation and purity and seeing how are they lining up. And we can see they're trending very well with a few outliers. But more importantly, we can also think about how do we use AI when it comes to interrogating our data. Now I wanna I'm gonna pause this for a moment because I'm gonna come back to this. But what I wanna share is is is that this is high level project data. What if we wanna drill into more specifics? What if we wanna see how our designs compare at a very specific level? So on any of these points, we can kind of drill down and get more extensive data about our our molecules. So as you can see over here on the left, we drill down, we can start to see the formats that we're looking at that are in any one of these peaks. We can also look at a high level to see what is our quality score. Again, we can configure this on this dashboard. This view is a culmination of things like average titer, purity, and aggregation, And we are getting this on a molecule by molecule basis. In addition to what we're doing for the production of of this data, we can also look at assay data like kinetics and look at binding data. We can do this across species, and all these views are brought in from that initial view from the summary data as we do a drill in and is comparing here our design to the data we're actually getting. And what I wanted to call attention to again here is AgentLuma and why this is important. And I'm gonna drill down a little further now on this. As many of us know when you're running a project, we oftentimes think we understand what are gonna be the key data elements that are gonna make our decision. And so when we build our our dashboards, we build them with the intention of being able to interrogate that information. But but most of us know is that as we are as the project evolves, there's gonna be new information that comes in that might trigger different decisions. With agent Luma, even if it's not built into the dashboard, we can very quickly ask questions of the data that's not that's both here to ask it to help us sort through many, rows, but we can also ask it to bring in new data that wasn't necessarily thought of as the project begins. And this is extremely powerful, again, because in most cases, we don't always know what is gonna be that deciding data factor. And how is that possible? As I said before, and, again, apologies for not having time to go into it in more detail. When we are working in Luma, we are building deep relationships with our data. So in this case here, what what I'm trying to show is that for one protein example, we have looked at how we can express it in multiple different cells under multiple different conditions and create multiple batches. All of this information is available to agent LUMA. So irrespective of what we started with on the dashboard, knowing that we have this relationship to those proteins, we can very quickly navigate through the multiple different nodes that we maybe didn't anticipate initially and be able to use that data through agent Luma to quickly add more information to what we had originally thought would be important. And I think this is really key. It is it is in in addition to this tool to be able to use and make operations. What's underlying that is this deep data model that is provided through Luma. So this is great, and I think what I wanted to share with you today is how we're looking at moving the design into the space where we're actually comparing it with our data, leveraging our data model, and our future tools of AI to be able to look at across multiple dimensions of data. And that's kind of always been the promise. But what do we wanna do in the future? And what we will what we're thinking about doing is moving the design to the earliest phases of analysis, and how might that look? And this is at the concept stage and something that we're working on as part of the road map. But what we wanna do is thinking about tools like protein metrics that the scientists are using to evaluate mass spec data at the earliest phases of analysis. We wanna be able to share with that scientist the actual data the the actual molecules and and and essentially the a glyph of the molecule that they're making. This really helps facilitate the analysis at at the bench because the scientist is able to see what it is that we're potentially trying to make. More importantly, we'd also like to be able to leverage the deep again, as I said before, scientific parts of a bioglyph to say, not only do we want to be able to take that data and share the glyph itself, we want to be able to leverage the important information about what will be our potential peaks and how it correlates in a mass spec analysis. And not just the primary peak, but we also wanna think about what are potential off peaks and be able to make sure that those are shared with the scientists as well. This helps facilitate the decision making when you're looking at the peaks at at the time analysis is being done. It makes the the analysis not only potentially more rapid, but it actually makes it a deeper and more rich experience because you can be looking not only at your scientific data that's coming out the instrument, but what you're actually predicting. And then to bring this full circle, we would envision that now once this analysis is done, it is shared in Bioglyph or in Biosphere with other with other data that is being generated for the production of these molecules. And what this allows us to start doing is it allows us to start looking at the molecules, the designs, in coordination with the different molecular properties that we are measuring in the labs and start to make decisions that are that allow us to have a more coordinated way of asking where are where do liabilities exist? And if the molecule is performing particularly well, how do we actually use that information to then move deeper into the bioglyph tool so we can say let's bring together sequence structure and function and be able to say knowing that this molecule has a high degree of value and knowing that I have this sequence, where does it actually reside in the molecule? So we can start to ask, is that a real liability? If it's a if it's a hydrophobic residue and it's buried in a pocket, we might say that's not as much of an issue. Whereas something that might be surface exposed might be something we wanna actually change or alter. And this is what this basically is doing is enabling that lab in the loop. We have knowledge that we're gaining at the bench. We're we're taking that back into the into the design phase and then essentially pushing it back out. And if we again, as I showed with the dashboarding, we can take these insights that we're learning very rapidly and move them ahead. And and, essentially, at the moment we're gaining those insights, use them to start impacting our overall design and, implementation. And so that's important information when you're looking at how everything you're doing moving forward. But what about when you're actually looking at your legacy data? The data that we that we as scientists have been learning, as I said, over the last thirty years. And that's a very important new feature that's being released in Bioglyph is this notion of format rescue. And what we can do here now with format rescue is we can actually import sequences, go through a validation step where we align and and look at the sequences, and then detect the format. And so the question that that comes out of this is that, well, this looks great, but how well does it actually work? And we've been actually able to do this with one of our customers. And in taking over two hundred thousand sequences, a hundred and ninety five thousand have been successfully annotated. That's roughly ninety percent. And for the pessimists out there like me, that says, well, that's great. But what about the rest? Well, the other twenty two thousand, we're able to partially annotate. And I think that's really important for two reasons. One is oftentimes when we start to think in terms of a curation effort, we start to think about what is it gonna take to curate that data that we've had for thirty years. We don't know where to begin. And so when you have something like format rescue, you can very quickly narrow it down to say, well, here's the delta. Now I know what the scope of my effort's gonna be. I know that I have these twenty two thousand that are partially annotated, and I want to go now into a full annotation. And so as a result, this gives you a scope for your data curation, and it gives you a way to start actually looking at it in terms of real again, going back to collaboration and having hard conversations. Do we have it in the budget? Is this a priority? And we can look at those sequences in terms of what was partially annotated and ask, are they relevant to what we're gonna be doing in the future? The other thing that it allows us to do, and this may be more important, and it is more important from a scientific perspective, is now, in many cases, we have had our data standing sort of, I would say, side by side. You have your historical data, and then you have your new data. And those are in two separate data schemas, and they're in different formats and in different data models. Now we can actually have through this process, we can actually merge our data in a way that we can now take our legacy data and be able to use it in the same way we are our new data and all the data moving forward. And, basically, what that allows us to do is allows us to apply all of the tools that I showed you earlier. The the the being able to look at two dimensional structures, being able to combine data, all that now, and then being able to extend the life of our of our legacy data to new projects and to new ideas and to be able to say, in the past, maybe we didn't have a modality for that. Now we do. And so now all this data is not standing side by side. It's actually living in a fully integrated environment. And that's and that's key for for all organizations. So kind of to wrap up, I wanna circle back to the beginning and when we talked about and I hope I've been able to share with you a little bit about our our plans with Luma and Lab in the Loop. But I wanna get back to what is Lab in the Loop actually give us? Lab in the Loop and and how is how is it actually built for Luma? It's built as an AI first technology. LUMA is built on bio on Databricks, and and the two go hand in hand to allow us to really use AI first and think in ways of making sure that our one of our newest and most powerful tools is sort of first and foremost in the way we we conduct our science. We do this by having a scientifically aware data model that tracks all our entities and processes and helps us track lineage throughout the entire process, which really allows us to bring all this data together even if we didn't know at the beginning we were gonna use it or how it will be used. And what this really empowers for the scientists is is first and foremost is collaboration. It allows scientists to have real communications about the data as it's being developed and so that we can use this throughout the entire process, not at the only at the end. It allows us to think about resource planning upfront, but also as the data's coming in. And what this kind of pushes us into the world of thinking about we can start to do is creating automated project updates and start to look at trends and projections as the data is being developed. And what this will future wise set us up for is how do we use our AI and ML models? How do we start thinking in terms of predictive modeling and potentially into simulation? And so to summarize, Dotmatics Lumen provides the foundation to accelerate the scientific discovery process. And I think it does that through allowing us to look at early insights that save time, resources, and accelerate our innovation. Also, it gives us the ability to leverage automated updates to enable collaboration between and across teams. And your data and your process are the foundation for your AI journey. So as you're going about doing your science, it's important that we are structuring it for that future and that we can leverage it. And, again, I think that's also important when you think about format rescue and how that can play a part in leveraging data moving forward. And I just wanna end with I think we've all heard that data data driven decision making, and it is and that it is super critical to what we do as scientists. But I think it's important to also note that timely data drives progress, and I think that's all what we're all moving towards. And thank you very much.
Antibodies are becoming powerful delivery vehicles, not just therapeutics. In this on-demand webinar, Jason Yasenchak, VP, Strategy for Multimodal Therapeutics at Dotmatics, uses a real-world example that delivers enzymes across the blood-brain barrier for Alzheimer's to show how the Dotmatics Luma platform connects design, process, and data with AI to speed up antibody discovery. He covers Bioglyph for molecular design, AI-powered dashboards, and how legacy data gets a second life with Format Rescue.
Our Latest on Science & Industry
Simplify your path to discovery.
See Luma in action by requesting a demo today.



