AI Alone will not Accelerate Science
Hi, my name is Christian Olsen. I'm the VP of Strategy for Protein Therapeutics at Dotmatics and I work primarily on our scientific platform, LUMA. Today's presentation, AI Alone Will Not Accelerate Science, a peek into how LUMA enables our scientists to leverage their rich biological scientific processes and AI at the same time. So brief agenda. I'll go through a little introduction. I'm going to ground this talk in an actual use case. I'll talk a bit about some wrong assumptions that we make. And then I'll spend the rest of the time talking about composability and why this is such a critical part of the Luma platform and what we're doing to enable innovating science. So this is our safe harbor statement. And really, the industry has been focused on trying to discover blockbuster drugs more quickly. So we're terming this faster molecule to market process where we've seen the drug discovery funnel here on the left in green, where we have thousands of candidates that we need to screen through. So this upstream research and development process is very intensive process wise, labor wise. The biology can be incredibly complex. But once you find some leads that you would like to test, then you have a whole other gauntlet of activities for development. And then when you make it through there, you've got the step for tepid transfer, and then you need to actually make the molecule so that you can get it into the patient's hands. And this process typically spans a decade to fifteen years. And so we have brought to our users a very powerful and flexible way of looking at antibody discovery. So I'm gonna talk about antibody formats and not just typical IgG molecules which are naturally found in our systems, but I'm going to talk about the non natural antibodies, those things that don't occur in nature. And so on the left I've got a PDB structure for a target, and on the right is a way that we render this multi specific antibody format. And this image on the right actually communicates quite a bit. And in the use case that I'd like to talk about, I'd like to talk about an Alzheimer's target where I need to actually have the molecule after I've designed it go across a physical barrier, the blood brain barrier, and I need it to carry an enzyme to treat A beta plaques. And so in this slide, really what this illustrates is I have a number of different formats of progressively smaller sizes, so if you'll focus on the very top arm of this slide, we see that I've got a typical IgG looking structure even though it's an FC building block on the bottom, and then the figure in the middle shows it's a little bit smaller. So I'm missing a fab building block. And then again smaller once you get to the far right. And for each of these arms, I'd like to test the same enzyme across these progressively smaller formats, but in both arms they're different enzymes. So how can I do this process with absolute precision, scientific data fidelity, as well as enabling our users to not have to go through lots of manual data joins and data uploads, manual file transfers, okay? So I'd like to go from a multi specific antibody. I'd like to optimize through this process of reducing the molecule size. And then quite likely, once I find some of those molecules of that size, after I've done my testing, I may wanna go through and I may want to send them to my protein engineering team to remove the liabilities and improve developability using both sequence data as well as structure data because we all know that structures function. And when we build these molecules, both IgG and multi specific, non natural, we take these molecules through a battery of tests. And this is just an example of some of the tests that you would perform on different parts of the molecule. For example, for the Fe regions, may want to use AI and in silico guided CDR optimization. You may want to then go through a process of affinity optimization all to the end so that you can increase potency or improve PKPD scores. At the very bottom, for the Fc building block, we may want to go through some activities of Fc engineering so that we can improve Fc binding, again, so that we can improve PKPD. So this is an example of the complexity of the assay steps or process steps that we need to go to to tune this molecule. And the way that we look at it, we look at it as when we build these molecules, when we design the structures, when we identify the sequences that are associated with the structures, there are annotations that are tied to the design goals. And so this is the way we look at the annotation design goal architecture where we have, you know, many goals. We have four goals here. We have affinity, humanization, developability, We have epitope. And then across the horizontal, across the top, we have different annotation types, different data inputs. Then we have scoring outputs for each of these design goals. And so this is what I'm calling this, just colloquially, a rationale record. This record has to be immutable per variant of the molecule per timestamp. And every annotation layer records the model's reasoning alongside its score. So the inputs that are used, the structural effects that are predicted, and which design goal is served, and then what the outcome is. So we track all of this. Now another way to look at this is to aggregate the former data into what I call a variant scoring panel, where I have four variants listed here and then I have all of the design goals listed horizontally across the top. So I have affinity scores, humanization scores, developability, epitope readouts. And so this is an example rationale record where we can see, for example, our our fourth variant, variant four, which is our lead. We have our affinity metric, right? And we can tell from our affinity data that if we mutate R94K, it removes the CDR H3 charge cluster. We can see the delta G as predicted by Rosetta. We can also incorporate humanization data developability and epitope so that we have a discrete record for this variant. And really this is what we call the innovation cycle. So for every loop around this innovation cycle, with each trip around the cycle, the pool of targets not only decreases, but you spend more time and the value of the data increases. So you may start out with hundreds of thousands of possibilities over the course of several days. As you go through your wet lab processes, you start to whittle down your candidates to less than a hundred, and this could take years. And so the value really increases dramatically as you spend more time focusing on these molecules. And so the question isn't whether AI makes biology faster. The question is whether your lab is structured to learn from every experiment so you can feed the learnings back into the next design cycle. So the goal isn't more experiments, though that's always nice. More data are always nice, but it's smarter experiments. And I think in this industry, we come to the table, we come into the lab with a number of assumptions, especially with the advent of AI. And so there's what we assume that AI just needs more data. And if you fix the data layer, your intelligence will emerge from that. And that's not necessarily true. The real insight, what's actually true is without model process and one hundred percent material fidelity, design rationale, context captured at every step, your data has no scaffold. And AI can only reason accurately about what is precisely defined. So you may have a Nobel laureate leading your program with all of this rich and amazing tacit knowledge in their brain, but if that knowledge is not structured and captured precisely and tied to model process and material fidelity, AI can't reason across what's locked up in that brilliant person's head. So even perfect data is useless without an explicit model of the scientific process that produced it. Without model process, the experiment's logic is implicit. It's not structured. The protocol steps live in the scientist's head or a Word doc or PDF perhaps that you shared at your lab meeting at the end of the week. And AI can't reason about a process and can't see. Without material fidelity, the molecule state, its lot, its lineage, and expression context aren't captured with the result. So a binding affinity number without its providence is not a data point. It's just a guess. And without decision traceability, you lose why a candidate was advanced or why it was dropped, why it was deprioritized. It's not recorded in structured form, so the institutional logic of your program is invisible to any downstream system, including AI. So the data may exist, but the structure and the connections often don't. And this is just an example of what I've run into quite a bit with all sorts of organizations from small startups to larger organizations where you have multiple teams. You have core groups. You have many different sites. And these data are siloed. We all know this reality. Silos are an artifact of our scientific and business processes. And what we're doing with LUMA is we're actually structuring all of the underlying data, all of the systems, all of the teams, all of the assays, all of the instruments so that AI has a has an environment that it can reason across. And really, that's what Luma is all about. And so just to briefly speak to this image over here to the right. So we we have many best of breed scientific tools that plug into LUMA like Genius Prime, Genius Biologics, Ohmic, GraphPad Prism. But this environment is an open environment, so using APIs, you can flow your data in and out of Luma. One of the most powerful things about this diagram is the plus sign at around the seven thirty or eight o'clock mark on the Albert orbit here. And this is meant to indicate that you can interact with Luma with your systems as you see fit, which means as your programs grow, perhaps you add a new modality next year, you add a different instrument, you add a whole new group, it's trivial to bring those new elements into your ecosystem. And so what I'm talking really about today is about two things. I'm talking about material representation, what was made, the thing in the tube, as well as the process representation, How was it made or tested? And that's really both of these you need to have in order to rationally and intelligently and efficiently do drug discovery. Without both, you don't have AI ready data. And without both, you have either the material or you have the process. Without one or the other, then you really don't have the total picture of context. Okay? So context is king. Every material, every fragment, every step and action, every transformation that your entity goes through has an address that's linked in a graph. Biology is essentially a graph set of challenges. Right? So the scientific intelligence platform, LUMA, is about bringing a number of things together. Okay, so we bring together a structured ELN, we integrate instruments, we include a registry, assay data are interleaved throughout this environment so that you can have an AI read data layer. So it's trustworthy, it's explainable, you have actionable recommendations. So every experiment is queryable. All of your materials are linked, and your provenance is preserved. So we go from static records to dynamic scientific models. So we're able to incorporate conditions and data and instruments, feedback. So really the experiments didn't change, right? The biology happens at the pace that biology happens. Assays take as long as assays take. What changed is the structure around it, and it's the structure that makes AI reasoning possible. So we go from disconnected, independent elements to real reasoning and insights so that we can make predictions, decisions. We can match all of those to our goals that we have in mind. So I'll talk a little bit about composability. And composability is really at the at the core of our philosophy where we have a number of discrete elements that individually are incredibly important, but in aggregate they all come together to do something, as illustrated here in this camera lens. So we're not talking about the lab of the future. The lab of the future, we know it requires structured science. We're talking about the lab of the now because patients can't wait. They need their therapies years ago. So we're talking about automation, AI scale, and speed. And so we're really talking about a common language for science. So we have model process language that of course is connected to human execution, this is the human in the loop, which is connected to automation, and then AI can take advantage of the aggregate of these three for both research, development, all the way down to manufacturing. So typically, biologics programs are running in an open loop model where AI generates a design, wet lab executes these designs, and the data are captured somewhere, but those data are rarely fed back into the next design cycle in a structured and automated way. What we're bringing forth with Luma is a closed loop R and D cycle. And we may start with AI assisted design where you can generate your candidates with annotations, what was optimized, what constraints were applied, what the model predicted, and the design intent is structured from the absolute start. But then you move into the wet lab execution portion of your pipeline. So protocols are templated and structured. The instrument data flow directly into the platform, so there's no manual transcription, no context is lost at the bench. We capture the data in a structured way so the results are recorded against the experiment's process model, not just as numbers, but as outcomes tied to the conditions and the materials and the decisions. And that's when we can have real insights appear. That's where we can update our models because the results are recorded against the experiment's process model, not just as numbers, but as outcomes tied to the conditions and the materials and the decisions. And that's really Luma. That's what Luma's all about. So it's infrastructure that makes AI reasoning trustworthy. How do you know what's in your tube? That's the fundamental question. So we're making knowledge with every step. So we are combining materials and process and data, as well as knowledge genealogy. So every node is connected. You can trace any decision back to its root experiments, your insights and decisions feedback into the new objectives and deliverables, and it really creates an organic record of the full development train of thought from beginning to end, from objective all the way to decision, go or no go. And so the composability element or the components, we're starting with our experiment and everything that's tied to the experiment, the processes, the materials, the equipment, to generating insight, what are the observations or the conclusions, to the decisions, and then the artifacts that are generated throughout this process. So really, we're weaving that digital thread. So we have an unbroken lineage from first experiment to last batch. Luma connects the data, the materials, the process, and the decisions across the full product life cycle. So knowledge travels with the molecule. So we're closing the gap between discovery and manufacturing. So we're making AI driven discovery an actuality, a deliverable, Right? So we're starting with the problem that AI is accelerating scientific reasoning, but the scientific work remains fragmented. And we're really introducing this strategic shift where we're structuring the scientific work. That's the prerequisite for meaningful AI driven digital transformation. And so the platform requirement for this is bridging or connecting both the wet lab and dry lab orchestration to enable AI agents to reason across the space to iteratively design and implement and refine experiments so that insight flows seamlessly between discovery, development and manufacturing. And so this is really the power that we have with the LUMA platform. So Siemens. MATICS can industrialize an end to end scientific intelligence platform so that we can get effective therapies into the hands of patients faster. So AI only becomes transformative when it's grounded in scientific structured work across all dimensions of this process. So we have structured representation of scientific processes, your materials and how they change, the decisions and the outcomes, and then experimental and operational context so that you have reproducibility in your process, in your science. You have orchestration. You have continuous learning. You have life cycle continuity, all to the end that you can explain and trust the outcome of your process. So we're not just talking about more data. We're aggregating data. We're talking about the structure of the data. And so the AI categories that we're really focused on right now are Luma application development, domain focused AI, assistive intelligence, and automation. And I've just got a couple of the sub domains that we're working on here, and we have a couple of very exciting ones, especially exciting to me anyway, when it comes to protein prediction and de novo design. So we're talking about lab in the loop, AI at every step. Your adaptable tasks that are required to perform multimodal analyses. Your AI powered ELN write ups allow scientists to spend more time in the science. So the ELN generates and unfolds as you do your work. Luma's AI agent can automatically perform analyses, very sophisticated analyses, and generate compliant documents by extracting and formatting the data, ensuring accuracy while maintaining full regulatory traceability. And then we have AI configurable user experiences. So, you can tailor these visuals, these dashboards according to the program or the molecule of interest. And really, this is under the end that you can have user interactive dry lab and wet lab operations, okay? So really this is about faster molecule to market. How do we go from research to development to manufacturing with absolute one hundred percent fidelity and precision? An end to end digitalization process. This is what Luma provides for our drug discovery partners. Thank you very much.
AI is accelerating scientific reasoning, but scientific work itself often remains fragmented, siloed across teams, instruments, and systems. In this on-demand webinar, Christian Olsen, VP of Strategy for Protein Therapeutics at Dotmatics, explores why AI alone can't accelerate biologics discovery unless it's grounded in structured, connected data. Using a real-world Alzheimer's antibody use case, Christian breaks down the difference between material representation (what was made) and process representation (how it was made and tested), and why both are essential for AI-ready data. He'll walk through Luma's approach to composability, the innovation cycle, and closed-loop R&D, showing how structured science, not just more data, is the prerequisite for trustworthy, explainable AI reasoning across discovery, development, and manufacturing.
Watch now to see how Luma is helping drug discovery teams move faster without sacrificing fidelity or precision.
Our Latest on Science & Industry
Simplify your path to discovery.
See Luma in action by requesting a demo today.



