Interview by Katherina Nguyen | Watch on @AGI House
---
Kat : Glad to have you here, Eric, welcome to AGI House. Tell us a little bit about Radical Numerics and what inspired you to found it.
Eric : Absolutely. Thanks for inviting me here. Radical Numerics is a new research lab. Actually, our team is known for working on a model called Evo, which is an AI system that could read and write DNA. And it actually got some traction, got the cover of Science and Nature. Then scientists actually used Evo recently to generate the first whole genome from scratch using AI. Turns out it's this organism called the bacteriophage, also known as a virus, which was a really big turning point for us because we realized that this was a very potentially powerful technology, both for potential human health applications, but also a huge concern on the biological risk standpoint.
So we started the company to work at the frontier of this intersection, both driving the capabilities and frontier of biological AI design. We work on biosecurity applications as well, making sure that the technology is safe and preventing it from misuse.
K : Wow, that's fascinating. So when you first founded it in 2025, what was the AI landscape like back then?
E : Yeah, kind of like right now, except perhaps even now even more feverish. But at that time when we started the company, we actually were trying to stay away from working on science and AI applications. There was this sense that people didn't know how to turn AI into a successful commercial effort in the life sciences.
After we started the company, folks did come to us asking about what we could do and help them make models… I think really at the time, it was this flavor of coding agents being the main thing. Is there value in working on the foundation models or being the layer on top?
We actually took a bet on building the model side, but in a different domain, bio. We're glad we did and I still believe that. The value being around the core models themselves is something we truly believe deeply. We think that's lacking in the biospace, the life sciences space. Very strong, capable models that can actually learn from data at scale.
You just don't see that unlocked in science yet. That's one of our aims – to actually show folks that we can learn from biological data at scale and perhaps work on an application that in my mind is the most important area of AI, life sciences in the physical world.
K: Your interdisciplinary background does seem to help in thinking about how to apply models in novel ways. Could you tell us more about your approach to multimodality?
E : Yeah, absolutely. When I think about multimodal in biology, I think biology is kind of made up of many different languages. So you have grammar around DNA.
DNA is our genetic code. It has all the instructions for making an organism.
And it basically follows some kind of grammar and rules. It turns out it's an extremely complex, perhaps one of the most complex languages ever. And it's just one of them in bio because there's also proteins.
There's also RNA. They all stem from DNA, but they all are pretty specialized. And but they all connect at the same time.
And so for us, that's been the challenge for folks. People thinking that we have to specialize in each of these domains. We have to have a model specialized for proteins, a model specialized for DNA, a model specialized for RNA.
But these things are all part of the same system. And when disease happens or when a drug is in your body, it's inherently a multimodal system. It has a cascade of effects that affects each other. We believe the ability to model that is what's been holding back the field. That's why it takes $2.5 billion to make a drug over 10 years. And that's why we can't figure out how a drug is going to behave in a body because it's inherently a multimodal system.
And so for us, we have now this incredible technology that is AI, the best pattern recognition machine in the world, the best ability to show that modalities can be merged. And at the same time, it's being applied to, sort of, in my mind, the low-hanging fruit, things like text and images, which need to be done. They should be done.
It's obviously super useful. But the next frontier is clearly, in my mind, it's the physical world. This is where the rubber meets the road. And in many ways, multimodality in bio is best set up. Its bio has an inherently multimodal problem. And so for folks interested in this space of, from a technical standpoint, multimodality, biology has the best setup, right?
It's massive-scale data, dozens of modalities, if not more, and potential for huge impact. And so that's been our bet because when we see the companies that have tried to do this, it's just, nothing seems to work.
There's just so much behind in the space that there's a long road ahead. But clearly, this is going to be the most impactful area of AI. Like, what is more important than human health?
So I think more and more, you'll see labs, people working in the space. And you already see it, like Anthropic and DeepMind, OpenAI, all having more initiatives in bio and life sciences. But it's a different ballgame.
I think they'll see that the challenges of trying to apply AI to the space has its own unique set of challenges. But at the same time, it's the most worthwhile thing to work on.
K : Right. I love your metaphor of having the proteins, I've never heard multimodality applied to human biology with that metaphor. But it makes a lot of sense because pattern recognition is how the multimodal is merging these paragraphs and forms. The grammar of the body is very interesting.
E : Yeah. And it's super important and super challenging. It's kind of like another analogy, moving away from the language one.
It's like when I think about how scientists have learned about biology and our bodies, if you were to make an analogy about a person's body as a car, let's say. If a biologist want to learn how a car works, they would study each part, like how an engine works, how a spark plug works, how the tire works, one at a time.
But if you ask an engineer how a car works, they're not going to study one piece at a time, they're going to try to rebuild the whole thing from scratch, each of the pieces, how they fit together.
I think that's, in many ways, what we're trying to do. When we first started working on DNA specifically, we worked on a model called Evo, right, as mentioned. What we were trying to do was understand how life is built from the ground up.
No one had built, no one had used AI to generate DNA from scratch. It was thought to be too difficult. The rule is unclear.
Humans don't need to understand it. How can you teach AI to do it? At the same time, if we could unlock that, I believed, that would change things.
Eventually, people did use Evo to generate a whole genome from scratch, as I mentioned before, bacteriophage virus. In that process of rebuilding things from the ground up, I do believe we are going to drive more deeper understanding because we're going to rebuild that car from the ground up and understand how all the pieces fit together, and not in isolation, but as a whole. I think that's super important.
K : I agree. I think it teaches appreciation from the ground up, and also teaches caution, right? That's why I think what you're doing is fantastic because you use this knowledge to reframe the situation and say, given that we know this can be done, how can we also create counterproductive protective effects?
And at Radical Numerics, you guys have already made quite a bit of progress. I remember you first told me like the Evo model had like context of 121,000...
E : Yeah, 131,000, yeah.
K : And then Evo 2 was 1 million, so the next one might be like 1 billion. That's amazing progress in terms of scaling context.
E : Yeah. Yeah, it's a measure of progress context length. It was one of the key things that we had worked on for the technology to unlock because before we worked on DNA language models, contexts were super short, like talking 2,000, 4,000 tokens.
We had developed this new architecture called Hyena. We could use it in extra long contexts, super long, and at the time text just didn't have long language, natural language didn't have that much interesting long context problems. And that's how we gravitated toward DNA. We thought, what's a domain that has the longest sequences?
Turns out DNA is pretty much the longest I can think of. The human genome, right, our bodies, it's made up of 3 billion letters with 3 billion tokens, right?
That's basically the equivalent of about 30,000 books. And in my mind, that's a key number, right? How can we get to capture the entire context so that we can understand all of the different sort of dependencies and motifs that kind of interact in the same way that we model language and code?
We care about the long context in those domains. But for our genome, our DNA, this is a space, this is a domain that makes us exactly how we are, right? It's the very DNA of ourselves, right?
At the same time, it's wild to me that we actually know very little about our own DNA and how it actually works. I find it interesting that we've been so caught up with using it on language… but using it to unlock how disease works in our body and how you can potentially make precision medicines to cure, like that, does it get more impactful and better than that? Like that's, in my mind, the coolest application, potential application of language models in AI systems.
K : Yeah, and I want to bring up that even though you're working on such amazing frontier tech, you already have existing partners and separate systems. So it's not like a lab experiment.
You have real progress and support for this.
E : Yeah, absolutely. So I just finished a PhD at Stanford and had a great experience and had the privilege to be able to publish the Evo work in the cover of Science and Nature. And I think the phase of academic accolades is important, but also to me, the real value is you got to bring this in the real world.
Such powerful technology, the potential of it, it's not enough for me to just showcase this is theoretically how it can be done. But our goal as a company, yeah, we want to reimagine what folks can think about what these models can do, but we also want to put it in their hands and we also want to save lives. That is one of our key metrics as a lab inside the company– how can we actually use our technology to save lives, concretely.
None of that theoretical, oh, we want to do good. No, we want to actively get this in the real world and to save people's lives.
K : Is there anything you wish to communicate or share out to the community? Are there any partners you wish to work with?
E : Everybody come work with us. We intentionally designed our models to be flexible to work with any type of biological data. So that's kind of generic.
But to be targeted about that, I think more broadly – there is a massive amount of data in the biological space, in the life science space… We internally have been seeing incredible amounts of modality transfer and scale happen, scaling laws happen, occur in this multimodal data that we're working with.
I think we're just at the beginning of starting to showcase to folks that this is possible. And so this year is going to be a year really focused on getting folks to, like I mentioned, reimagine what's possible from data at scale in life sciences.
K : That's awesome. And we're very happy to have people like you working on it.
E : Thank you. Appreciate it. It's been a pleasure chatting with you.


