September 23, 2026·6 min read·AIgentic.media

Nature Is AI's New Training Ground

ai-drug-discoveryresearchai-fundingnvidiabiotech
Nature Is AI's New Training Ground

Basecamp Research collects genetic samples from rainforests, volcanoes, and deep-sea vents across all seven continents, creating what it calls the largest biological AI training dataset ever assembled.

For years, the AI industry's training data has followed a predictable pattern: text from the internet, images from the web, code from GitHub. The richest untapped training set on Earth has been sitting under our feet the entire time.

Basecamp Research, a London-based startup, just raised $140 million from Nvidia, Anthropic's Anthology Fund, the NATO Innovation Fund, and others to prove that the most powerful training signal for AI isn't on the open web - it's in the genomes of microorganisms living in rainforest soil, deep-sea vents, and volcanic hot springs.

The thesis is as audacious as it is simple. Evolution has been running experiments for 4 billion years. Every organism alive today is a solution to a biological problem. Basecamp's AI, called EDEN, learns the patterns behind those solutions and designs new medicines from them.

The Scale Problem in Biology

CTO Philip Lorenz draws a striking comparison. Language models train on roughly five quadrillion words - essentially the sum of all text ever written. Basecamp estimates the number of nucleotides on Earth at ten to the power of 37. "If you took a stack of poker cards of 10 to the 37th power, that stack would surround the observable universe a million times," Lorenz says.

The gap between available biological data and what AI needs is staggering. Even more problematic: public genome databases are dramatically skewed. About 68 percent of all sequencing data in the Sequence Read Archive comes from just five species. Humans alone account for 54 percent.

"If you were to train an LLM only on newspaper articles from 1975, it would be a really, really bad model," Lorenz told The Decoder in an interview. "That is kind of where we are in biology."

Basecamp's solution is to collect its own data, working with local research partners in more than 30 countries across all seven continents. The dataset now holds roughly 15 trillion DNA tokens - comparable in scale to the text datasets used to train frontier language models like GPT. By the end of next year, the company plans to pass one quadrillion tokens as part of its Trillion Gene Atlas project with Anthropic, Nvidia, PacBio, and Ultima Genomics.

EDEN: The Model That Learns From Evolution

EDEN - Basecamp's flagship 28-billion-parameter model - doesn't process words. It processes strings of DNA building blocks, learning which genetic patterns correlate with useful biological properties. The first generation was trained on 9.7 trillion DNA building blocks from more than one million newly sequenced species, running on Microsoft Azure with compute that reportedly matched GPT-4's training run.

The results are early but striking. In antimicrobial peptide tests, 97 percent of EDEN-designed candidates showed activity in the lab. One candidate, EDEN-7, worked about as well as a last-resort antibiotic in mice infected with multidrug-resistant bacteria - with no iterative tweaking before the test.

For gene therapy, EDEN designs large serine recombinases - enzymes that can insert therapeutic DNA at precise locations in the human genome. Half of the generated recombinases were active in human cells. In tests with CAR-T cells - immune cells modified to fight cancer - the approach cleared more than 90 percent of tumor cells in the lab.

The Irony Beneath the Headline

The funding announcement reads like a standard biotech Series C. What makes it genuinely surprising is the inversion at its core. The AI industry has spent years building models that learn from human output - our text, our images, our code. Basecamp argues that human output is a narrow slice of the intelligence that exists on this planet. Life itself has been generating solutions to problems we are only beginning to understand, and those solutions are encoded in genomes we have barely cataloged.

This is not cheap. The company has built a global supply chain for genetic material that spans five dozen countries, requires consent agreements with local governments, and involves extracting and sequencing DNA from environments most humans will never visit. Every token in EDEN's training set is traceable to its exact geographic origin and the consent that authorized its collection.

The Long Road to the Clinic

Basecamp's ambition is nothing less than making biology programmable. "You prompt on a disease and out comes a molecule that will address that," Lorenz says. The company is transparent about how far that goal remains. None of its six therapeutic programs has advanced beyond lead optimization - the stage where promising candidates are refined before entering preclinical studies and eventually human trials.

The pipeline includes in vivo CAR-T cell therapies for blood cancer and autoimmune disease, a liver gene therapy for the metabolic disorder phenylketonuria (PKU), antimicrobial peptides against resistant pathogens, and early-stage work on solid tumor CAR-T therapy and diabetes treatments.

Delivery, toxicity, and the body's immune responses remain open challenges. The Arc Institute's Patrick Hsu has pointed out that a genetic edit that works in the lab doesn't answer how the necessary components reach the right cells in the body.

Why This Matters Now

Basecamp's $140 million raise comes at a moment when the AI industry is questioning whether scaling laws apply outside language. Lorenz's answer is a qualified yes. Basecamp's internal tests show that scaling curves hold for biological models - but with different ratios than language models. The company reserves a third of its GPUs for reinforcement learning experiments, betting that the same techniques that drove recent gains in models like GPT-6 will also work for designing molecules.

The company also warns against trusting technical metrics alone. In tests comparing model architectures, StripedHyena sometimes reached better perplexity scores, but Llama - the architecture Basecamp ultimately chose - performed better on actual biological tasks. "Not to overpromise what AI can or can't do, but just do the experiments, figure out what it does," Lorenz says.

The story of Basecamp Research is a story about data. Not synthetic data, not web-scraped data, not data generated by another AI - but data collected from the real world, from environments where life has been solving problems longer than humans have existed. For an industry increasingly aware that it has scraped the internet dry, that is perhaps the most interesting signal of all.

Sources

Want to learn more?

Let's discuss how AI can transform your business.

Get in Touch