Weather     Live Markets

The Data Dilemma: How One Biotech Founder Is Building a Bridge Between Scientific Discovery and AI Innovation

In the bustling world of biotechnology, where groundbreaking discoveries often hide behind layers of complex data and proprietary research, a fascinating story is unfolding. Harlan Robins, a name synonymous with innovation in the field of immunology, has spent years at the helm of Adaptive Biotechnologies, where he helped build massive datasets that decoded the intricate workings of the human immune system. His journey from chief scientific officer to founder of a new venture represents more than just a career pivot; it embodies a growing frustration that many in the scientific community have been feeling as artificial intelligence continues to reshape the landscape of life sciences. The problem, as Robins sees it, is stark and increasingly urgent: those who invest heavily in generating high-quality research data rarely see fair compensation when AI developers leverage those valuable assets to build commercial tools. It’s a disconnect that has left countless datasets languishing in digital isolation, their potential unrealized and their creators uncompensated.

The birth of Harell Data, Robins’s new Bellevue-based startup, came from this recognition of a fundamental market failure. With $15 million in fresh funding from prominent investors including Fuse and Cercano Management, the company aims to solve a puzzle that has stumped the industry for years. “The goal of the company is to connect AI modelers to proprietary training sets to enable the solution of challenging scientific problems,” Robins explained in a recent interview. His vision is clear: entities that generate proprietary training data using their own technology and resources currently have no good way to commercialize that data, so they effectively sit on it, watching as its value depreciates with each passing day. It’s a situation that makes no sense from an economic perspective, let alone a scientific one. The data could be accelerating breakthroughs in medicine, materials science, and countless other fields, but instead, it remains locked away, accessible only to the privileged few who created it.

To understand why this matters so deeply, consider the remarkable success stories that have emerged when data flows freely. AlphaFold, the revolutionary protein structure prediction system, achieved its stunning results largely because it drew upon decades of publicly available, experimentally derived protein structures. This was data that had been generated through painstaking laboratory work, often funded by government grants and academic institutions, then made available to anyone who wanted to use it. The results speak for themselves: a breakthrough that has transformed our understanding of biology and opened new avenues for drug discovery. But as Robins points out, not all critical data is so readily accessible. In many vital areas of biology and medicine, generating high-quality datasets requires millions of dollars in investment and years of dedicated labor. The organizations that own this data—whether they’re pharmaceutical companies, biotech firms, or research institutions—currently have no secure, profitable way to share it. The result is a vast network of data silos, each one holding pieces of a puzzle that could potentially solve some of humanity’s most pressing health challenges.

Harell Data’s solution to this problem is elegantly simple in concept, though technically sophisticated in execution. The company has built a secure cloud platform that serves as a trusted intermediary between data creators and AI modelers. Here’s how it works: organizations host their proprietary datasets on Harell’s platform, where machine learning groups can train their models without the raw underlying data ever leaving the secure environment. This addresses the legitimate concerns about data security and intellectual property protection that have long prevented collaboration. But the innovation goes deeper than just security. The company has devised a direct revenue-sharing model that changes the economics of data sharing entirely. Instead of waiting years for speculative downstream drug royalties—a process that can take a decade or more and often yields nothing—data owners earn a direct share of the compute revenue generated during training runs. It’s a pay-as-you-go model that provides immediate, tangible returns on data investment. And perhaps most intriguingly, the system creates a compound value effect: as models improve through training on these rich datasets, the intrinsic worth of the underlying data becomes even more valuable, creating a virtuous cycle that benefits everyone involved.

While Harell Data is initially focusing on problems that Robins knows intimately—particularly around computational medicine—the implications of this approach extend far beyond biology. The data-silo problem he’s tackling is universal across scientific disciplines. Materials scientists sit on vast databases of experimental results that could accelerate the development of new alloys and composites. Medical imaging researchers hold millions of scans that could train diagnostic AI systems. Climate scientists have decades of environmental data that could improve predictive models. In every case, the same fundamental issues arise: Who owns the data? How can it be shared securely? And how do we ensure that those who generated it are fairly compensated for their investment? “If the business works right, we should be able to enable solutions to really important problems,” Robins said, and the ambition behind those words is clear. He’s not just building a company; he’s attempting to create an entirely new market for scientific data that could accelerate innovation across multiple fields.

The team behind Harell Data reflects the seriousness of the undertaking. Robins has assembled an 11-person workforce split between Bellevue and Palo Alto, bringing together expertise in both technology and business. The leadership team includes Chief Technology Officer Rakesh Nair, who brings deep experience in building secure computing platforms; Head of Operations Saray Covey, who manages the complex logistics of data onboarding and quality control; and Head of Sales Analise Polsky, who is tasked with convincing data-rich organizations that this new model is worth embracing. The early access for the platform is launching this week, with initial datasets provided by prestigious partners including Seattle-based A-Alpha Bio and Adaptive Biotechnologies. The fact that Adaptive—a publicly-traded company with a market value of $3.9 billion—has signed on as an early partner speaks volumes about the credibility of Robins’s vision. Notably, Adaptive recently spun out another venture called Digital Biotechnologies, which is developing DNA sequencing technology, suggesting that the company sees strategic value in staying at the forefront of computational biology.

Behind the technical innovations and business models, though, lies a human story that makes this venture feel personal. When Robins decided to start Harell Data, he did what many parents do: he asked his 8-year-old son, Ellis, for advice. The boy’s response came within seconds—suggesting “Harell,” a charming combination of “Harlan” and “Ellis.” It’s the kind of moment that reminds us that even the most sophisticated technological ventures often begin with simple human connections. “No offense to the large cap cloud compute companies,” Robins joked, “but it sounds better to me than any of their names, and things seem to have worked out OK for them so I went with it.” The story captures something essential about the entrepreneurial spirit: the willingness to bring family into your work, to take risks on unconventional ideas, and to find joy in the process of creation. As Harell Data begins its journey, it carries not just the weight of ambitious scientific and business goals, but also the innocent wisdom of a child who saw a path forward where others saw only obstacles. Whether the venture achieves all that Robins hopes remains to be seen, but one thing is certain: the conversation about how we value and share scientific data has just become a great deal more interesting.

Share.
Leave A Reply

Exit mobile version