Patient data sits underneath nearly every AI use case in life sciences, and yet the complaints about access, completeness, and quality sound much the same as they did twenty years ago. In this DeciBio Q&A, Blanca Baez, founder of DBounce and a two-decade veteran of pharma, biotech, and AI for drug R&D, walks through how the working definition of patient data has expanded, what the roughly 350 companies her team mapped have gotten right and wrong on business model, and where she thinks the real culprits sit. Her view is that regulation takes more of the blame than it deserves, and that value to patients, data adequacy, and an over-reliance on pharma explain far more of the stall.
Thanks for joining, Blanca. To get started, could you share a bit about your background and what led you to found DBounce?
Thank you for having me. I have been in the life sciences space for over twenty years, across pharma, biotech, and AI for drug R&D, and the themes running through all of it have been big data, then data for drug R&D, and then data as the fuel for artificial intelligence. We are living through a revolution in what data lets us achieve that we could not achieve before, and I have been passionate about that whole field of possibility for a long time.
DBounce came out of the unmet needs I kept running into around patient data. My firm belief, and I know you and I have discussed this, is that we do not have enough data, and we are not producing or generating enough of it to fuel all the algorithms we could be building against the most pressing biomedical and healthcare unmet needs we face right now.
We throw the term "data" around quite a bit, so it is worth anchoring what we actually mean by it. What are the different types of patient data, and how has that definition shifted over time?
That question is essential to anchor, because "data," much like "AI," gets used in a very vague manner, and without precision we do not find solutions to anything.
In the big data and real-world data era, we typically meant records captured when a person visited a doctor or hospital, plus clinical trial readouts, so the market was, in practice, a question of how to buy or access electronic health records. Genomics and omics then added many more sources, and at the AAIH (Alliance for Artificial Intelligence in Healthcare) we began working on standards for merging clinical data with sequencing data, because putting the two together gives you far more discovery and inference power. In the last five years, home devices and sensors have joined the mix, a ring, a Whoop, a glucose monitor, along with the last piece, patient self-reported data. Chronic diseases are on the rise and patients want their data at their fingertips. Data captured only when someone visits a clinic is episodic and reactive, and it misses most of what we need. It has to become more real-time, holistic, proactive, and personalized.
As those data types proliferated, it became a question of incentives, of who uses the data and who pays for it. From your vantage point, what business models have entered this space, and what has worked and what has not?
The patient data market is one of the largest in life sciences and one of the ones carrying the most unmet need. Billions are spent generating, curating, and accessing it, and the buyers keep multiplying: precision medicine, big data in healthcare, digital and telehealth, and more recently wellness, where longevity, microbiome research, and epigenetics are converging with precision and preventative medicine. Even so, a lot of companies have not delivered on the promise, and a lot of people put that down to regulation. At DBounce we researched roughly 350 companies globally, segmented by business model and use case, and I firmly believe this goes well beyond regulation and sits squarely with business models.
Setting aside electronic health records, which followed the usual path of incumbents, fragmentation, and consolidation, the frustration sits where the patient is accessed directly. Disease registries feeding pharma with real-world evidence, CorEvitas for example, have been solid, and digital healthcare delivery has produced strong players like Huma and Welldoc. Companies that went directly to patients and tried to give them agency over their data have had a harder time, StuffThatWorks, LunaDNA, and 23andMe among them. Online health communities capture the patient voice, though it never connects to the patient's records; consumer genetic profiling suffered from data that was never connected to everything else happening for that person; and health-marker self-monitoring quieted down after the hype, held back by poor medical education at population scale. The newest segment, data brokers out of the decentralized science movement, gives people a digital data wallet to commercialize their own data. There is real promise in that angle, but the technology has not let these players thrive yet.
That maps closely to the buckets we see. What strikes me is that this area has been around a long time without producing a clear winner or a go-to-market approach that works universally, within each bucket and not just across them. What do you see as the key culprits keeping us from leveraging patient data fully?
Both sides, the corporation that needs the data and the person who owns their health, have pain points we have not addressed. Companies have tried sharing value back with patients, Evidation gave people cash-convertible rewards for their fitness data, and that helps, but patient centricity has become a trite and empty term. The business models and incentives on both sides have not been in the right place. Across both, we see the shortcomings clustering in three categories.
The first is value to patients. Patients of several companies that shut down or pivoted told us they gave their data and got back far less than they believed they were giving. And because Western research sees everything in silos, one organ, one diagnosis, one biomarker at a point in time, a person gets tired of being asked about one slice of their health by one company and a different slice by another. That fragmented version of health feels artificial.
The second is data adequacy. Surprisingly few companies in this space have a data strategy, so there is a lot of reactive scrambling for datasets that end up incomplete or below research grade. Collection systems should be designed so the data is ready to plug into algorithms, which generative AI and machine learning now make possible, and many players have gone too far toward the individual profile or too far toward the anonymous cohort, when the value of the data depends on both.
The third is business model. Pharma's deep pockets have created a myopic view in which everything comes back to whether pharma will buy this, and it is a fallacy that the only route to a lucrative, meaningful company runs through a pharma play. Hospitals, payers, regulators, academia, and smaller biotechs are all intertwining in data exchanges and partnerships, and chasing an exclusive pharma deal is hard to scale and stifles innovation. As our friend Paul Howard puts it, we should be competing on algorithm use cases and AI tools. The database is raw material. These unmet needs have moved very little in the twenty years I have been in this industry, which tells me we need much more innovation across all three points.
Silos come up repeatedly here, and my read is that a good part of that is driven by pharma being the party with the money and treating data as a competitive asset. What do we do to break those silos down?
We need new paradigms and new ways of designing data collection systems, and Western medicine has to recognize that chronic disease is holistic in how it manifests, how it comes about, and in its outcomes and treatment. The industry has put billions into Alzheimer's, autoimmune disorders, and rare disease over decades without advancing the way people need us to, partly because the angle has to be more preventative. Waiting until someone has a tumor the size of a hand and then asking which drug intervenes is always going to be playing catch-up.
Three tenets are core for us. The first is continuous capture of a person's state of health and disease patterns, so the data reflects real-world health. People with chronic diseases do this on their own already. They are not waiting until they see their doctor to figure out what is going on as they already spend hours and days on end managing their disease and tracking its patterns. They capture everything they can get their hands on, through wearables, through medical devices at home, through their own observations, and they build their own formulas out of every factor they believe bears on their disease. They are far more sophisticated than anyone gives them credit for, and the heavier the disease weighs on their quality of life, the more sophisticated they get. None of it is taken into account. If I sit in one more meeting where someone says patient data is very hard to get, honestly, stop focusing on the problem and start designing the datasets your research programs need.
The second is a holistic approach: lifestyle, psycho-emotional, socio-economic, and environmental indicators are all collected in silos today, if at all, and have to be aggregated with a person's clinical and genomic profile in order to truly understand their disease case. The third is multimorbidity. The medical literature has captured these patterns well over time with a primary disease at the center, and around it the comorbidities that typically follow, with their time to onset and their incidence. If the science already knows those co-occurrences, we should be collecting data across all of them, because a lot of comorbidities are preventable. Not every psoriatic patient will develop diabetes or depression, but too many do.
To round out the discussion, what other pillars of innovation do you see beyond those three, and how close do you think we are to realizing them?
The effort has to be collective. We have chosen to focus on collecting chronic disease data that is not being collected today, to fuel algorithms in precision, preventative, and longevity medicine over the next ten years and empower people to use their own data to avoid unnecessary suffering. There is already enough hardship in living with a chronic disease. Some outcomes cannot be avoided, but many can. Beyond that, other deep tech could get us to answers faster: synthetic datasets, digital twins, and some very interesting ex-vivo models that pair an AI approach with an organism approach. We also need evidence we do not have. You and I have talked about young people in their late teens and early twenties being diagnosed with all kinds of tumors despite lifestyles considerably healthier than a lot of baby boomers, and we do not know why. The dialogue is much larger than data alone, and where I think we are shortest is on imagination and creativity. That is the part we need to enlarge.
That is a good note to close on, even if I would rather we were not short on it. There is clearly a lot more here, and perhaps next time we take the more optimistic view of what becomes possible once these problems get solved. Blanca, thank you for your time.
Thank you for the time and for chatting with me!
Comments and opinions expressed by interviewees are their own and do not represent or reflect the opinions, policies, or positions of DeciBio Consulting or have its endorsement. Note: DeciBio Consulting, its employees or owners, or our guests may hold assets discussed in this article/episode. This article/blog/episode does not provide investment advice, and is intended for informational and entertainment purposes only. You should do your own research and make your own independent decisions when considering any financial transactions.








