
Real-world patient data with the depth your question demands
Electronic health records for more than 4 million patients, enriched with the specialist measures generic datasets leave behind, across journeys a third of this population has built over a decade.
- 4M+
- patients of EHR data
- 250K+
- whole genomes
- 11M+
- imaging studies
- 1.9M+
- digital ECGs
- 1B+
- clinical notes
Data provenance you can trust
Every NashBio data product is derived from the electronic health records of Vanderbilt University Medical Center, an academic medical center with 6,650+ providers across 70+ locations caring for roughly 3 million patients a year, from routine primary care through nationally ranked cardiology, oncology and transplant programs.
That clinical range creates the depth: specialist-verified diagnoses, specialty lab panels and advanced imaging captured at the point of care, in records reaching back to the integrated EHR VUMC launched in 2001.

Integrated breadth and depth
Start with the structured clinical foundation, then add the layers your question needs. Every layer is de-identified to HIPAA Safe Harbor, with date-shifting that preserves the timing of the patient journey.
Make sure every cohort is fit-for-purpose
NashBio data products support RWE, HEOR, epidemiology and target discovery teams whose cohort definitions turn on a specialist measure: a staging code, an ejection fraction, a fibrosis stage or a treatment-response signal recorded in a note.
Feasibility comes first. The Data Dashboard and NB Person Catalog let your team size a population and confirm data availability before a license is signed.

Ready to dig deeper? Bring us your use case and we will get the data you need.
No license required to run a feasibility review.
Frequently Asked Questions
- Where does NashBio's real-world data come from?
- NashBio data is derived from the electronic health records of Vanderbilt University Medical Center, a single academic medical center with 6,650+ providers across 70+ locations caring for roughly 3 million patients a year. Because it comes from one health system rather than aggregated sources, every layer describes the same patients under the same care standards.
- How many whole genome sequences are linked to clinical data?
- NashBio holds 250K+ whole genome sequences, each linked to that patient's longitudinal electronic health record. The cohort carries an 11-year median EHR length and a median of 50 clinical visits per patient, which is the clinical context needed to connect a rare variant to a real treatment outcome. Sequences are 30X coverage on GRCh38.
- What clinical measures are missing from most real-world datasets?
- The specialist measures that decide clinical endpoints are usually absent, including ejection fraction, LVOT gradient, pathological TNM stage, Lp(a) and liver fibrosis stage, all of which are recorded in free-text reports rather than structured fields. NashBio curates them into analysis-ready tables, with 75+ echocardiogram and 25+ ECG features across 960K+ cardiology patients.
- Can I check whether NashBio has data for my cohort before licensing it?
- Yes. The Data Dashboard lets your team size a cohort and test inclusion criteria across the full 4M+ patient population, and the NB Person Catalog shows patient-level availability of genomic, imaging and waveform data. Both are available before a license is signed, so you confirm the data answers your question first.
