NashBio
An abstract white rendering of a column of cubes dissolving upward into scattered particles, flecked with blue.

Real-world patient data with the depth your question demands

Electronic health records for more than 4 million patients, enriched with the specialist measures generic datasets leave behind, across journeys a third of this population has built over a decade.

4M+
patients of EHR data
250K+
whole genomes
11M+
imaging studies
1.9M+
digital ECGs
1B+
clinical notes

Data provenance you can trust

Every NashBio data product is derived from the electronic health records of Vanderbilt University Medical Center, an academic medical center with 6,650+ providers across 70+ locations caring for roughly 3 million patients a year, from routine primary care through nationally ranked cardiology, oncology and transplant programs.

That clinical range creates the depth: specialist-verified diagnoses, specialty lab panels and advanced imaging captured at the point of care, in records reaching back to the integrated EHR VUMC launched in 2001.

The Vanderbilt Health sign on a medical centre building, framed by autumn leaves.

Integrated breadth and depth

Start with the structured clinical foundation, then add the layers your question needs. Every layer is de-identified to HIPAA Safe Harbor, with date-shifting that preserves the timing of the patient journey.

Structured clinical data
Depth4M+ patients4.4-year median history · 32% with 10+ years
What is in itDemographics, diagnoses, procedures, labs, medications, vitals and encounters in OMOP 5.3.1
HIPAA Safe HarborDate-shifted, patient-level linkage across every layer

Make sure every cohort is fit-for-purpose

NashBio data products support RWE, HEOR, epidemiology and target discovery teams whose cohort definitions turn on a specialist measure: a staging code, an ejection fraction, a fibrosis stage or a treatment-response signal recorded in a note.

Feasibility comes first. The Data Dashboard and NB Person Catalog let your team size a population and confirm data availability before a license is signed.

The NashBio Data Dashboard showing cohort counts, demographics and top codes for a diabetes cohort.

Ready to dig deeper? Bring us your use case and we will get the data you need.

No license required to run a feasibility review.

Frequently Asked Questions

Where does NashBio's real-world data come from?
NashBio data is derived from the electronic health records of Vanderbilt University Medical Center, a single academic medical center with 6,650+ providers across 70+ locations caring for roughly 3 million patients a year. Because it comes from one health system rather than aggregated sources, every layer describes the same patients under the same care standards.
How many whole genome sequences are linked to clinical data?
NashBio holds 250K+ whole genome sequences, each linked to that patient's longitudinal electronic health record. The cohort carries an 11-year median EHR length and a median of 50 clinical visits per patient, which is the clinical context needed to connect a rare variant to a real treatment outcome. Sequences are 30X coverage on GRCh38.
What clinical measures are missing from most real-world datasets?
The specialist measures that decide clinical endpoints are usually absent, including ejection fraction, LVOT gradient, pathological TNM stage, Lp(a) and liver fibrosis stage, all of which are recorded in free-text reports rather than structured fields. NashBio curates them into analysis-ready tables, with 75+ echocardiogram and 25+ ECG features across 960K+ cardiology patients.
Can I check whether NashBio has data for my cohort before licensing it?
Yes. The Data Dashboard lets your team size a cohort and test inclusion criteria across the full 4M+ patient population, and the NB Person Catalog shows patient-level availability of genomic, imaging and waveform data. Both are available before a license is signed, so you confirm the data answers your question first.