The whole evidence base.
Every number behind InVision Precision Cardiac Amyloid1 — the FDA validation, peer-reviewed international validation across five sites in two countries, and an independent head-to-head run by investigators outside InVision. Including the metrics where we score lower.
FDA 510(k) K243866
Cleared 21 May 2025 as an adjunctive diagnostic aid. Software as a Medical Device, product code SW0016.
| Item | Value |
|---|---|
| Validation cohort | 1,221 unique echocardiogram studies |
| Sites | 3 US sites |
| Case-to-control ratio | 1:2 |
| Reference standard | Confirmatory imaging or pathology |
| Views used | Apical four-chamber and parasternal long-axis |
| Intended population | Adults aged 65 and older |
| Clearance type | Adjunctive diagnostic aid |
| Subgroup analyses | Age, gender, BMI, race/ethnicity, amyloidosis type, site, and imaging manufacturer — all met acceptance criteria |
| Human factors validation | n=31 (16 non-cardiologists, 15 cardiologists) |
| Prior designations | Breakthrough Device Designation (January 2024); FDA TAP enrolment |
Peer-reviewed international validation
JACC: Advances, 2025. Temporally and geographically distinct external validation, conducted by investigators at five institutions across two countries.
| Metric | Value | Confidence interval |
|---|---|---|
| Overall AUC | 0.893 | 95% CI 0.874–0.912 |
| Sensitivity | 0.648 | 95% CI 0.607–0.687 |
| Specificity | 0.982 | 95% CI 0.971–0.989 |
| PPV at 2:1 control-to-case | 0.954 | 95% CI 0.928–0.972 |
| NPV at 2:1 control-to-case | 0.826 | 95% CI 0.804–0.848 |
The cohort
| Cases | 520 patients with cardiac amyloidosis |
| Controls | 903 matched controls, all aged 65 and older |
| Countries | United States and Japan |
| Case mix | ATTR 74.2%, AL 22.7% |
| Reference standard | Bone-tracer nuclear imaging, monoclonal gammopathy testing, genetic testing, and/or tissue biopsy |
| Sites | Cedars-Sinai Medical Center · Keio University, Tokyo · Northwestern Medicine · Kumamoto University · Yale-New Haven |
How to read sensitivity and specificity
Two devices can both be well built and behave completely differently, because they are cleared for different jobs. A screening aid is tuned to catch more and accepts more false flags. A diagnostic aid is tuned so that nearly every flag is real. InVision Precision Cardiac Amyloid is cleared as an adjunctive diagnostic aid, at 99.0% specificity.
PPV = (sens × prev) ÷ [ (sens × prev) + (1 − spec) × (1 − prev) ]
NPV = (spec × (1 − prev)) ÷ [ (spec × (1 − prev)) + (1 − sens) × prev ]
Computed at the K243866 operating point: sensitivity 0.607, specificity 0.990. Predictive values depend on the prevalence of the population tested, so the table shows a range rather than a single number.
| Prevalence | PPV | NPV | True flagsper 1,000 | False flagsper 1,000 | Missedper 1,000 |
|---|---|---|---|---|---|
| 1% | 38.0% | 99.60% | 6.1 | 9.9 | 3.9 |
| 2% | 55.3% | 99.20% | 12.1 | 9.8 | 7.9 |
| 5% | 76.2% | 97.95% | 30.4 | 9.5 | 19.7 |
| 10% | 87.1% | 95.78% | 60.7 | 9.0 | 39.3 |
| 14% | 90.8% | 93.93% | 85.0 | 8.6 | 55.0 |
The corresponding figure from the peer-reviewed international validation is a PPV of 0.954 (95% CI 0.928–0.972) at a 2:1 control-to-case ratio.
The honest counterpart, stated plainly: a higher-specificity operating point misses more cases. At 60.7% sensitivity, a negative result does not exclude cardiac amyloidosis.
Independent head-to-head
JACC: Advances 2025. Northwestern-led. Integrated health system, 2010–2022. 176 confirmed ATTR-CM cases and 3,192 heart failure controls, matched to a target prevalence of 5%. We host this study because it tested our cleared model and we would rather you read the whole table than half of it. Each model evaluated at its own developer-specified threshold. InVision's is ≥0.8.
| Metric | InVision Precision Cardiac Amyloid |
EchoGo Amyloidosis |
Mayo ATTR-CM Score |
|---|---|---|---|
| F1 | 0.610.55–0.67 | 0.490.43–0.53 | 0.310.26–0.35 |
| Accuracy | 0.960.96–0.97 | 0.910.90–0.92 | 0.850.83–0.86 |
| Sensitivity | 0.570.50–0.64 | 0.850.79–0.90 | 0.650.59–0.73 |
| Specificity | 0.980.98–0.99 | 0.910.90–0.92 | 0.860.84–0.87 |
| PPV | 0.660.58–0.74 | 0.340.30–0.39 | 0.200.17–0.23 |
| NPV | 0.980.97–0.98 | 0.990.99–0.99 | 0.980.97–0.98 |
| False negative rate | 0.430.35–0.51 | 0.150.11–0.21 | 0.350.27–0.42 |
| Average precision | 0.620.55–0.69 | 0.670.59–0.74 | 0.160.13–0.20 |
| AUC | 0.880.84–0.91 | 0.920.89–0.95 | 0.790.76–0.83 |
Shading marks the higher value where the 95% confidence intervals do not overlap. The AUC confidence interval for InVision appears as 0.84–0.91 in Table 3 and 0.85–0.91 in the body text. Table 3 is cited here.
Read it in both directions. Confidence intervals do not overlap for PPV, specificity, F1, and accuracy, favouring InVision. They also do not overlap for sensitivity and false negative rate, favouring the other model. On average precision — which the authors note may be more informative than AUC in data sets with class imbalance — the intervals overlap.
Both deep learning models outperformed the Mayo ATTR-CM score, with DeLong P < 0.001 for each. The paper reports no DeLong test between the two deep learning models.
The authors also observe that both deep learning models outperformed the Mayo score on positive predictive value, and that a reduction in unnecessary downstream testing from fewer false positives could help offset the higher upfront cost of a more complex model.
Subgroups
Moderate or greater left ventricular hypertrophy (n=504, prevalence 23.8%)
| Metric | InVision | EchoGo Amyloidosis |
|---|---|---|
| Sensitivity | 0.660.57–0.74 | 0.920.87–0.96 |
| Specificity | 0.960.94–0.98 | 0.850.82–0.88 |
| Average precision | 0.810.74–0.87 | 0.830.76–0.89 |
| AUC | 0.910.88–0.94 | 0.930.90–0.95 |
LVEF below 40% (n=713)
| Metric | InVision | EchoGo Amyloidosis |
|---|---|---|
| Sensitivity | 0.590.43–0.77 | 0.820.69–0.94 |
| Specificity | 0.980.97–0.99 | 0.880.85–0.90 |
| Average precision | 0.650.45–0.79 | 0.670.49–0.83 |
| AUC | 0.880.79–0.95 | 0.890.79–0.95 |
In both subgroups the AUC and average precision intervals overlap heavily. The pattern is stable: our specificity is higher, their sensitivity is higher.
Study limitations. Retrospective, single health system. Prevalence set artificially to 5%, so absolute predictive values vary with the prevalence of the population tested. The authors note that relative comparisons of predictive value between models remain useful.
What a false flag actually costs
Every flag triggers a workup: bone-tracer scintigraphy, serum and urine immunofixation, free light chains, sometimes biopsy or genetic testing. False flags consume nuclear cardiology capacity and clinic slots, and they expose patients to tests they did not need.
~10 false flags per 1,000 studies
The operating point InVision Precision Cardiac Amyloid is cleared at.
~100 false flags per 1,000 studies
A tenfold difference in downstream workup burden, for the same 1,000 studies.
In the independent cohort above, the arithmetic works out as follows. Computed directly from Table 3 against the study's own 176 cases and 3,192 controls. InVision roughly 164 patients flagged to find 100 cases. The comparator model roughly 437 patients flagged to find 150 cases. The 50 additional cases cost roughly 223 additional false positives — about 4.5 unnecessary workups per additional case found.
Both halves are true. We find fewer cases. We generate far fewer unnecessary workups. Which operating point suits your service depends on your nuclear cardiology and amyloid clinic capacity — which is exactly why the two devices are cleared for different jobs.
Subgroups and fairness
In the FDA validation, subgroup analyses across age, gender, BMI, race/ethnicity, amyloidosis type, site, and imaging manufacturer all met the acceptance criteria. The independent study went further and ran a bias audit.
Equal opportunity ratio 1.05
Met the study's equal opportunity criterion for patients who self-identified as Black, statistically significant for parity. ATTR-CM prevalence in that group was about five times higher than in other groups in this cohort.
Demographic parity 0.64
Flagged by the paper as a disparity under the 80% rule. Published here alongside the favourable result.
We publish the flagged metric next to the favourable one. An AI governance committee that only ever sees a vendor's good number has learned nothing about the vendor.
Why the same model reports different AUCs
You will see several numbers for this model: 0.83 internal and 0.79 external in 2022, 0.88 in the independent heart failure cohort, 0.893 across five international sites, and 0.833 to 0.944 by individual site. That spread is not instability. It is cohort construction.
InVision published the paper that demonstrates this. In Impact of Case and Control Selection on Training Artificial Intelligence Screening of Cardiac Amyloidosis (JACC: Advances, 2024), model AUCs ranged from 0.660 to 0.898 on matched held-out test sets, and from 0.467 to 0.898 in a general patient population — varying by nothing except how cases and controls were selected.
An AUC is a property of a model and a cohort, not of a model alone. Comparing an AUC from one study against an AUC from another study with different case and control selection is not a valid comparison. That is why the head-to-head section above matters: it is the one place where every model was measured on the same patients.
Provenance and open science
The model family behind Precision Cardiac Amyloid was developed and published in the peer-reviewed literature, and the underlying dataset was released publicly. Independent groups have tested it without our involvement and published what they found.
| JAMA Cardiology 2022 | Value |
|---|---|
| Patients | 23,745 |
| Training videos | 24,804 echocardiogram videos |
| Cardiac amyloidosis classification AUC | 0.83 |
| Hypertrophic cardiomyopathy AUC | 0.98 |
| External validation, amyloidosis | AUC 0.79 |
| External validation, HCM | AUC 0.89 |
| Wall thickness mean absolute error | 1.4 mm (95% CI 1.2–1.5) |
| LV diameter mean absolute error | 2.4 mm (95% CI 2.2–2.6) |
| Posterior wall mean absolute error | 1.2 mm (95% CI 1.1–1.3) |
| Open dataset released | 23,212 annotated echocardiogram videos |
These figures are provenance, not current performance. They are superseded by the 2025 international validation above.
What this device does not do
InVision Precision Cardiac Amyloid is an adjunctive diagnostic aid. It does not diagnose.
Confirmation requires bone-tracer nuclear imaging, monoclonal gammopathy testing, genetic testing, and/or tissue biopsy. The software produces a structured starting point; a physician authors and signs every report.
Sensitivity at the cleared operating point is 60.7%. A negative result does not exclude cardiac amyloidosis. The intended population is adults aged 65 and older.
Both validations described on this page are retrospective case-control studies. Predictive values shift with the prevalence of the population actually tested, which is why the table in Reading sensitivity and specificity shows a range.
Citations
1. Published in the peer-reviewed literature as EchoNet-LVH. EchoNet-LVH and InVision Precision Cardiac Amyloid are the same model — the literature uses the research name, the FDA clearance uses the product name. Every external validation of EchoNet-LVH is therefore an external validation of the cleared device.
- International Validation of Echocardiographic Artificial Intelligence Amyloid Detection Algorithm. JACC: Advances. 2025. doi:10.1016/j.jacadv.2025.102067 ↗
- Evaluating the Performance and Potential Bias of Predictive Models for Detection of Transthyretin Cardiac Amyloidosis. JACC: Advances. 2025;4(8). doi:10.1016/j.jacadv.2025.101901 ↗
- High-Throughput Precision Phenotyping of Left Ventricular Hypertrophy With Cardiovascular Deep Learning. JAMA Cardiology. 2022;7(4):386–395. doi:10.1001/jamacardio.2021.6059 ↗
- Impact of Case and Control Selection on Training Artificial Intelligence Screening of Cardiac Amyloidosis. JACC: Advances. 2024. doi:10.1016/j.jacadv.2024.100998 ↗
Questions about the evidence?
Our clinical team will walk through any figure on this page, including the ones where we score lower, and what the operating point means for your service.