# Diagnostic Accuracy of Web-Based COVID-19 Symptom Checkers: Comparison Study

**Authors:** Nicolas Munsch, Alistair Martin, Stefanie Gruarin, Jama Nateqi, Isselmou Abdarahmane, Rafael Weingartner-Ortner, Bernhard Knapp

PMC · DOI: 10.2196/21299 · 2020-10-06

## TL;DR

This study compares how well different online tools can correctly identify whether someone has COVID-19 based on their symptoms.

## Contribution

The study is the first to rigorously evaluate the diagnostic accuracy of web-based COVID-19 symptom checkers using statistical methods.

## Key findings

- Symptom checkers varied widely in their ability to correctly identify COVID-19 cases.
- Only two tools achieved a good balance between correctly identifying positive and negative cases.

## Abstract

A large number of web-based COVID-19 symptom checkers and chatbots have been developed; however, anecdotal evidence suggests that their conclusions are highly variable. To our knowledge, no study has evaluated the accuracy of COVID-19 symptom checkers in a statistically rigorous manner.

The aim of this study is to evaluate and compare the diagnostic accuracies of web-based COVID-19 symptom checkers.

We identified 10 web-based COVID-19 symptom checkers, all of which were included in the study. We evaluated the COVID-19 symptom checkers by assessing 50 COVID-19 case reports alongside 410 non–COVID-19 control cases. A bootstrapping method was used to counter the unbalanced sample sizes and obtain confidence intervals (CIs). Results are reported as sensitivity, specificity, F1 score, and Matthews correlation coefficient (MCC).

The classification task between COVID-19–positive and COVID-19–negative for “high risk” cases among the 460 test cases yielded (sorted by F1 score): Symptoma (F1=0.92, MCC=0.85), Infermedica (F1=0.80, MCC=0.61), US Centers for Disease Control and Prevention (CDC) (F1=0.71, MCC=0.30), Babylon (F1=0.70, MCC=0.29), Cleveland Clinic (F1=0.40, MCC=0.07), Providence (F1=0.40, MCC=0.05), Apple (F1=0.29, MCC=-0.10), Docyet (F1=0.27, MCC=0.29), Ada (F1=0.24, MCC=0.27) and Your.MD (F1=0.24, MCC=0.27). For “high risk” and “medium risk” combined the performance was: Symptoma (F1=0.91, MCC=0.83) Infermedica (F1=0.80, MCC=0.61), Cleveland Clinic (F1=0.76, MCC=0.47), Providence (F1=0.75, MCC=0.45), Your.MD (F1=0.72, MCC=0.33), CDC (F1=0.71, MCC=0.30), Babylon (F1=0.70, MCC=0.29), Apple (F1=0.70, MCC=0.25), Ada (F1=0.42, MCC=0.03), and Docyet (F1=0.27, MCC=0.29).

We found that the number of correctly assessed COVID-19 and control cases varies considerably between symptom checkers, with different symptom checkers showing different strengths with respect to sensitivity and specificity. A good balance between sensitivity and specificity was only achieved by two symptom checkers.

## Linked entities

- **Diseases:** COVID-19 (MONDO:0100096)

## Full-text entities

- **Diseases:** loss of appetite (MESH:D001068), H1N1 influenza (MESH:D007251), cough (MESH:D003371), smoking (MESH:D015208), knee problems (MESH:D007718), dyspnea (MESH:D004417), respiratory distress (MESH:D012128), asthma (MESH:D001249), viral infection (MESH:D014777), COVID-19 (MESH:D000086382), fracture (MESH:D050723), infected (MESH:D007239), respiratory disorder (MESH:D012131), fever (MESH:D005334), pulmonary disease (MESH:D008171), heatstroke (MESH:D018883), knee pain (MESH:D046788), Corona symptom (MESH:D018352),  (MESH:D011024)
- **Species:** Homo sapiens (human, species) [taxon 9606], Gammacoronavirus (genus) [taxon 694013]

## Figures

5 figures with captions in the complete paper: https://tomesphere.com/paper/PMC7541039/full.md

---
Source: https://tomesphere.com/paper/PMC7541039