TL;DR
Divisi is an interactive tool that enables scalable exploratory subgroup analysis in complex datasets, helping data scientists discover surprising patterns and better understand data features and their interactions.
Contribution
It introduces Divisi, a novel interactive notebook-based system with a fast approximate subgroup discovery algorithm and visualization tools for comprehensive subgroup exploration.
Findings
Divisi helps uncover surprising data patterns.
It encourages thorough exploration of complex subgroups.
Practitioners find it improves understanding of data features.
Abstract
Analyzing data subgroups is a common data science task to build intuition about a dataset and identify areas to improve model performance. However, subgroup analysis is prohibitively difficult in datasets with many features, and existing tools limit unexpected discoveries by relying on user-defined or static subgroups. We propose exploratory subgroup analysis as a set of tasks in which practitioners discover, evaluate, and curate interesting subgroups to build understanding about datasets and models. To support these tasks we introduce Divisi, an interactive notebook-based tool underpinned by a fast approximate subgroup discovery algorithm. Divisi's interface allows data scientists to interactively re-rank and refine subgroups and to visualize their overlap and coverage in the novel Subgroup Map. Through a think-aloud study with 13 practitioners, we find that Divisi can help uncover…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
