Knowledge-Based Learning through Feature Generation

Michal Badian; Shaul Markovitch

arXiv:2006.03874·cs.LG·June 9, 2020

Knowledge-Based Learning through Feature Generation

Michal Badian, Shaul Markovitch

PDF

Open Access

TL;DR

This paper presents a novel feature generation algorithm that leverages auxiliary datasets to improve learning performance, especially in small data scenarios, by inducing new features from external knowledge sources.

Contribution

The paper introduces a new algorithm for generating features using auxiliary datasets without assuming task similarity, enhancing learning with external knowledge.

Findings

01

Significant performance improvements in text classification tasks.

02

Effective feature generation from diverse auxiliary datasets.

03

Enhanced generalization in small sample learning scenarios.

Abstract

Machine learning algorithms have difficulties to generalize over a small set of examples. Humans can perform such a task by exploiting vast amount of background knowledge they possess. One method for enhancing learning algorithms with external knowledge is through feature generation. In this paper, we introduce a new algorithm for generating features based on a collection of auxiliary datasets. We assume that, in addition to the training set, we have access to additional datasets. Unlike the transfer learning setup, we do not assume that the auxiliary datasets represent learning tasks that are similar to our original one. The algorithm finds features that are common to the training set and the auxiliary datasets. Based on these features and examples from the auxiliary datasets, it induces predictors for new features from the auxiliary datasets. The induced predictors are then added to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning and Data Classification · Text and Document Classification Technologies · Topic Modeling