Feature Extraction and Feature Selection: Reducing Data Complexity with   Apache Spark

Dimitrios Sisiaridis; Olivier Markowitch

arXiv:1712.08618·cs.DB·December 27, 2017

Feature Extraction and Feature Selection: Reducing Data Complexity with Apache Spark

Dimitrios Sisiaridis, Olivier Markowitch

PDF

TL;DR

This paper presents a scalable approach for feature extraction and selection in security analytics of heterogeneous network data using Apache Spark's pyspark API, aiming to improve efficiency in cyber threat detection.

Contribution

The paper introduces a novel implementation of feature extraction and selection tailored for heterogeneous security data using Apache Spark, enhancing processing efficiency.

Findings

01

Efficient handling of heterogeneous data sources.

02

Implementation in Apache Spark with pyspark.

03

Improved processing speed for security analytics.

Abstract

Feature extraction and feature selection are the first tasks in pre-processing of input logs in order to detect cyber security threats and attacks while utilizing machine learning. When it comes to the analysis of heterogeneous data derived from different sources, these tasks are found to be time-consuming and difficult to be managed efficiently. In this paper, we present an approach for handling feature extraction and feature selection for security analytics of heterogeneous data derived from different network sensors. The approach is implemented in Apache Spark, using its python API, named pyspark.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.