BIG DATA AND SMART DATA ANALYTICS

Irene Finocchi, Fabio Angeletti

Obiettivi formativi

As datasets grow to Petabyte scale, traditional analysis models and computation paradigms become obsolete. The course will focus on fundamental algorithmic and programming issues posed by big-data analytics, tackling major problems and techniques for extracting knowledge from massive amounts of data. By the end of the course the students will gain an understanding of modern methods and systems for big data analytics, also in a distributed setting.

Prerequisiti

Computer programming skills (knowledge of Python is required for the project work). Basic knowledge of fundamental algorithms.

Risultati di apprendimento attesi

Knowledge and understanding: Upon successful completion of the course, the students will be familiar with data mining problems and techniques and computational models and frameworks for analyzing and extracting insights from massive, possibly distributed or rapidly changing amounts of data at a large scale. Applying knowledge and understanding: After this course, the students will be able to develop efficient data analytics solutions. They will be also able to implement the proposed solutions on top of industry-standard frameworks, e.g., Apache Spark, in order to tackle real-world problems such as those typically faced by big tech companies. Making judgements: Throughout the entire course, students will be invited to assess critically strengths and weaknesses of all the different methods and tools presented in class. After this course, they will be able to analyze different solutions to big data problems and to demonstrate an in-depth, critical understanding of the scope and challenges of different data-driven analytics techniques. Communication skills: This course will give the students the possibility to acquire and to understand major terms and concepts so as to communicate effectively their ideas, findings, proposals, analysis, and critical reasoning in the area of data-driven analytics. A special emphasis will be given to oral presentations and pitches in project group works. Learning skills: This course will provide the students with the ability to learn cutting-edge design and analysis tools and to apply them to real-world data analytics problems. The method of study will make the students able to break down complex problems arising in specific applications into manageable pieces and to apply different patterns in order to design rigorous and documentable solutions. A strong emphasis will be given to the direct application of the techniques and tools covered in this course to complex problems that are typical of today’s data-driven industry.

Contenuti Del Corso

Big data characteristics, architectures, and analytics workflows. Distributed computing models and algorithm design for large-scale data. MapReduce and Apache Spark for scalable data processing. Data acquisition, storage, management, and querying at scale. Mining and analyzing massive networks. Scalable data analytics in cloud environments. Real-time analytics of streaming data.

Testi Di Riferimento

Mining of Massive Datasets. HYPERLINK "https://www.amazon.com/s/ref=dp_byline_sr_ebooks_1?ie=UTF8&field-author=Jure+Leskovec&text=Jure+Leskovec&sort=relevancerank&search-alias=digital-text" Jure Leskovec, Anand Rajaraman, Jeffrey David Ullman. Second edition. Apache SparkTM manual. Lecture notes, research papers, and other course material made available on the e-learning platform.

Metodologie Didattiche

The course consists of traditional lectures complemented by hands-on lab sessions and industrial testimonials, that will guide the students on the use of good analytics practices and industry-standard practices.

Modalità di verifica dell'apprendimento

There will be a written test, a research paper presentation, and a group software project. Non-attending students: oral exam on the course contents + discussion of a topic selected with the professor.

Criteri per l’assegnazione dell’elaborato finale

Thesis assigned (upon specific request to the professor) to students who demonstrate a serious and motivated interest in the course topics.

Settimana 1

Introduction. Big data characteristics. Big data systems. Hw/sw stack for big data.

Settimana 2

The MapReduce programming model. Software design principles in distributed big data systems. Lab: from MapReduce to Apache Spark.

Settimana 3

Analyzing distributed computations: key metrics. Designing efficient algorithms in MapReduce. Exercises. Lab: hands-on Apache Spark.

Settimana 4

Data acquisition, storage, and management.

Settimana 5

Mining massive networks. Example: finding dense structures in distributed settings.

Settimana 6

Big data databases and data management. Lab: Spark SQL

Settimana 7

Big data databases and data management. Lab: Spark SQL

Settimana 8

Midterm. Industry testimonial: data analytics in practice.

Settimana 9

Scalable data analytics in the cloud. Lab.

Settimana 10

Real-time analytics over data streams. Lab.

Settimana 11

Course recap. Research directions. Lab.

Settimana 12

Research paper presentations.