M.Sc. in Computer Science · Sapienza University of Rome

Big Data Computing

Academic Year 2026/27 — 6 CFU
The course introduces big data computing as an advanced structural paradigm of modern computing platforms, at the intersection of distributed computing architectures, high-dimensional data representations, and scalable machine learning systems. Students learn how the scale-out model handles volume, velocity and variety of data, and the methodological abstractions needed where classical exact computation becomes intractable.

COURSE START: The first lecture is scheduled for Wednesday, September 23.

Lecture Schedule

Wednesday

8:00 – 11:00

Room 1L — Via del Castro Laurenziano 7a (RM018) (map).

Thursday

10:00 – 12:00

Room 2L — Via del Castro Laurenziano 7a (RM018) (map).

Google Classroom

Enrolling on the course's Google Classroom page is highly recommended to stay up to date on all course-related communications.

Enroll to Classroom

Contacts

Office Room 106, 1st floor, Building E — Viale Regina Elena 295
Office Hours By appointment via email

Always use the following pattern in the email subject line:

[BDC 2026/27]: <SUBJECT>

Learning Objectives

Distributed Computing High-Dimensional Data Scalable ML

The course introduces big data computing as an advanced structural paradigm of modern computing platforms, at the intersection of distributed computing architectures, high-dimensional data representations, and scalable machine learning systems. Students learn how the scale-out model handles volume, velocity and variety of data, and the methodological abstractions needed where classical exact computation becomes intractable.

Prerequisites

Required

  • Algorithms and data structures
  • Operating systems
  • Basics of distributed systems
  • Introductory ML mathematics (linear algebra, calculus, basic optimization)

Recommended

  • Familiarity with data analytics concepts
  • Concurrent programming abstractions

Course Contents — 4 Modules

1. Big Data Infrastructure — scale-out storage, fault tolerance via lineage, MapReduce, memory-centric DAG execution (Spark), streaming.
2. High-Dimensional Data Representations — geometry of high-dimensional spaces, dimensionality reduction (Johnson–Lindenstrauss), columnar layouts, text/image/graph embeddings.
3. Classical Non-Learning Tasks — web-scale indexing, sampling theory, Approximate Query Processing, sub-linear sketching (HyperLogLog, Count-Min), Locality-Sensitive Hashing.
4. Scale-Out Learning Systems — large-scale ERM, parallel/asynchronous SGD, Parameter Server, Federated Learning, memory-bandwidth optimization.

Textbooks & References

Free online

Mining of Massive Datasets — Leskovec, Rajaraman & Ullman (freely available online)

Foundations of Data Science — Blum, Hopcroft & Kannan (freely available online)

Designing Data-Intensive Applications — Kleppmann — systems and architecture reference.
Introduction to Information Retrieval (Manning, Raghavan & Schütze), Understanding Machine Learning (Shalev-Shwartz & Ben-David) — plus canonical research papers referenced throughout the lecture material.

Assessment Method

Research Seminar — mandatory

oral presentation 100% of final grade min. 18/30 to pass

A critical seminar on a recent, high-impact research paper, chosen autonomously by the student and approved by the instructor, from top-tier venues (e.g., SIGMOD, VLDB, OSDI, SOSP, NeurIPS, ICML, KDD). Honors (30 cum laude) require exceptional mastery, deep critical insight, and flawless academic communication.

Part I — BIG DATA INFRASTRUCTURE

Lecture 1.1 2026/09/23

Big Data Infrastructure I. The Scale-Out Paradigm & Distributed Storage — vertical vs. horizontal scaling, the Google File System (GFS) and HDFS.

Lecture 1.2 2026/09/24

Big Data Infrastructure II. The MapReduce Execution Model — Map, Shuffle, Sort and Reduce phases; straggler tracking and fault tolerance.

Part II — HIGH-DIMENSIONAL REPRESENTATIONS

Part III — SCALING CLASSICAL TASKS

Part IV — SCALING LEARNING TASKS

Material and announcements from previous editions of the course.