Big Data Infrastructure I. The Scale-Out Paradigm & Distributed Storage — vertical vs. horizontal scaling, the Google File System (GFS) and HDFS.
M.Sc. in Computer Science · Sapienza University of Rome
Academic Year 2026/27 — 6 CFU
The course introduces big data computing as an advanced structural paradigm of modern computing platforms,
at the intersection of distributed computing architectures, high-dimensional data
representations, and scalable machine learning systems. Students learn how the
scale-out model handles volume, velocity and variety of data, and the methodological abstractions needed
where classical exact computation becomes intractable.
COURSE START: The first lecture is scheduled for Wednesday, September 23.
Enrolling on the course's Google Classroom page is highly recommended to stay up to date on all course-related communications.
Enroll to ClassroomAlways use the following pattern in the email subject line:
The course introduces big data computing as an advanced structural paradigm of modern computing platforms, at the intersection of distributed computing architectures, high-dimensional data representations, and scalable machine learning systems. Students learn how the scale-out model handles volume, velocity and variety of data, and the methodological abstractions needed where classical exact computation becomes intractable.
Required
Recommended
Free online
Mining of Massive Datasets — Leskovec, Rajaraman & Ullman (freely available online)
Foundations of Data Science — Blum, Hopcroft & Kannan (freely available online)
Research Seminar — mandatory
A critical seminar on a recent, high-impact research paper, chosen autonomously by the student and approved by the instructor, from top-tier venues (e.g., SIGMOD, VLDB, OSDI, SOSP, NeurIPS, ICML, KDD). Honors (30 cum laude) require exceptional mastery, deep critical insight, and flawless academic communication.
Big Data Infrastructure I. The Scale-Out Paradigm & Distributed Storage — vertical vs. horizontal scaling, the Google File System (GFS) and HDFS.
Big Data Infrastructure II. The MapReduce Execution Model — Map, Shuffle, Sort and Reduce phases; straggler tracking and fault tolerance.
Material and announcements from previous editions of the course.