Applications for the September 2026 Qualifier will open on June 29, 2026. Notify me
Degree Level Course

Introduction to Big Data

This course will introduce students to practical aspects of analytics at a large scale, i.e. big data. The course will start with a basic introduction to big data and cloud concepts spanning hardware, systems and software, and then delve into the details of algorithm design and execution at large scale.

Code BSDA5001
Credits 4 Credits
Type Elective
Prerequisites None
Core Competencies

What You'll Learn

View Course Videos
  • This course will introduce students to practical aspects of analytics at a large scale, i.e. big data. The course will start with a basic introduction to big data and cloud concepts spanning hardware, systems and software, and then delve into the details of algorithm design and execution at large scale.

  • Introduction to Cloud Concepts: Cloud-Native architecture, serverless computing, message queues, PaaS, SaaS, IaaS

  • Introduction to Big Data concepts: divide- and-conquer, parallel algorithms, distributed virtualized storage, distributed resource management, real-time processing.

  • Data Processing Fundamentals: data formats, sources and their semantics, processing patterns for large data (the ETL vs ELT difference), processing + storage options on cloud, lakehouse architecture

  • Technology deep-dive on Open Source as well as Google Cloud

  • echnologies covered: Spark (PySpark, Spark ML, Spark Streaming), SQL (SparkSQL), Kafka, Google Pub/Sub, Google Dataproc, Google Cloud Functions

12-Week Roadmap

Course Structure & Syllabus

For details of standard term assessment timelines and exam structures, visit our Academics page.

WEEK 1
Introduction: ​Big data concepts & GCP Platform Setup
WEEK 2
Cloud concepts​: ​Cloud-Native architecture, serverless computing, message queues, PaaS, SaaS, IaaS
WEEK 3
Types of Data​:​ Data formats, sources & their semantics, processing & storage options on Cloud. Use of serverless to get started (e.g. Google Cloud Functions)
WEEK 4
Intro to Big Data Engineering​:​ Hadoop and PySpark
Faculty & Experts

About the Instructors

Rangarajan Vasudevan

Rangarajan Vasudevan

Co-Founder & Chief Data Officer , Lentra.ai

Rangarajan Vasudevan is the Co-Founder & CDO of Lentra.ai, India’s fastest growing lending cloud. He did “big data” & “data science” before it was fashionable, building data-native applications across industries and geographies over 15+ years.

Ranga joined Lentra by way of an acquisition in June 2022 of his company TheDataTeam, creators of Cadenz.ai customer intelligence platform. Prior to founding TheDataTeam, Ranga served as Director, Big Data with Teradata Corporation’s international business unit. Ranga joined Teradata via the acquisition of Aster Data Systems, where he was a founding engineer and co-invented a company-defining, patented, pattern recognition algorithm. He is a recipient of both the Distinguished Engineer (R&D) and Consulting Excellence awards while at Teradata.

Ranga has degrees in Computer Science from the University of Michigan and IIT Madras.