NYU Tandon · Fall 2026

Big Data Management & Analysis

Learn how to collect, manage, analyze, and communicate insights from data at scale—through practical tools, thoughtful questioning, and hands-on projects.

CUSP-GX 8083 / ECE-GY 9113

From raw data to useful answers.

This course is an introduction to the principles, technologies, and challenges of big data management and analysis. We will work across the full data lifecycle: framing questions, collecting data, preparing it for analysis, choosing appropriate technologies, and presenting findings clearly.

Core technologies

SQL, NoSQL, MapReduce, Spark, and modern data platforms.

Practical techniques

Data preprocessing, querying, benchmarking, visualization, and text analysis.

Real applications

Hands-on work grounded in real-world datasets, distributed computing, and machine learning.

Learning outcomes

  • Compare core technologies for managing large-scale data.
  • Choose and optimize data-management approaches for different scenarios.
  • Extract insights, visualize results, and communicate a convincing data story.

Prerequisites

  • Basic Python knowledge.
  • Basic data-analysis skills, such as working with spreadsheets.
  • A willingness to investigate messy, real-world data.
↑ Back to top

A high-level roadmap.

We meet every Thursday from 6:00 to 8:30 p.m. Topics and activities may evolve as the class progresses; readings, materials, and project details will be shared during the semester.

Attend in person or online

You may join class in person or online—no advance notice is needed. Use the class document to access the Zoom link.

Participation is essential

Ask thoughtful questions and join the discussion. After a good question or response, use the class document to access the participation form and record your contribution.

Upcoming classes

September 10
End-to-end data analysis — practical analysis from raw data to insight.
September 17
Data pipelines and visualization — cleaning, aggregation, and communicating results.
September 24
NoSQL and data modeling — flexible data structures, databases, and trade-offs.
October 1
Web data collection — scraping, data quality, and ethical considerations.
October 8
Storage and performance — formats, indexes, benchmarking, and scale.
October 15
Data analysis and project development — exploration, validation, and project ideas.
October 22
Distributed computing foundations — MapReduce, Spark, and scalable processing.
October 29
Spark in practice — building and optimizing distributed data workflows.
November 5
Text data and data challenges — unstructured data and large-scale analysis.
November 12
LLMs and urban analytics — guest perspectives and emerging applications.
November 19
AI/ML for big data — modern methods, evaluation, and responsible use.
November 26
No class — Thanksgiving recess.
December 3
Current big data trends — streaming, modern platforms, and project progress.
December 10
Final project presentations — findings, methods, and lessons learned.

Past classes

September 3
Introduction to big data Course foundations, questions, and data realities.
  • Interactive format: live demos, Mentimeter, and real-time questions
  • Big data beyond scale: people, documentation, policy, and privacy
  • Data cleaning, labeling, provenance, and validation as central work
  • Data cascades: early collection errors and bias amplified downstream
  • Storage trade-offs: files, databases, single powerful machines, and distributed systems
  • AI encouraged—with critical judgment and human oversight
  • Course structure: participation, one open-resource quiz, data challenge, final project
↑ Back to top

Applied learning, assessed broadly.

Assessment follows the same practical spirit as the 2025 course: active participation, a contextual quiz, a data challenge, and a final project. Exact point allocations for components other than the quiz will be announced during the semester.

ComponentWeightHow it is assessed
Class participationTBDThoughtful questions, contributions, and engagement in class.
Online pollsTBDSynchronous participation in occasional in-class polls.
In-class quiz~30%One quiz; date TBD. Scenario-based and focused on course concepts.
Big Data ChallengeTBDA timed data challenge emphasizing accuracy, efficiency, and scalable code.
Final projectTBDA collaborative, original data-driven project and final presentation.
Extra creditTBDOpportunities, if offered, will be announced during the semester.
Participation matters. Danny cares deeply about active participation, especially asking thoughtful questions during class. Use the class document to access the participation form and record a good question or response after you contribute.
About the quiz. The quiz will be completed synchronously in class. It will use practical scenarios discussed during the course and will reward understanding the main ideas, not memorization. Details and the date will be announced later.

Big Data Challenge

Work individually or in a small group to solve data tasks. Submissions will be assessed for correctness and performance, with opportunities to refine and resubmit.

Final project

Develop an original, data-driven project. You will document your data pipeline, reflect on data quality and bias, and present a clear story with your findings.

↑ Back to top

Ask early. Build together.

Course communication

Use the class Slack workspace for questions, project discussion, and sharing useful resources. An invitation and course materials will be provided before the semester begins.

Office hours

Danny Y. Huang will be available after class on Thursdays, 8:30–9:00 p.m., in person or on Zoom using the class link.

↑ Back to top