Project

General

Profile

Actions

Activity #3743

open

Module #3693: Big Data Analytics Platform and Tools for Epidemic Forecasting and Mitigation

Module #3696: Center for Epidemic Forecsting(EPIFORM)

Spark Technology Exploration

Added by Prasidh J S about 1 year ago.

Status:
Resolved
Priority:
Normal
Assignee:
Start date:
05/01/2025
Due date:
% Done:

100%

Estimated time:
Planned Due Date:

Description

Studied Apache Spark architecture: Driver, Executors, Cluster Manager.

Explored Spark modules: Core, SQL, Streaming, MLlib, GraphX.

Understood DAG, lazy evaluation, and RDD/DataFrame transformations.

Hadoop vs Spark – Key Differences
Spark supports in-memory computing, faster than disk-based Hadoop MapReduce.
Spark handles batch + streaming, Hadoop supports batch only.
Spark supports high-level APIs in Python, Scala; easier development.
Spark offers built-in machine learning and graph processing libraries.

No data to display

Actions

Also available in: Atom PDF