Activity #3748
openModule #3693: Big Data Analytics Platform and Tools for Epidemic Forecasting and Mitigation
Module #3696: Center for Epidemic Forecsting(EPIFORM)
Exploration and Finalization of Efficient Hardware Requirements for Big Data Technologies
100%
Description
Studied the hardware requirements for setting up a reliable and scalable Big Data lab.
Focused on configurations needed for Hadoop Distributed File System (HDFS) and Apache Spark in a multi-node cluster environment.
Identified the Challenge
Recognized that determining optimal hardware configurations for Big Data technologies is complex due to:
Varying workloads (batch, streaming, ML)
Rapidly evolving technologies
Budget and space constraints in lab environments
Research and Literature Review
Reviewed technical papers, whitepapers, and industry reports on:
Hadoop and Spark cluster deployment best practices
Performance bottlenecks in distributed storage and processing
Hardware configurations used in research and commercial environments
Studied comparative analyses of disk I/O, network bandwidth, and memory utilization across node types.
Comparison with HPC Systems
Compared Big Data cluster needs with traditional High-Performance Computing (HPC) hardware:
Big Data requires scalable storage, fault tolerance, and commodity hardware
Noted that Big Data clusters benefit more from horizontal scaling (more nodes) than vertical scaling (stronger single nodes)
Identified Role-Based Hardware Profiles
Finalized a cost-efficient, scalable cluster design for Big Data education and experimentation.
Ensured compatibility with Hadoop, Spark, Delta Lake, NiFi, and future ML tools.
Built confidence in sizing nodes appropriately without over-provisioning, while maintaining upgrade paths.