Building Batch Data Analytics Solutions on AWS
Official partner
AWS
Course Description
Do you need to operate, and scale your big data environments by automating time-consuming tasks? The Building Batch Data Analytics Solutions on AWS cou-rse will help you learn how to build batch data analytics solutions using Amazon EMR to optimize cost and performance.
This course is part of the Building Modern Data Analytics Solutions on AWS collection of four, one-day, intermediate-level classroom training courses.
Course Summary
Module A: Overview of Data Analytics and the Data Pipeline
Module 1: Introduction to Amazon EMR
Module 2: Data Analytics Pipeline Using Amazon EMR: Ingestion and Storage
Module 3: High-Performance Batch Data Analytics Using Apache Spark on Amazon EMR
Module 4: Processing and Analyzing Batch Data with Amazon EMR and Apache Hive
Module 5: Serverless Data Processing
Module 6: Security and Monitoring of Amazon EMR Clusters
Module 7: Designing Batch Data Analytics Solutions
Module B: Developing Modern Data Architectures on AWS
Prerequisites for this course
• Students with a minimum one-year experience managing open-source data frameworks such as Apache Spark or Apache Hadoop will benefit from this course
• We suggest the AWS Hadoop Fundamentals course for those that need a refresher on Apache Hadoop
• We recommend that attendees of this course have :
Completed either AWS Technical Essentials or Architecting on AWS Completed either Building Data Lakes on AWS or Getting Started with AWS Glue
Audience for this course
This course is intended for: • Data platform engineers • Architects and operators who build and manage data analytics pipelines
