Our UK training centre is reopening in June. Learn more about it on our blog.

Cloudera CCA Spark and Hadoop Developer Certification

- Only 3 Days

Learn how to import data into an Apache Hadoop cluster and process it using modern data analysis tools like Spark, Flume, Hive, Impala, Sqoop and more – in just 3 days.

On this accelerated CCA Spark and Hadoop Developer course, you’ll identify which data analysis tools to use in a given situation and gain hands-on development experience using those tools.

Experience Firebrand’s Lecture | Lab | Review methodology as you prepare for the real world challenges faced by Hadoop developers. You’ll learn:

Read more...

  • How data is distributed, stored, and processed in a Hadoop cluster
  • Data distribution in Apache Spark
  • How to use Sqoop and Flume to ingest and model data as tables

You’ll also learn how to choose the best data storage format and study best practices for data storage.

Plus, you’ll sit the CCA Spark and Hadoop Developer exam (CCA175) during your accelerated course. This exam is covered by your Certification Guarantee.  

See Benefits...

See prices now to find out how much you could save when you train at twice the speed.

Seven reasons why you should sit your course with Firebrand Training

  1. Two options of training. Choose between residential classroom-based, or online CCA courses
  2. You'll be CCA certified in just 3 days. With us, you’ll be CCA trained in record time
  3. Our CCA course is all-inclusive. A one-off fee covers all course materials, exams, accommodation and meals. No hidden extras
  4. Pass CCA first time or train again for free. This is our guarantee. We’re confident you’ll pass your course first time. But if not, come back within a year and only pay for accommodation, exams and incidental costs
  5. You’ll learn more. A day with a traditional training provider generally runs from 9am – 5pm, with a nice long break for lunch. With Firebrand Training you’ll get at least 12 hours/day quality learning time, with your instructor
  6. You’ll learn CCA faster. Chances are, you’ll have a different learning style to those around you. We combine visual, auditory and tactile styles to deliver the material in a way that ensures you will learn faster and more easily
  7. You’ll be studying CCA with the best. We’ve been named in Training Industry’s “Top 20 IT Training Companies of the Year” every year since 2010. As well as winning many more awards, we’ve trained and certified 75,595 professionals, and we’re partners with all of the big names in the business

Think you are ready for the course? Take a FREE practice test to assess your knowledge!

Benefits of Training with Firebrand

  • Two options of training - Residential classroom-based, or online courses
  • A purpose-built training centre – get access to dedicated Pearson VUE Select facilities
  • Certification Guarantee – pass first time or train again free (just pay for accommodation, exams and incidental costs)
  • Everything you need to certify – you’ll sit your exam on the course and return home certified
  • No hidden extras – one cost covers everything you need to certify

See Curriculum...

Introduction

  • Introduction to Hadoop and the Hadoop Ecosystem
  • Problems with traditional large scale systems
  • Hadoop!
  • Data storage and ingest
  • Data processing
  • Data analysis and exploration
  • Other ecosystem tools
  • Introduction to the hands-on exercises

Hadoop architecture and HDFS

  • Distributed processing on a cluster
  • Storage:
    • HDFS architecture
    • Using HDFS
  • Resource management:
    • YARN architecture
    • Working with YARN

Importing relational data with Apache Sqoop

  • Sqoop overview
  • Basic imports and exports
  • Limiting results
  • Improving Sqoop’s performance
  • Sqoop 2

Introduction to Impala and Hive

  • Introduction to Impala and Hive
  • Why use Impala and Hive?
  • Querying data With Impala and Hive
  • Comparing Hive and Impala to traditional databases

Modelling and managing data with Impala and Hive

  • Data storage overview
  • Creating databases and tables
  • Loading data into tables
  • HCatalog
  • Impala metadata caching

Data formats

  • Selecting a file format
  • Hadoop tool support for file formats
  • Avro schemas
  • Using Avro with Impala, Hive and Sqoop
  • Avro schema evolution
  • Compression

Data file partitioning

  • Partitioning overview
  • Partitioning in Impala and Hive
  • Capturing data with Apache Flume
  • What is Apache Flume?
  • Basic Flume architecture:
    • Flume sources
    • Flume sinks
    • Flume channels
    • Flume configuration

Spark basics

  • What is Apache Spark?
  • Using the Spark Shell
  • RDDs (Resilient Distributed Datasets)
  • Functional programming in Spark

Working with RDDs in Spark

  • Creating RDDs
  • Other general RDD operations

Writing and deploying Spark applications

  • Spark applications vs Spark Shell
  • Creating the SparkContext
  • Building a Spark application (Scala and Java)
  • Running a Spark application
  • The Spark application web UI
  • Configuring Spark properties
  • Logging

Parallel processing in Spark

  • Review: Spark on a Cluster
  • RDD partitions
  • Partitioning of File-based RDDs
  • HDFS and data locality
  • Executing parallel operations
  • Stages and tasks

Spark RDD persistence

  • RDD Lineage
  • RDD persistence overview
  • Distributed persistence
  • Common patterns in spark data processing
  • Common Spark use cases
  • Iterative algorithms in Spark
  • Graph processing and analysis
  • Machine learning
  • Example: k-means

DataFrames and Spark SQL

  • Spark SQL and the SQL context
  • Creating DataFrames
  • Transforming and querying DataFrames
  • Saving DataFrames
  • DataFrames and RDDs
  • Comparing Spark SQL, Impala, and Hive-on-Spark

See Exam Track...

You'll sit the following exam at the Firebrand Training Centre, covered by your Certification Guarantee:

  • CCA175 - CCA Spark and Hadoop Developer Exam

Additional details:

  • Number of questions: 10-12 performance-based (hands-on) tasks on a CDH5 cluster
  • Time limit: 120 minutes
  • Passing score: 70%

For each CCA question you must solve a particular scenario. You will be required to use tools like Hive and Impala, as well as coding in Scala or Python.

See What's Included...

Your accelerated course includes:

  • Accommodation *
  • Meals, unlimited snacks, beverages, tea and coffee *
  • On-site exams **
  • Exam vouchers **
  • Practice tests **
  • Certification Guarantee ***
  • Courseware
  • Up-to 12 hours of instructor-led training each day
  • 24-hour lab access
  • Digital courseware **
  • * For residential training only. Doesn't apply for online courses
  • ** Some exceptions apply. Please refer to the Exam Track or speak with our experts
  • *** Pass first time or train again free (just pay for accommodation, exams and incidental costs)

See Prerequisites...

You should be a developer/engineer with programming experience in Scala or Python. Basic knowledge of Linux command line and SQL is also recommended.

Prior experience with Hadoop is not required for this accelerated course.

Unsure whether you meet the prerequisites? Don’t worry. Your training consultant will discuss your background with you to understand if this course is right for you.

See Dates...

Cloudera CCA Course Dates

Start

Finish

Status

Location

Book now

24/2/2020 (Monday)

26/2/2020 (Wednesday)

Finished

-

 

29/6/2020 (Monday)

1/7/2020 (Wednesday)

Wait list

Nationwide

 

10/8/2020 (Monday)

12/8/2020 (Wednesday)

Limited availability

Nationwide

 

21/9/2020 (Monday)

23/9/2020 (Wednesday)

Open

Nationwide

 

2/11/2020 (Monday)

4/11/2020 (Wednesday)

Open

Nationwide

 

14/12/2020 (Monday)

16/12/2020 (Wednesday)

Open

Nationwide

 

Here's the Firebrand Training review section. Since 2001 we've trained exactly 75,595 students and asked them all to review our Accelerated Learning. Currently, 96.76% have said Firebrand exceeded their expectations.

Read reviews from recent accelerated courses below or visit Firebrand Stories for written and video interviews from our alumni.


"Course was very good, I liked the analogies given to explain complex networking areas it really helped. The instructor was very good at bringing information to a level everyone could understand and then further on that. "
Ewan jamieson, Mr. (17/2/2020 to 22/2/2020)

"Knowledgeable, well spoken and relatable instructor. The instructor explained difficult concepts clearly and provided good examples to help us understand the material. He also discussed real-world examples to help solidify how the concepts are applied in the field. There were a few technical issues with the remote learning and remote labs, but overall a great experience, especially so when considering the current situation with COVID-19."
JE, Mitchells & Butlers. (4/5/2020 to 7/5/2020)

"I would highly recommend taking up a course with Firebrand because the way the course is delivered is very helpful and useful for someone who either has no knowledge or some knowledge. The instructors do everything they can to ensure everyone is learning and is always on hand to help anyone who is struggling with anything. Everyone on the course is also very helpful and in the same position as they want to learn to. So there is not just the instructor but the class as a whole to assist. The facility is very nice and I will definitely be coming back for any future courses I attend. "
CP, Innovate Ltd. (4/5/2020 to 7/5/2020)

"Firebrand have very passionate instructors."
Anonymous (4/5/2020 to 7/5/2020)

"Extremely useful course, detailed understanding and trainer had good grasp on the information and was able to explain it in different ways to help our understanding and help us to relate this to our own working environment."
SA, Lloyds Banking Group. (4/5/2020 to 6/5/2020)

Latest Reviews from our students