This four-day hands-on training course delivers the key concepts and knowledge developers need to use Apache Spark to develop highperformance, parallel applications on the Cloudera Data Platform (CDP). Hands-on exercises allow students to practice writing Spark applications that integrate with CDP core components. Participants will learn how to use Spark SQL to query structured data, how to use Hive features to ingest and denormalize data, and how to work with “big data” stored in a distributed file system. After taking this course, participants will be prepared to face real-world challenges and build applications to execute faster decisions, better decisions, and interactive analysis, applied to a wide variety of use cases, architectures, and industries.
What Skills You Will Gain During this course, you will learn how to:
• Distribute, store, and process data in a CDP cluster
• Write, configure, and deploy Apache Spark applications
• Use the Spark interpreters and Spark applications to explore, process, and analyze distributed data
• Query data using Spark SQL, DataFrames, and Hive tables
• Deploy a Spark application on the Data Engineering Service
Download DENG-254: Preparing with Cloudera Data Engineering 250624
| Available Options: |
|   |  
|
| Include Exam Voucher: | |
Upon completion of the training, you will receive a Training Certificate of Completion.