Skip to main content
  • Book
  • © 2015

Big Data Analytics with Spark

A Practitioner's Guide to Using Spark for Large Scale Data Analysis

Apress

Authors:

  • Highlights the role of the Spark in the Big Data Landscape and proves the worth of having strong Spark skills in order to increase subject knowledge and boost your career
  • Introduces readers to the concept of functional programming in Scala, one of the core languages supported by Spark
  • Key fundamental concepts of Spark are highlighted through practical examples for different scenarios, each showcasing the features of Spark's core and its add-on libraries, Spark SQL, Spark Streaming, GraphX, and MLlib and so on

Buy it now

Buying options

eBook USD 44.99
Price excludes VAT (USA)
  • Available as EPUB and PDF
  • Read on any device
  • Instant download
  • Own it forever
Softcover Book USD 59.99
Price excludes VAT (USA)
  • Compact, lightweight edition
  • Dispatched in 3 to 5 business days
  • Free shipping worldwide - see info

Tax calculation will be finalised at checkout

Other ways to access

This is a preview of subscription content, log in via an institution to check for access.

Table of contents (12 chapters)

  1. Front Matter

    Pages i-xxiii
  2. Big Data Technology Landscape

    • Mohammed Guller
    Pages 1-15
  3. Programming in Scala

    • Mohammed Guller
    Pages 17-33
  4. Spark Core

    • Mohammed Guller
    Pages 35-61
  5. Interactive Data Analysis with Spark Shell

    • Mohammed Guller
    Pages 63-70
  6. Writing a Spark Application

    • Mohammed Guller
    Pages 71-78
  7. Spark Streaming

    • Mohammed Guller
    Pages 79-102
  8. Spark SQL

    • Mohammed Guller
    Pages 103-152
  9. Machine Learning with Spark

    • Mohammed Guller
    Pages 153-205
  10. Graph Processing with Spark

    • Mohammed Guller
    Pages 207-230
  11. Cluster Managers

    • Mohammed Guller
    Pages 231-242
  12. Monitoring

    • Mohammed Guller
    Pages 243-264
  13. Bibliography

    • Mohammed Guller
    Pages 265-267
  14. Back Matter

    Pages 269-277

About this book

Big Data Analytics with Spark is a step-by-step guide for learning Spark, which is an open-source fast and general-purpose cluster computing framework for large-scale data analysis. You will learn how to use Spark for different types of big data analytics projects, including batch, interactive, graph, and stream data analysis as well as machine learning. In addition, this book will help you become a much sought-after Spark expert.

Spark is one of the hottest Big Data technologies. The amount of data generated today by devices, applications and users is exploding. Therefore, there is a critical need for tools that can analyze large-scale data and unlock value from it. Spark is a powerful technology that meets that need. You can, for example, use Spark to perform low latency computations through the use of efficient caching and iterative algorithms; leverage the features of its shell for easy and interactive Data analysis; employ its fast batch processing and low latency features to process your real time data streams and so on. As a result, adoption of Spark is rapidly growing and is replacing Hadoop MapReduce as the technology of choice for big data analytics.

This book provides an introduction to Spark and related big-data technologies. It covers Spark core and its add-on libraries, including Spark SQL, Spark Streaming, GraphX, and MLlib. Big Data Analytics with Spark is therefore written for busy professionals who prefer learning a new technology from a consolidated source instead of spending countless hours on the Internet trying to pick bits and pieces from different sources.

The book also provides a chapter on Scala, the hottest functional programming language, and the program that underlies Spark. You’ll learn the basics of functional programming in Scala, so that you can write Spark applications in it.

What's more, Big Data Analytics with Spark provides an introduction to other big data technologies thatare commonly used along with Spark, like Hive, Avro, Kafka and so on. So the book is self-sufficient; all the technologies that you need to know to use Spark are covered. The only thing that you are expected to know is programming in any language.

There is a critical shortage of people with big data expertise, so companies are willing to pay top dollar for people with skills in areas like Spark and Scala. So reading this book and absorbing its principles will provide a boost—possibly a big boost—to your career.

Reviews

“Programmers seeking to learn the Spark framework and its libraries will benefit greatly from this book. … The book is well written, with a good balance between presenting simple computer science concepts, such as functional programming, and introducing Scala, the Spark core language. … the book provides substantial information on cluster-based data analysis using Spark, a prominent framework used by data scientists. It is very nicely written, with interesting contemporary considerations and several source code examples.” (Andre Maximo, Computing Reviews, computingreviews.com, June, 2016)

About the author

Mohammed Guller is the principal architect at Glassbeam, where he leads the development of advanced and predictive analytics products. He is a big data and Spark expert. He is frequently invited to speak at big data–related conferences. He is passionate about building new products, big data analytics, and machine learning.


Over the last 20 years, Mohammed has successfully led the development of several innovative technology products from concept to release. Prior to joining Glassbeam, he was the founder of TrustRecs.com, which he started after working at IBM for five years. Before IBM, he worked in a number of hi-tech start-ups, leading new product development.


Mohammed has a master's of business administration from the University of California, Berkeley, and a master's of computer applications from RCC, Gujarat University, India.

Bibliographic Information

Buy it now

Buying options

eBook USD 44.99
Price excludes VAT (USA)
  • Available as EPUB and PDF
  • Read on any device
  • Instant download
  • Own it forever
Softcover Book USD 59.99
Price excludes VAT (USA)
  • Compact, lightweight edition
  • Dispatched in 3 to 5 business days
  • Free shipping worldwide - see info

Tax calculation will be finalised at checkout

Other ways to access