Apress Cyber Monday SALE

Pro Apache Hadoop

2nd Edition

By Sameer Wadkar , Madhu Siddalingaiah , Jason Venner

  • eBook Price: $31.99
  • Print Book Price: $44.99
Buy eBook Buy Print Book
Pro Apache Hadoop, Second Edition brings you up to speed on Hadoop – the framework of big data – helping you build resilient and reliant compute clusters capable of analyzing large volumes of data in amazingly short times.

Full Description

  • Add to Wishlist
  • ISBN13: 978-1-4302-4863-7
  • 444 Pages
  • User Level: Intermediate to Advanced
  • Publication Date: September 15, 2014
  • Available eBook Formats: EPUB, MOBI, PDF

Related Titles

  • Pro Couchbase Server
  • Next Generation Databases
  • Big Data Analytics with Spark
  • Beginning Big Data with Power BI and Excel 2013
Full Description

Pro Apache Hadoop, Second Edition brings you up to speed on Hadoop – the framework of big data. Revised to cover Hadoop 2.0, the book covers the very latest developments such as YARN (aka MapReduce 2.0), new HDFS high-availability features, and increased scalability in the form of HDFS Federations. All the old content has been revised too, giving the latest on the ins and outs of MapReduce, cluster design, the Hadoop Distributed File System, and more.

This book covers everything you need to build your first Hadoop cluster and begin analyzing and deriving value from your business and scientific data. Learn to solve big-data problems the MapReduce way, by breaking a big problem into chunks and creating small-scale solutions that can be flung across thousands upon thousands of nodes to analyze large data volumes in a short amount of wall-clock time. Learn how to let Hadoop take care of distributing and parallelizing your software—you just focus on the code; Hadoop takes care of the rest.

  • Covers all that is new in Hadoop 2.0
  • Written by a professional involved in Hadoop since day one
  • Takes you quickly to the seasoned pro level on the hottest cloud-computing framework

What you’ll learn

  • Build a resilient and scalable Hadoop compute cluster.
  • Analyze large volumes of data in amazingly short time.
  • Optimize Hadoop tasks like a seasoned professional.
  • Implement bulletproof patterns that are proven successful.
  • Scale out using the new HDFS Federations feature set.
  • Chunk large problems into highly-parallel, MapReduce modules

Who this book is for

This book is aimed at I.T. professionals investigating Hadoop and implementing it in their organizations. Existing Hadoop users will deepen their toolkits and come up to speed on what’s new Hadoop 2.0. New Hadoop users will quickly move to the seasoned professional level in their use of the toolset.

Table of Contents

Table of Contents

1. Motivation for Big Data

2. Hadoop Concepts

3. Getting Started with the Hadoop Framework

4. Hadoop Administration

5. Basics of MapReduce Development

6. Advanced MapReduce Development

7. Hadoop Input Output

8. Testing Hadoop Programs

9. Monitoring Hadoop

10. Data Warehousing using Hadoop

11. Data Processing using Pig

12. HCatalog and Hadoop in the Enterprise

13. Log Analysis using Hadoop

14. Building Real-Time Systems using HBase

15. Data Science With Hadoop

16. Hadoop in the Cloud

17. Building a YARN Application

18. Appendix A

19. Appendix B

20. Appendix C

Source Code/Downloads

Downloads are available to accompany this book.

Your operating system can likely extract zipped downloads automatically, but you may require software such as WinZip for PC, or StuffIt on a Mac.


If you think that you've found an error in this book, please let us know by emailing to editorial@apress.com . You will find any confirmed erratum below, so you can check if your concern has already been addressed.
No errata are currently published


    1. The Definitive Guide to MongoDB: Third Edition


      View Book

    2. Digital Asset Management


      View Book

    3. Pro Apache Hadoop


      View Book

    4. Beginning CouchDB


      View Book