Front Matter

Pages i-xxiii

PDF
MapReduce and Its Abstractions
- Balaswamy Vaddeman
Pages 1-20
Data Types
- Balaswamy Vaddeman
Pages 21-31
Grunt
- Balaswamy Vaddeman
Pages 33-40
Pig Latin Fundamentals
- Balaswamy Vaddeman
Pages 41-67
Joins and Functions
- Balaswamy Vaddeman
Pages 69-87
Creating and Scheduling Workflows Using Apache Oozie
- Balaswamy Vaddeman
Pages 89-101
HCatalog
- Balaswamy Vaddeman
Pages 103-113
Pig Latin in Hue
- Balaswamy Vaddeman
Pages 115-122
Pig Latin Scripts in Apache Falcon
- Balaswamy Vaddeman
Pages 123-136
Macros
- Balaswamy Vaddeman
Pages 137-145
User-Defined Functions
- Balaswamy Vaddeman
Pages 147-155
Writing Eval Functions
- Balaswamy Vaddeman
Pages 157-169
Writing Load and Store Functions
- Balaswamy Vaddeman
Pages 171-186
Troubleshooting
- Balaswamy Vaddeman
Pages 187-199
Data Formats
- Balaswamy Vaddeman
Pages 201-208
Optimization
- Balaswamy Vaddeman
Pages 209-223
Hadoop Ecosystem Tools
- Balaswamy Vaddeman
Pages 225-248
Back Matter

Pages 249-274

PDF

About this book

Learn to use Apache Pig to develop lightweight big data applications easily and quickly. This book shows you many optimization techniques and covers every context where Pig is used in big data analytics. Beginning Apache Pig shows you how Pig is easy to learn and requires relatively little time to develop big data applications.

The book is divided into four parts: the complete features of Apache Pig; integration with other tools; how to solve complex business problems; and optimization of tools.

You'll discover topics such as MapReduce and why it cannot meet every business need; the features of Pig Latin such as data types for each load, store, joins, groups, and ordering; how Pig workflows can be created; submitting Pig jobs using Hue; and working with Oozie. You'll also see how to extend the framework by writing UDFs and custom load, store, and filter functions. Finally you'll cover different optimization techniques such asgathering statistics about a Pig script, joining strategies, parallelism, and the role of data formats in good performance.

What You Will Learn

• Use all the features of Apache Pig
• Integrate Apache Pig with other tools
• Extend Apache Pig
• Optimize Pig Latin code
• Solve different use cases for Pig Latin

Who This Book Is For

All levels of IT professionals: architects, big data enthusiasts, engineers, developers, and big data administrators

Keywords

Authors and Affiliations

Hyderabad, India

Balaswamy Vaddeman

About the author

Balaswamy Vaddeman, Thinker, Blogger, Serious and Self-motivated Big data evangelist with 9 years of experience in IT and 4 years of experience in Big data space. My Big data experience covers multiple areas like delivery of analytical applications, product development, consulting, training, book reviews, hackathons and mentoring and helping people on forums. I have proved myself while delivering analytical applications in retail, banking and finance domain in 3 aspects (Development, Administration and Architecture) of Hadoop related technologies. At Startup Company, I had developed a Hadoop based product that was used for delivering of analytical applications without writing code.
In 2013 I had won Hadoop Hackathon event for Hyderabad conducted by Cloudwick technologies. Being top contributor at stackoverflow.com, I helped many people on big data at multiple websites like stackoverflow.com and quora.com. With so much passion on big data I went ahead as independenttrainer and consultant to train hundreds of people and to set big data teams in couple of companies.

Bibliographic Information

Book Title: Beginning Apache Pig
Book Subtitle: Big Data Processing Made Easy
Authors: Balaswamy Vaddeman
DOI: https://doi.org/10.1007/978-1-4842-2337-6
Publisher: Apress Berkeley, CA
eBook Packages: Professional and Applied Computing, Apress Access Books, Professional and Applied Computing (R0)
Softcover ISBN: 978-1-4842-2336-9Published: 16 December 2016
eBook ISBN: 978-1-4842-2337-6Published: 10 December 2016
Edition Number: 1
Number of Pages: XXIII, 274
Number of Illustrations: 34 b/w illustrations, 35 illustrations in colour
Topics: Open Source, Database Management, Data Storage Representation, Data Mining and Knowledge Discovery, Information Storage and Retrieval

Publish with us

Policies and ethics

Authors:

Sections

Buy it now

Buying options

Other ways to access

Table of contents (17 chapters)

Front Matter

Back Matter