Please use this identifier to cite or link to this item: http://hdl.handle.net/123456789/910
Title: Fast Data Processing with Spark
Authors: Sankar, Krishna
Karau, Holden
Keywords: Fast Data Processing with Spark
Issue Date: Mar-2015
Publisher: Packt Publishing
Series/Report no.: 1250315;
Abstract: Apache Spark has captured the imagination of the analytics and big data developers, and rightfully so. In a nutshell, Spark enables distributed computing on a large scale in the lab or in production. Till now, the pipeline collect-store-transform was distinct from the Data Science pipeline reason- model, which was again distinct from the deployment of the analytics and machine learning models. Now, with Spark and technologies, such as Kafka, we can seamlessly span the data management and data science pipelines. We can build data science models on larger datasets, requiring not just sample data. However, whatever models we build can be deployed into production (with added work from engineering on the "ilities", of course). It is our hope that this book would enable an engineer to get familiar with the fundamentals of the Spark platform as well as provide hands-on experience on some of the advanced capabilities.
URI: http://hdl.handle.net/123456789/910
ISBN: 978-1-78439-257-4
Appears in Collections:E-Books

Files in This Item:
File Description SizeFormat 
Fast Data Processing with Spark Second Edition.pdf8.06 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.