Loading...
Ebook details
Log in if you are interested in the contents of the item.
Serverless ETL and Analytics with AWS Glue. Design scalable data lakes, optimize ETL pipelines, and accelerate analytics on AWS - Second Edition
Noritaka Sekiyama, Albert Quiroga, Tomohiro Tanaka, Subramanya Vajiraya, Akira Ajisaka, Ishan Gaur
Loading...
EBOOK
Loading...
Building a modern data platform is no longer just about moving data. Organizations must scale reliably, control costs, enforce governance, and accelerate analytics. This book shows you how to design and operate production-grade data platforms using AWS Glue and related AWS analytics services.
You will begin with core data management concepts before moving into ingestion from diverse sources, data preparation strategies, metadata management, security controls, and cross-account data sharing. Learn how to design efficient data layouts, orchestrate pipelines, implement CI CD practices, and manage the full lifecycle of data integration workloads.
This updated edition expands coverage of open table formats such as Apache Hudi, Delta Lake, and Apache Iceberg, along with performance tuning, observability, cost optimization, and real-world troubleshooting. You will also explore integrations with machine learning and generative AI workflows powered by Glue and SageMaker.
Written by AWS engineers and architects with deep hands-on experience in large-scale enterprise data lakes, this guide blends architecture principles with real-world implementation insight.
By the end of this book, you will be able to design, deploy, monitor, and optimize scalable serverless ETL pipelines and governed data platforms on AWS.
You will begin with core data management concepts before moving into ingestion from diverse sources, data preparation strategies, metadata management, security controls, and cross-account data sharing. Learn how to design efficient data layouts, orchestrate pipelines, implement CI CD practices, and manage the full lifecycle of data integration workloads.
This updated edition expands coverage of open table formats such as Apache Hudi, Delta Lake, and Apache Iceberg, along with performance tuning, observability, cost optimization, and real-world troubleshooting. You will also explore integrations with machine learning and generative AI workflows powered by Glue and SageMaker.
Written by AWS engineers and architects with deep hands-on experience in large-scale enterprise data lakes, this guide blends architecture principles with real-world implementation insight.
By the end of this book, you will be able to design, deploy, monitor, and optimize scalable serverless ETL pipelines and governed data platforms on AWS.
- 1. Data Management: Introduction and Concepts
- 2. Introduction to Important AWS Glue Features
- 3. Data Ingestion
- 4. Data Preparation
- 5. Data Layouts
- 6. Data Management
- 7. Metadata Management
- 8. Data Security
- 9. Data Sharing
- 10. Data Pipeline Management
- 11. Monitoring
- 12. Tuning, Debugging and Troubleshooting
- 13. Data Analysis
- 14. Machine Learning and Generative AI Integration
- 15. Architecting Data Lakes for Real World Scenarious and Edge Cases
- 16. End-to-end development lifecycle
- 17. Open Table Format
- 18. Cost Optimization
- Title:Serverless ETL and Analytics with AWS Glue. Design scalable data lakes, optimize ETL pipelines, and accelerate analytics on AWS - Second Edition
- Author:Noritaka Sekiyama, Albert Quiroga, Tomohiro Tanaka, Subramanya Vajiraya, Akira Ajisaka, Ishan Gaur
- Original title:Serverless ETL and Analytics with AWS Glue. Design scalable data lakes, optimize ETL pipelines, and accelerate analytics on AWS - Second Edition
- ISBN:9781835468012, 9781835468012
- Date of issue:2026-09-11
- Format:Ebook - EPUB
- Item ID: e_4ysl
- Publisher: Packt Publishing
Loading...
Loading...