Stop Optimizing the Wrong Part of Your Spark Job

High utilization doesn’t necessarily mean an efficient Spark workload – and the layer where waste shows up is often not where you need to fix it. In this webinar, we’ll walk through a practical framework for investigating and optimizing Spark workloads across infrastructure, runtime, data access, and query plans, using real production cases to show the signals, root causes, and optimizations at each layer. We’ll also discuss what it takes to collect the runtime signals needed to apply this framework effectively, and the trade-offs between different approaches to monitoring Spark workloads.


