Ammar Chalifah
Senior Data Engineer

Modash

Estonia

About

Ammar Chalifah is a data engineer focused on building scalable, optimized data pipelines. He has reduced compute costs and wall-clock time through compute and storage tuning, generating more than €1 million in savings for the organizations he has worked with. He currently builds an influencer-marketing platform at Modash.
Workshop

Ammar Chalifah | Spark Pipeline Optimization

Apache Spark, Data Pipelines, Performance Optimization, Query Optimization
<p>1. Abstract<br /> Despite its limitations, Apache Spark is still the go-to choice for big data workloads across organizations in the industry. However, organizations around the world waste money and productive time by running inefficient Spark jobs. The difference between an efficient Spark pipeline and an inefficient one could be an order of magnitude greater in terms of both compute cost and wall-clock time, and investing in an efficient pipeline could yield more than 75% savings in money and time.</p> <p>In this workshop, Ammar Chalifah will cover best practices for optimizing a Spark job, from reading the physical plan, minimizing shuffle and skew, avoiding UDFs, choosing the right storage format and storage layout, and right-sizing the cluster.</p> <p>2. Agenda</p> <ul> <li>Brief introduction to the topic: problems around Spark pipelines (10 minutes)</li> <li>Setting up repository for attendees (10 minutes)</li> <li>Reading Spark UI (10 minutes)</li> <li>Optimization case 1: shuffle. Demonstration + practice (15 minutes)</li> <li>Optimization case 2: shuffle, StoragePartitionedJoin (10 minutes)</li> <li>Optimization case 3: lazy execution, solving it through cache/checkpoint/materialization (15 minutes)</li> <li>Optimization case 4: UDF vs native Spark (10 minutes)</li> <li>Optimization case 5: native acceleration, Apache Gluten (15 minutes)</li> <li>Closing, questions (5 minutes)</li> </ul> <p>3. Objectives<br /> Attendees understand the biggest bottlenecks in Spark pipelines, know how to identify them, are able to implement an optimization technique, and are aware of production best practices.</p> <p>4. Target audience and Prerequisites</p> <ul> <li>Data Engineer in early-to-mid career</li> <li>Engineers looking to deepen their expertise in Spark</li> </ul>

2026-11-24

09:00

17:00