Yusuf Ganiyu
Senior Data Engineer

AstraZeneca Ltd

UK

About

Yusuf Ganiyu is a Senior Data Engineer at AstraZeneca, where he architects AI-powered big data solutions that transform pharmaceutical data into actionable insights at enterprise scale. As Founder of Data Mastery Lab—recognized as London's Best Data Engineering and AI Training Platform in 2025—he has established himself as a leading voice in big data education.With over 50,000 students taught globally through platforms like Udemy, YouTube (CodeWithYu, 1M+ views), and his own training platform, Yusuf excels at making complex big data concepts practical and implementable. His end-to-end projects, ranging from real-time streaming pipelines to complete data platform implementations, serve as reference architectures for engineering teams worldwide.Holding an MSc in Computational Intelligence from Cranfield University, Yusuf is triple-certified across AWS, Azure, and GCP. His expertise spans the complete big data stack, including Apache Kafka, Spark, Airflow, Cassandra, Elasticsearch, and modern cloud data services.As an active contributor to the global big data community, Yusuf was a 2023 Elastic Silver Contributor with a 2M+ reach on Stack Overflow. His unique position, bridging enterprise implementation and large-scale education, provides practical insights into what truly works at production scale.
Workshop

Yusuf Ganiyu | End-to-End Real-Time Data Pipeline: From Ingestion to Insights (Hands-On Workshop)

Apache Kafka, Real-Time Streaming, Data Engineering, Hands-On
This intensive hands-on workshop guides participants through building a complete production-grade streaming pipeline from scratch, using patterns proven at AstraZeneca and refined through teaching 50,000+ engineers at Data Mastery Lab.THE CHALLENGE: Organizations struggle to move from batch to real-time processing. Most streaming tutorials demonstrate "hello world" examples that collapse under production load. This workshop addresses the gap with enterprise-tested patterns.WHAT PARTICIPANTS BUILD: Using Apache Kafka, Spark Structured Streaming, and Docker, participants construct an end-to-end solution including: - Multi-source data ingestion (APIs, databases, files) - Stream processing with exactly-once semantics - State management and windowed aggregations - Writing to multiple sinks (Cassandra, Elasticsearch, data lakes) - Production monitoring and alertingMEASURABLE OUTCOMES: Participants leave with: - Working code repository (ready for production adaptation) - Reference architecture diagrams - Checklist for streaming project evaluation - Before/after performance benchmarks from real implementationsSTEP-BY-STEP STRUCTURE: Morning: Architecture foundations, Kafka setup, producer/consumer patterns Afternoon: Spark Streaming transformations, state management, deployment with CI/CDThis curriculum has achieved 4.8+ ratings across 50,000 students, with documented success stories of engineers deploying streaming systems within weeks of completing the training.

2026-11-24

09:00

17:00