
About
Kate Gawron | Building a GenAI-Ready Lakehouse on AWS: From Relational Data to RAG-Optimised Datasets
Most organisations want to build GenAI applications, but their data platform wasn’t designed for it. Relational databases and data warehouses are excellent for structured analytics, but GenAI introduces new requirements: handling unstructured content, supporting rapid iteration, enabling governed access to sensitive datasets, and producing “retrieval-ready” data that can power search and RAG workflows.
In this 6-hour hands-on workshop, participants will build a GenAI-ready lakehouse on AWS. We’ll start with a traditional relational dataset and a set of unstructured documents, then design a lakehouse architecture using Amazon S3, AWS Glue Data Catalog, Athena, and Apache Iceberg. Participants will implement ingestion and transformation patterns that create both analytics-friendly tables and GenAI-friendly datasets, including chunked text outputs, metadata enrichment, and quality checks that improve retrieval performance.
The workshop is structured as a real end-to-end case study: we benchmark the starting point, build the lakehouse step-by-step, and demonstrate measurable outcomes such as faster dataset iteration, improved searchability, and better governance. Attendees will leave with reference architectures, a GenAI data readiness checklist, and templates they can apply immediately.