Marcin Szymaniuk
CEO | Senior Data Engineer | International Conference Speaker

TantusData

Poland

About

Specialising in helping clients monetise big data since the early 2000s, Marcin Szymaniuk leads a team of seasoned data engineers with expertise in data engineering, machine learning (ML), machine learning operations (MLOps), and cloud technologies.Marcin is adept at solving both non-standard challenges and everyday problems that require fast, practical solutions. His experience spans a wide range of industries and project sizes, with a strong focus on artificial intelligence (AI), ML, and deployment strategies.He has presented at numerous industry events, including Infoshare, J On The Beach, Devoxx, Huawei Eco-Connect Poland 2023, Berlin Buzzwords, Codestar, GeeCON, and Java Day Istanbul.
Talk

Marcin Szymaniuk | From No Tests to Trust: Evaluating End-to-End GenAI Systems

AI Evaluation, AI Testing, RAG, LLM-as-a-Judge, Synthetic Testing, AI Reliability
<p>The talk begins with a company where generative AI (GenAI) is already in production, but nothing is tested. Every change is risky, and nobody can tell whether the system has improved or simply broken something.</p> <p>Marcin Szymaniuk explains how to bring structure from three perspectives:</p> <ul> <li>Developer: how to build evaluation loops, datasets, and basic tests for retrieval-augmented generation (RAG), retrieval, and document pipelines</li> <li>Team lead/manager: what is realistic to test, what is expensive, and how to plan work around uncertainty</li> <li>Business: how to move fast without guessing and how real user feedback is more valuable than any single metric</li> </ul> <p>He then covers how to evaluate non-deterministic systems:</p> <ul> <li>RAG and document-quality metrics, including relevance, faithfulness, and completeness</li> <li>Conversation testing using synthetic scenarios</li> <li>Using large language models (LLMs) as judges when rule-based checks fail</li> </ul> <p>Marcin also discusses the trade-offs: not everything should be automated, and not everything needs perfect scoring. Testing GenAI is not about achieving perfect correctness - it is about building enough confidence to ship, iterate, and avoid breaking what already works.</p>

2026-11-26

10:10

10:55

Data Meets AI