Schedule
1:45 pm
Spark, Kafka and Endless Java Stack Traces: A Streaming Ingestion Case Study
At UMC Utrecht, the current on-premise dataplatform supports a large variety of use-cases: research, BI, in-house applications, training AI models, you name it. However, the large variety of demands on a rather simple on-premise system has been causing issues for some time, and for about a year we have been working on a cloud platform that is able to do all of it better. Building a cloud data platform that supports near-real-time data ingestion seemed straightforward, until we actually had to make all the pieces work together…
In this case study, we relay our experience of building a near-real-time ingestion platform on Azure and Databricks. We start at a SQL Server database (connected to the main system in the hospital) that uses Change Data Capture (CDC) to capture changes and then stream them through Debezium and Azure Event Hubs into Databricks. We’ll dive into the trials and tribulations of designing, implementing, and operating this pipeline: the architectural choices we made, the problems we encountered, and the lessons learned along the way. Expect practical insights into CDC, streaming ingestion, reliability, and, inevitably, a heap of Java stack traces.
Guests

Axel Blom
Data Engineer
UMC Utrecht