Featured

Building Real-Time Data Governance at Scale with Apache Kafka ft. Tushar Thole



Published
https://cnfl.io/podcast-episode-205 | Data availability, usability, integrity, and security are words that we sometimes hear a lot. But what do they actually look like when put into practice? That’s where data governance comes in. This becomes especially tricky when working with real-time data architectures.

Tushar Thole (Senior Manager, Engineering, Trust & Security, Confluent) focuses on delivering features for software-defined storage, software-defined networking (SD-WAN), security, and cloud-native domains. In this episode, he shares the importance of real-time data governance and the tool—Stream Governance, which his team has been building to debug and monitor Apache Kafka® data streams on Confluent Cloud for real-time insights and observability.

With the increase of data volume, variety, and velocity, data governance is mandatory for trustworthy, usable, accurate, and accessible data across organizations, especially with distributed data in motion.

When it comes to choosing a tool to govern real-time distributed data, there is often a paradox of choice. Some tools are built for handling data at rest, while open source alternatives lack features and are not managed services that can be integrated with the Kafka ecosystem natively.

To solve governance use cases by delivering high-quality data assets, Tushar and his team have been taking Confluent Schema Registry, considered the de facto metadata management standard for the ecosystem, to the next level. This approach to governance allows organizations to scale Kafka operations for real-time observability with security and quality while remaining compliant with data regulations, such as GDPR.

The fully managed, cloud-native Stream Governance framework is based on three key workflows:
Stream catalog: Search and discover data in a self-service fashion
Stream lineage: Understand the complex data relationships with interactive, end-to-end maps of event streams
Stream quality: Deliver trusted, high-quality event streams to the organization

Tushar also shares use cases around data governance and sheds light on the Stream Governance roadmap.

EPISODE LINKS
► Stream Governance – How it Works: https://cnfl.io/stream-governance-episode-205
► Data Mess to Data Mesh | Jay Kreps: https://www.youtube.com/watch?v=G_qRkDZfX00
► Demo: Stream Governance: https://www.youtube.com/watch?v=2KNP1P9Wk-E
► Data Governance for Real Time Data: https://cnfl.io/data-governance-for-real-time-data-episode-205
► Kris Jenkins: https://twitter.com/krisajenkins
► Join the Confluent Community: https://cnfl.io/join-confluent-developer-community-episode-205
► Learn more with Kafka tutorials, resources, and guides: https://cnfl.io/confluent-developer-episode-205
► Live demo: Intro to Event-Driven Microservices with Confluent: https://cnfl.io/event-driven-microservices-demo-episode-205
► Use PODCAST100 to get $100 of free Confluent Cloud usage: https://cnfl.io/try-confluent-cloud-episode-205
► Promo code details: https://cnfl.io/podcast100-details-episode-205

TIMESTAMPS
0:00 - Intro
2:02 - What is Stream Governance?
7:58 - Friendly UI
9:13 - Ensure Data Quality
11:44 - Stream Quality
16:42 - Stream Lineage
21:12 - Data Regulation Compliance: GDPR & CCPA
23:02 - Stream Catalog
29:04 - All your data in one place
33:43 - Roadmap
39:18 - It's a wrap!

CONNECT
Subscribe: https://youtube.com/c/confluent?sub_confirmation=1
Site: https://confluent.io
GitHub: https://github.com/confluentinc
Facebook: https://facebook.com/confluentinc
Twitter: https://twitter.com/confluentinc
LinkedIn: https://www.linkedin.com/company/confluent
Instagram: https://www.instagram.com/confluent_inc

ABOUT CONFLUENT
Confluent is pioneering a fundamentally new category of data infrastructure focused on data in motion. Confluent’s cloud-native offering is the foundational platform for data in motion – designed to be the intelligent connective tissue enabling real-time data, from multiple sources, to constantly stream across the organization. With Confluent, organizations can meet the new business imperative of delivering rich, digital front-end customer experiences and transitioning to sophisticated, real-time, software-driven backend operations. To learn more, please visit www.confluent.io.

#datagovernance #apachekafka #kafka #confluent
Category
Management
Be the first to comment