
43
How to leverage real-time data processing for enterprises

Andrew Sellers
Head of Technology Strategy at Confluent
In this episode, David is joined by Andrew Sellers, Head of Technology Strategy at Confluent, to discuss the dynamic world of data streaming and its transformative impact in the enterprise as well as everyday applications. Andrew sheds light on the evolution of data infrastructure and how Confluent is constantly working to revolutionize data management with its platform.
Join us as we discuss:
The pivotal role of technologies like Kafka and Apache Flink in application architectures
How data streaming translates into tangible business value, from improved decision-making to enhanced operational efficiency
Foresight into the evolving trends in real-time data processing, including why data streaming will ultimately make generative AI better
David Joy:
What is up, everyone? And welcome to the Big Ideas in App Architecture podcast. In today's episode, I speak with Andrew Sellers, the head of strategy at Confluent. Andrew and I get into his career, to talking about some amazing things he and his team are working on at Confluent. We get into some really interesting ideas around Kafka, Flink, and AI, and how it'll all come together in the future. So pump up that volume and get ready for a fun conversation with Andrew Sellers.
Welcome to the podcast, Andrew. How are you doing today?
Andrew Sellers:
David, I'm doing great. I'm really excited to be here, so thank you so much for having me.
David Joy:
Before we get into any of the details on some of the interesting stuff you have been working on, let me know or let the people know really, what you're doing at Confluent and a little bit about yourself.
Andrew Sellers:
Sure. So yeah, my name's Andrew Sellers, I work at Confluent. Confluent manages the world's premier streaming service. We have a data streaming platform that allows one to stream, process, govern, and connect data from wherever it is in the enterprise. And my role there is I'm the head of technology strategy. We have three mandates on this team with the technology strategy group. The first thing is our product team has the roadmap very well contemplated for the next three to four quarters. My team considers what we do after that, what should we be doing in three years, in five? That first mandate is heavily influenced by the second, which is we look at the competitive landscape.
So we look at what everyone else is doing, we look at their marketing, we look at their docs. If it's open source and the license allows it, we benchmark. We figure out what out there is really good architecture verse what's real technology that we have to have a response for? What's the market really clamoring for and what is helping add, create business value wherever people are using data infrastructure? We're not in the field, we're not on the product team. So we get to, in a sense, to have a little more objectivity in a way. The third mandate is then we do a lot of thought leadership, we do a lot of blogging, public speaking, things like that. So that's maybe a little bit of what you were picking up on. So yeah, anyway, it's a great role.
David Joy:
I've been following Confluent for a while. But before I talk about Confluent, I do have to say this was such a humble introduction, because you were also the CTO at this great company called Complex for about almost eight years. And before that, you were a senior cyberspace operations officer, so like a CTO role and worked in, you were also an academic for a bit. So you've been in this space for a while and it's been interesting to know all of that actually.
Andrew Sellers:
It's been interesting to get to see a lot of things. And I think that does make me better at this role, the breadth of things, just because there's a perspective you get when you come from other communities. Cybersecurity is very closely related to streaming in a way. I like to say that my streaming began at the last role, not this one. Cybersecurity is very much a real-time problem. If you think about what it is, there's a lot of heterogeneous telemetry, depending on the enterprise, generated from all over the globe that you have to bring together high velocity and reason about. If you put it at rest, you can't really process it affordably. And so there's tremendous value in these event-driven architectures that are able to scale very effectively by decoupling systems, teams, and technologies. Not everything is a low-latency problem, but a lot of things fall into that category, which is why data streaming can be tremendously beneficial for a lot more wider use cases than many people think. It's far more than this niche capability.
David Joy:
100%. No, I agree with that. I think today one of the benchmarks is latency experience. If you get poor latency or if your systems are down, you are going to see somebody unhappy, tweeting about it, and saying how bad of a system it is. 15 years ago, nobody really cared about it that much to go and tell the people that, I really want to tell people. But things have changed so much.
Andrew Sellers:
And David, if I can, just to comment on that. I think you're exactly right. I think 15 years ago, as you said, I think the difference was technology was a supporting function. That's how we saw it. You had Blockbuster Video, where the business was distributing these movies and discs and people came in. And it was still interesting to see which movies were the most viewed and they'd help predict which ones I'd buy in the future, but there wasn't a real-time component to that. And the difference now, to be competitive technology really has become the business. Blockbuster was eaten by Netflix. And there, real-time really matters in terms of delivery, the recommendation engines, having content delivery that is robust and reliable and isn't keeping people waiting because people won't put up with it. And so you're exactly right that the entire landscape has changed. And so while these things, events were always real-time in a sense, business has become real-time in a way that was just never true before.
David Joy:
True. True. And this is funny that you brought up Netflix, because I don't think a lot of people understand or give credit to how much of a pioneer they have been and broadening infrastructure, architecture, and the kind of systems we have all adopted to. For example, Netflix did a bunch of experimentation that people got inspired by. So they do get credit for whatever they did with the movie market, the streaming market, for television and stuff like that. But on the infrastructure side of things also, they did so much. So I'm glad you know that story and you love talking about it.
Andrew Sellers:
It's one of my favorite tech blogs. One thing that I'm very grateful to that community for is they share a lot of the knowledge, and they open source a lot of the infrastructure work that they do, and I have learned a tremendous amount from them.
David Joy:
Let's jump into, from that segue, into Confluent itself, a little bit on Confluent. So I personally am a fan of the company. We also partner with Confluent
Andrew Sellers:
And we appreciate your partnership.
David Joy:
Yeah. Yeah, I know. It's great. It's such a fantastic product. But before I talk about my origins of learning about Kafka and what Confluent did, tell me how you got introduced to Kafka and where did your story begin?
Andrew Sellers:
We talked a little bit about how I came from the cybersecurity community. I'm a bit of a transplant into streaming. And there, it just turns out that cybersecurity, I think, if there is a canonical streaming problem, it might be cybersecurity, just because real-time really matters. These events get generated and they all look a little bit different. And so being able to handle them in a way that's very flexible is really important. And new indicators of compromise are always coming out, and so time to market is really important. And David, you and I have been doing this a long time. One of the, I would say the absolute trues of this business is the bigger data gets, the more specialized one has to be, and how data is organized and queried. Well, with streaming, the specialization is the data is meant to be consumed, and that's really handy with something like cyber because as a new indicator of compromise comes out, time to market is just a matter of instantiating a new consumer group or writing a new query against the stream. And that is an incredibly powerful formalism for staying ahead of the hackers.
And so my previous role, as you said, was I worked at a really cool cybersecurity identity analytics company. And really what we did then is we assured our customers that the people in their environment are who they say they are. In order to do that, you have to validate authentication transactions. Well, a big global enterprise will generate tens of thousands of authentications every second, and that can be really challenging if you put that data at rest. Like a lot of startups, when we're very small, you do the thing that you can, which is the fastest time to market, get the thing in people's hands and really test how it works. And then we ended up having that problem that you want to have when you're a startup, in that you achieve product market fit and the demand out-scales your current tech stack.
And so I ended up reading a book by a gentleman named Ben Stopford, on designing event-driven applications and completely re-rationalized our business. So instead of putting everything at rest, we would look at everything once and do stateful stream processing. So the vast majority of our critical microservices were actually Flink jobs running on top of persistent Kafka. And it was an incredible pattern because we were able to just do these magical things,. All of a sudden with just a fraction of the computing footprint we were using before, we were able to power all of our customers globally and we could scale arbitrarily. And it was just an incredible thing to be a part of, just because just by changing the way we thought about things, just by realizing that most of the time it's simply enough to look at every piece of data once and then capture some characterization of the state, we were able to just do some really remarkable things with it. And so ever since, I was absolutely in love with streaming. And so when this position opened up, I couldn't resist.
David Joy:
Oh, yeah. It's amazing. What a great story. My introduction is slightly different. I had this thing, when around 2017, I had an Alexa at my home and I was like, "How does Alexa really work? Is this really listening to me?" That was one of the things I was thinking about. And I was like, "Why don't I build my own Alexa or my Siri?" That was what I was thinking. So I started basically building my own Alexa. And I actually ended up building a small project called Apollo, that I eventually shared with a bunch of people and say, "If you want to do something, you want to learn how this works, you do it." But one of the problems that I had was I wanted to, when you start having conversation with Alexa or Apollo, I wanted to generally figure out, I want to see how many interactions are happening. I wanted to see views.
And I didn't want to read the database to determine that. I wanted everything real-time. And I naturally ended up in a situation, in an architectural pattern that required some real time streaming. And I had no idea how to do it. So I started hitting up forums and Stack Overflow at the time, and keyword searches. And then I eventually saw Kafka pop up everywhere, and then that's how I started using Kafka. And then I was like, I had to re-architect everything just for this small project. And my brain just changed that I can have a publisher consumer model and I can decouple my application without being dependent on the database. That was amazing for me. So I felt like what Kafka did really was of course, and what you guys have been doing at the company for almost a decade as leadership in the streaming space, but also showing the world how to scale high throughput, low latency streaming experience and have this decoupled experience throughout so many use cases. So that's what I give Kafka credit for, to help me understand this [inaudible 00:10:40].
Andrew Sellers:
David, I love hearing that story, so thank you for sharing that with me. That's infectious. That's why we do what we do. And I think you're exactly right, that the magic of it is all of a sudden you're not coupled to this database query anymore. As much as I always love wide column stores, you almost have to create a column family for every query you want to ask. And so there's just very little ability to ask a different question later. And we've all been in those application meetings where some executive says something like, "Oh, it would just be so cool if I could just see this." And it's like, well, the data model doesn't support that, and that's just not how streaming works. And it's just the flexibility, what it allows people to create and reimagine and to do it again and again and again as they go, it's just such an important technology I think for enabling so many of the experiences we've come to take for granted today. Every time we watch a Netflix again or take an Uber, it's all powered by that.
David Joy:
One of the thoughts I've had is that even though the core idea of event-driven architecture is centered around real-time experience, things have changed in the last decade. So how do you see the changes that have happened at Confluent? What's happening in Confluent land that you are excited about?
Andrew Sellers:
The old Silicon Valley model, I think the way that venture capital and things worked is it was all about building this very niche capability. We're going to be the best time series processing engine ever, or we're going to be the best wide column store for financial trading data or whatever. And those are really great things and they find markets and then you can really specialize and build a great company around that in a sense. But the problem with that is then all of the complexity becomes the burden of that end consumer, that end enterprise that is simply trying to run their business and not wrestle with all this data infrastructure tooling and not have to constantly think about how all these things integrate together. In some ways that's where most of these big projects fail, it generates a lot of business risk that people don't want to consume.
And so one thing that we at Confluent have been focused on is on building and enabling what we call a data streaming platform. And that is what we consider to be the core capabilities of process, of the actual stream itself, of governance, and of connecting data wherever it needs to be throughout the enterprise, in order to create data that can really be reusable, discoverable, and trustworthy. And so that way your application developers and your analysts and your data engineers are focused on building really differentiated business applications, and they're not having to wrestle with data infrastructure. And I think that as time goes on, we're going to have even more work to do there to make it even better. I think there's, in addition to the consolidation of platforms, we've worked really hard to provide effective abstractions as well. And so if you're just using open source Kafka and you're an operator of that, you've got to spend a lot of time thinking about how brokers get tuned and about partition counts and things.
And some of that is hard to change later on. But I think as we get better at some of the work we're doing, we announced our core capability, which is our cloud-native Kafka-compatible stream processing broker, that there's more and more we can do around dynamic allocation and really leveraging the sophisticated primitives of the cloud so that devs can more and more just instantiate a topic and not think about brokers and partitions and considerations like that. That if it's, I just need a topic or a queue and it works and performs as I need it to, and then everything becomes more seamless and intuitive. And I think that we've made a lot of progress there. We've got a lot more work to go, but I think that that's something five to 10 years, I think we're just ruthlessly focused on how we make this a better experience for our customers, and how more and more we connect them to their data. And they're thinking about what their data actually is, and far less about the transport of it and how it actually moves and how it actually persists.
David Joy:
Now, one of the thoughts I've had is the behavior or data pattern itself has changed. Obviously, we have the challenge of growth of data, and that challenge has been faced by everybody in the infrastructure space. You go from compute to storage to databases, everybody's like, "How can I scale my database?" And even Confluent, "How can I stream this amount of data and move it around?" But your story is also intertwined with the story of Apache Flink, and you've written extensively about some of your thoughts about it. So one of the things that'll be great for you to help people understand is how do you compare the roles of Apache Flink in Kafka, and what is the synergy there for the real-time data processing landscape? And do you feel like there are some specific scenarios in which they work really well?
Andrew Sellers:
They work really well together and they also work very well independently. The thing that, to clarify, for those that may not be as familiar with the technologies, they're really very complimentary things. Kafka is an incredible data broker for streams. It improved upon what I think was a very good conceptual model of the message queuing. But the conceptual model of Kafka is built around the immutable log, and so the idea then that data becomes replayable. With traditional, the pub-sub pattern, I suppose, data was basically discarded post-acknowledgement, so queues were meant to be ephemeral. And what Kafka really changed about it was this idea that no, it turns out you actually might have multiple subscribers for this data, and you might want to actually go back and revisit history later on. And so in addition to being a really wonderful real-time capability for what's happening right now, it can also provide a really great way to consume what happened in the past.
And in many ways, that Kafka as a conceptual model, it's protocols, they've clearly won. If you look at every major streaming technology, streaming broker, they've all adopted Kafka compatibility to some extent. The things that live in CSPs, event hubs, and PubSub+. And even aspects of RabbitMQ has replayability now in some deployment options. And so the conceptual model of Kafka has really won, and that's I think really, really cool. And I think that as we have more and more out there that find the incredible capabilities of this, I think it's only going to solidify Kafka's preeminent place in the modern data processing stack. Flink is also another wonderful technology, but it doesn't persist data, it processes it. And not just streaming data, but can also process batch data as well. And it can do it in a way... One thing that deeply impressed me with Flink is that there was other things that preceded it, but they tended to be these leaky abstractions where you actually could do streaming in batch, but you had to know what was back there or else it wouldn't work effectively.
I think Flink has done a very elegant job of making a lot of that work a lot more seamlessly. But as a distributed data processing engine, it can then consume from a streaming engine like Kafka or another one, or it can basically process and run analytics or do materialized views or do these streaming pipelines from something like Iceberg, that's an actual table view of persistent data. Which is really wonderful there because then Kafka can serve as your processing engine wherever it is. But similarly, Kafka can be your streaming broker and that you can bring different technologies to bear to process it as well. But they happen to work very well together. The communities are there. At Confluent as part of our data streaming platform, we've done a lot of work to integrate the common identity plane and to make it so that aspects like governance and the schemers and all the metadata that can help power a really great experience are all provided in one place. And the technologies really do work very well together.
David Joy:
I'm an avid reader of DoorDash blogs. DoorDash is also one of Cockroach Labs', they use us. And obviously, I love some of the blogs they have released. And I saw how they're using Flink and Kafka in their architecture. And it's fantastic, the brilliance in which it all works together. And it's really, really cool. One of the other exciting things is that the transformative part of this data is very critical. Nobody thinks about this or had thought about it for a while. I know there are other projects who were introducing these ideas, like Spark was a project. And Spark introduced Spark Streaming and then they had Spark SQL and a bunch of other things, but didn't mature exactly how Kafka and Apache Flink worked together in that sense. So I've seen some really interesting use cases where we tried using that, and then switched it off and moved Kafka and said, let's just use... We even used, I think at one point, Kafka SQL, KSQL to do certain things, but I eventually moved to Flink kid of an architecture with Kafka. So that was fun.
Andrew Sellers:
I was just going to say, I think that you're exactly right that there are a lot of good data modeling patterns that the Flink Kafka stack really enables, that have work across industries. And I think they're just incredibly powerful for the kinds of work that you need to do. It's really easy ways to enrich data, do these factor dimension kinds of joins. Events, because they tend to be consumed one at a time, and you want to do that even if it's in a stateful way, you still want to try and address them one at a time to whatever extent you can. You tend to de-normalize them.
Whereas the originating operational data stores tend to be highly normalized, just because 30 years ago we actually cared a lot about storage and we didn't want to expose application developers to deletion and update anomalies. Well, what's really nice is that Flink provides really standard great ways of taking data out of Kafka and joining streams, and providing these events then that can help consumers be ready to go as they are. And I think that's really its power is that now all of a sudden, this application developer doesn't have to think about all this data mess, they just have what they need. And that's an incredibly powerful thing that the data stream platforms provide.
David Joy:
I agree. I agree. And I think there's been so much that has been happening in this space, especially with what you say, data pipeline and what you can do with data, enriching the data, processing it, feature engineering, a bunch of other things. But another pattern that has been blowing up and basically is sitting right in front of you and me is generative AI. And I know you've dabbled with it and you've spoken about it in the vlog. I dabbled with generative AI and AI capabilities a lot myself. But the transformative potential for enterprises with having so much data, having so much data real-time, what is your thought on where all of this is going? And let's start there. How do you see all of this coming together right now?
Andrew Sellers:
So it's interesting. I think that there's a wonderful virtuous cycle, in that there's a bilateral relationship between generative AI and streaming. So streaming can make generative AI-enabled applications better. I think that, and again, as you said, I was an academic before, I actually studied artificial intelligence in graduate school, and in some form for the last 20 years have been involved in either researching AI or bringing AI-enabled applications to market. And in all that time, it always felt like it was three to four years away, that we were right on the cusp of a big transformation, but the inflection point wasn't quite here yet. And a lot of that I think came because the technologies themselves, as powerful as they were, they weren't particularly accessible. So Google and Meta and some other big winners had them, but they weren't really democratized or accessible.
And so this moment feels different, that the inflection moment, inflection point feels a lot closer now. And a lot of that comes because traditional, we'll say predictive ML modeling, particularly with some of the deep learning architectures, you needed the legion of PhDs in statistics and data science to help build and curate. As you said, there's a lot of activities in an MLOps pipeline feature engineering and model training and model deployment. And those are tremendously hard things and they're in some ways more art than science. What generative AI does a little bit different is we have these foundation models now. So someone else spends a lot of time and money training these things, and then we all get to benefit. But the problem is the data engineering challenges don't go away. If anything, they become more real-time.
And so one of the best patterns we know of to make a generative AI-enabled application useful is to enrich it at real -time with domain-specific information. This is often the retrieval augmented generation or prompt engineering kinds of patterns. What they boil down to is providing additional context to the LLM so it can give you a better answer and prevent some of these solutio-nations and things. And in order to deliver the reactive sophisticated experiences that consumers have come to expect, whether they're your internal customers or your internal developers that you're augmenting your knowledge workers or they're an external customer, people aren't willing to put up with waiting for an hour for the answer to come back.
Well, in the modern enterprise, as you know, the relevant domain-specific data, it could be spread out and siloed in some disparate operational database and not really there where it's ready. And so one of the real superpowers of data streaming is that in real time we connect data across the enterprise, and then we can then stage it into a vector store wherever it needs to be so it's accessible to that generative AI-enabled application. And there's a lot of patterns I think when we build those applications, that work really well as data streaming patterns. Because again, we get to decouple the teams and the technologies and systems that come into play when we build those applications. And they're very multidisciplinary, generative AI-enabled application. You need some full-stack people to do the web part to handle some sessions, but you don't want them necessarily thinking about the seconds-long interval between responding BLM and back, just because that's an attorney-distributed system.
So you want to make that async, which is what a data streaming application does. And by decoupling that from the backend part, well now you can have ML engineers that can scale confidently and independently of all the other teams. And from a deployment standpoint, that can scale independently as well, if it's not a big monolith and the only thing that connects them is a well-governed topic inside something like Kafka. You have compliance teams and you do the same thing. And the GenAI part is going to evolve a lot. But if you decouple the LLM because it's used connected just through a Confluent topic, well, now it's a lot easier to swap out the LLM when Databricks comes out and all of a sudden they announce the better LLM, or the embedding model changes and all of a sudden you need something that favors higher dimensionality than what you had before. Or just all the other things that can happen. Maybe your vector store, you need a different vector store.
And because that evolves so rapidly, you want to treat those things like modular components. Then at the same time, I think GenAI can really help data streaming as well. One thing that's really important for data streaming is we talked a little bit about it, good governance of data. The thing about governance with data is it's actually not rocket science. It's not that hard. It's just a matter of building the right metadata. It's basically having good schemas, of having good field descriptions, of having good business metadata about what this thing actually is and who owns it and who can change it. And that today ends up being a lot of manual road work. But if a data engineer can have GenAI that can induce from maybe 50 examples, here's what the machine thinks the schema is, here's the field descriptions that the machine guesses that are there, and here's some suggested business metadata. Now a data engineer is just validating. They're clicking, yes, or maybe making a few small edits. Now this all becomes very possible.
In my life as an academic, one of the painful things we lived through was the semantic web 15 years ago. And one of the reasons it failed, as you know, is because it was incumbent upon data creators to create all the metadata that powered the thing. And nobody really did it. It's a time to market thing and it's like they didn't see the business value. I think the same thing happens in this fast-paced world we live in now, where we generate all this data and if governance is hard, then it gets skipped. But if we make governance easy, which GenAI can help us do, then data streaming becomes even more powerful. So yeah, basically that was a long answer, but they help each other.
David Joy:
No, no, it's true. Because there's no one specific answer. Generative AI is such a broad... It's not broad really. We know very clearly what we are capable of, what technology we have, but the applications of it is in so many different places. I know you mentioned Databricks releasing DBRX, which I was actually in the code yesterday checking it out. It's really cool what they've done and a lot of homages to some other models that have come before them. But it's pretty interesting the capabilities. And I feel like one of the examples you gave around data governance, where because this is a challenge for me as a developer. As a developer, I'm trying to develop something and all I know is this is where all my data is coming from. If I can have a recommendation coming from GenAI or from my Copilot saying, "Hey, look, this is what the data looks like, this is the kind of pipeline you can build through it or you can do these feature engineering," I save so much time considerably.
So there is so much value and I can see where you guys are trying to go with this. So where do you think overall, with all of these capabilities coming in, obviously the nature of data streaming is also going to change in the next five to 10 years. There is this, I'll give you an example. Some of the things I'm thinking is one of my perspectives is the app store led to the creation of applications and modern data feeds and architectures changed and what we were doing. And now we are at a space or a time where we are moving away from, "Hey, I don't want apps actually. I don't want to interact with apps at all. I just wanted to talk to them."
So we are talking about there's this idea of large action model, where you just communicate with natural language, the app really exists, but you don't interact with it. And that adds an extra layer of streaming, because you have to now condense or read that natural language toward that information and that behavior somewhere else. So data streaming behavior is also going to change. So I'm just curious as to how you are looking at all of this in the five to 10 year timeline. Where do you think it's going? And maybe 10 years after, I'll come back in and say, "This is what you said, Andrew." So this is the reason for the question.
Andrew Sellers:
Yeah, it's good and I like this. So I think you need to be bold with these kinds of things, so I need to say something that's actually refutable. I think sometimes people offer platitudes. And so yeah, let me take a crack at something and I sure might be wrong. So I think when we first started, there was a lot of hype around we're going to build the LLM to rule them all, that there's going to be one that's just all things to all people. And I think already it's clear that's not going to be the case. And so I think what's going to happen is I think we're going to get first of all better at training these and better at fine-tuning them. I think a lot of what's missing there is a reasonable QA stack. Like regression testing for fine-tuning an LLM is not something that we really know how to do well at all.
People have ventured proof of concept kinds of things, but I think that's really scary about, especially as you go down market, if you're not Google that can afford to hire a thousand people to do this. It's really tough if you're let's say a regional bank or something to fine-tune LLM and just hope it works out. Because the problem with modifying these things is maybe you're teaching it something new, but you're almost necessarily letting it forget something. And that is a real challenge. And so I think that we're going to get better at that, about actually curating these things ourselves, whether they're some of the small language models and things so they can be deployed on things like our mobile phones. But I do think that what's going to happen then is rather than one big winner, I think we're going to get better at specialization and training and making these things work better and faster.
And so just as we said earlier, that technology is the business now, I think that generative AI will become an important enabling capability for most of our workflows. Before I think they were like, "Oh, they're coming for our jobs." I just don't think that's the case. I think it's going to make us better at our jobs. 100 years ago, other people have made this example, but you maybe own two pairs of clothes. And we got better at how we harvested and manufactured textiles. And it's not that all the people that did that lost their jobs. What happened instead is now we have a closet full of clothes. I think it's going to be similar. I think one example I think about is legal review in a business today. I don't go bother the general counsel, unless there's something really important going on, because it's expensive. Particularly, I have to use external counsel, it's really hard.
But if there was some kind of workflow where generative AI was involved, it becomes a lot cheaper to just almost have the lawyer review any external communication. Because it runs through a filter, it can say something. Maybe it's like, "Oh, this is weird enough, I need to flag this for a human." But I think those are the kinds of things we're going to get a lot better at. And so these experiences are going to get built into everything we do and augment almost every aspect of what particularly knowledge workers, but I think many others do as well. I think that when we manufacture things, I think that incorporate these applications to provide good feedback, to process and to operations in real-time is going to be really important. And so for the refutable part, then, yeah, I think that we're going to use these technologies a lot, but we're going to do it again and again.
So having common operating models and development patterns for how we build and test the LLM or the SLM or whatever and where it goes. But I think we're going to get further and further away from these really expensive general purpose ones. I think that we'll have learned and It's like this idea that like, oh, ChatGPT and we're waiting for the new GPT whatever to come out. That I think will still always matter from a GWIS standpoint, but I think there's going to be less and less of these really big ones. And I think we're just going to get better at having more specialization.
David Joy:
No, I was going to say. I agree with some of the thoughts you have. I like, to your analogy of we got better things and there was more opportunities and things like that. I feel like LLM are only going to enable more people to be able to doing things in a different way. I like Walter. I don't know if you follow Andrej Karpathy, he's one of those open source figures in AI, he used to work at Tesla, the OpenAI founder. He likes to say, LLM OS, this whole idea of LLM becoming a new operating system and you building integrations and systems on top of that. So I would love to just go and ask an LLM, "Hey, build me a Confluent pipeline that does this for me, as simple as that, and it can do something similar."
Andrew Sellers:
And then deploy it. And that kind of thing.
David Joy:
For sure. That's what I'm excited about myself. We'll see how it all comes together in a decade, but it's super interesting.
Andrew Sellers:
You can come back and tell me I got it wrong. We'll have a follow-on episode, hopefully with many more before that, but we'll have one then too.
David Joy:
Yeah, for sure. I don't know what we'll be doing then, but definitely it'll be fun to talk about how it all shapes. But I want to transition quickly and come back into some of the really cool stuff that you guys are working on at Confluent. For people who are listening right now, what are some of the things that you're excited to preview or excite people about with what Confluent is doing to shape up the future and where things are going?
Andrew Sellers:
So one really exciting feature that we announced about a week and a half ago at Kafka Summit London was our new Tableflow feature. And what it does is we always had this idea of this operational estate where we had backend developers, and then we had this analytic estate where we had data engineers and data scientists that would then find something out about the business. And information tended to flow in one direction, and they tended to use different tools and different data primitives. And it was all so really expensive and hard and error-prone to move that data around. And so with TableFlow, what we've done is we've taken another emerging standard. We have Kafka and we have Flink, but we've looked at what Iceberg has accomplished as this really amazing headless data architecture for doing OLAP kinds of work. And what we've released to the world now is early access that we're going to rapidly bring to market is a GA capability of where we're actually materializing Kafka topics as Iceberg tables.
And so for the people that want to consume their data as a stream, they can do that. For those that want to consume it as a table to do OLAP kinds of things, they can do that too. And what's really cool then is that we handle all the consistency. And so it was something that came out of a lot of the design of, again, our core, our cloud native Kafka compatible offering, where we work really hard to make that as fast and as cost-effective as possible. And that involved tiering the storage. And so you separate, compute, and store and you put whatever you can into object store, these sophisticated cloud primitives that are offered now that make storage a lot more affordable. And then that way then your streaming data no longer has to be ephemeral. You can store it forever.
You can do really cool patterns like event sourcing because you can replay it from where it all began. But then looking at the way we serialize that data, Iceberg ends up being a really good way to approach that. And so over time, we'll expose more and more to get to the point where you can even just subscribe to data from the Iceberg table itself. And so it'll be just one and the same. And I think that's going to be a really important capability for unifying those estates. And so reducing that friction that does exist today between the people that work in the data lake and the people that work on the operational apps.
David Joy:
That'll be really cool. I've not checked it out. I'm going to check it out myself. That sounds like a really good feature. If you guys are listening, go check out this update from Confluent on Tableflow to Iceberg. One of my thoughts while you were saying that is there's a project around that I've been following, and I have a bunch of friends who also move to the project and it's called Apache Pinot. Now, how does that project fit in overall data streaming story and data analytics for you guys?
Andrew Sellers:
So yeah, Pinot is an important technology and they're an important partner of ours. And so it allows you to do another way to do OLAP processing over Kafka data. And so it's another cool technology that makes sense for a lot of applications, and it fits very well into this ecosystem, and it's something that we're very proud to partner with. One thing that we look very hard at when we do this work is we don't want to have opinions on how our customers do their jobs. We want to make sure that in a sense we're Switzerland, we integrate with wherever data is, we can consume from it, we can publish to it. And so Pinot is just another one of those great technologies that we partner with. And for customers, it does a lot of magic things. And so we're very proud to be integrated with them as well.
David Joy:
Yeah. That's a great point of view. At the end of the day, what matters is what is it the use case is and what product or what feature really the consumer or the user wants and what product works. And you are focused on still the publishing part of it and the consumption part and wherever this needs to go. So you do the piping and wherever you want to build a pipe, we let you build it. That's the idea. I had one question to ask, and then I know we have to be cognizant of the time as well, but I just realized it's 30 minutes in. There have been a lot of changes happening around regulations, especially around data regulations and data privacy and how it's going to impact companies who are architecting these solutions. So how are you at Confluent preparing for that and also helping shape that data privacy and regulation part of things?
Andrew Sellers:
Yeah. So it's a challenge not just for us, but I think everyone in the data infrastructure landscape, that so much of how these technologies came to be, the way they were designed came from a different time. We were optimizing for availability and resilience, rather than thinking a lot about data locality and traceability. And so you look at the very definition of what the cloud was, the whole idea of it was, that's why that metaphor was the cloud. It was like the data's just out there, you don't know where, it's just out there. And that worked really well until we started thinking about data privacy. And so when you have a distributed system, things like replication and not always being deliberate about data locality are challenging for data privacy frameworks, things like GDPR and CCPA. If I have an immutable log, what does the right to forget actually look like?
How do I accomplish that? And so at Confluent, what we do is we want to put the power in our customers' hands. We want to provide them industry best practices that enable whatever flexible data modeling methodology they want to use, and to give them the features to assure the privacy of their data. So we recently announced client-side field level encryption. We have Flink actions that we just announced again a week and a half ago was a big announcement, enabling things like masking and encryption and synchronization over parts of data, so that however our customers want to protect that data and however they want to make sure that they can not just check a box for compliance, but really be assured that they're serving their customers in the way that their customers have put the trust into them, that that's something that we support.
And so I think that that's only going to become more and more important. And so I think with the way we think about where data goes and availability zones and how we architect the different offerings is something that is always top of mind for us. And I think the features that we offer are so critical to make sure that our customers can keep their promises to their customers regarding data privacy.
David Joy:
Right, yeah. And this is a topic that I know it's a hard problem to solve, especially for companies like Confluent as well as for CockroachDB. We are building distributed system that can scale. And I like your perspective on consistency and scale and resilience, but now we are reaching a point where this is essential. So I don't know if you know this, but in CockroachDB, what we can do is when you create a table, you can declare at the database level, you want to keep all your data locally or do you want to keep it globally?
Andrew Sellers:
I'm a huge fan of Cockroach, by the way. I didn't say that earlier. I should have. I think it's an amazing technology. And the way that it just runs, it's just incredible. So I've used it before and I'm a huge fan. I should have said that earlier.
David Joy:
Before I came to the company, I was really, again, it was like the Apollo, Siri, Alexa kind of problem. I was trying to do a project and solve it, and I was like, "Hey, I need consistency and I can't find it at the scale." And that's how I came across CockroachDB. I was a fan of the product before I came over here. But it's fascinating though. The data locality, that ability is fantastic, and I'm glad that everybody is pushing for that kind of an architecture. Where do you recommend, what are your recommendations for people who are trying to learn data streaming and event streaming and where everything is going, where do you think they can go and learn these skills? Is it that Confluent has a place to go learn this? Are there blogs? What is your recommendation?
Andrew Sellers:
We've got an amazing page I encourage everyone to check out, developer.confluent.io, where we have a bunch of free courses that are available to go through many of the fundamentals but also beyond. So there's a Kafka 101, Flink 101, but we go all the way through practical event modeling and data governance and data security and networking and everything that we hope people need to know in order to both operate and use data streaming technologies. And so you can go there too and sign up for Confluent Cloud for free.
We provide some credits as well to aid in the education, so you can actually do the labs right in the cloud. And that's, I think, a really wonderful place to start. It's where I learned a lot is from those technologies. Because it's really hands-on and we really try and emphasize the actionability, a path to applicability, I should say, in a way that is, I think, really refreshing. I think that, that is something that is... I know I've learned a lot and I came in knowing a lot of the technologies already, and I still learned things from those courses. So that was really cool. Yeah.
David Joy:
Oh, yeah. Great. My favorite place to learn is documentation and Slack channels.
Andrew Sellers:
Yeah, that's true. The community is really active and there's nothing like a community. To be honest, it's why I love the technology strategy group at Confluent so much, because we challenge each other, we push each other, we make each other better. And I think there's no substitute for getting involved with humans. And so yeah, check out the Kafka community too, because it's a great place to start your journey with people that have already been down the road before.
David Joy:
Very cool. Now, you brought up you learning. In your role, obviously you have to be thinking about Confluent and Confluent strategy, event streaming. But there are subsequently important changes in the industry is happening, and you have to learn about what's happening in GenAI. You have to learn about what's happening in the database market. So how do you keep up? What's your method to consume information and be on top of everything?
Andrew Sellers:
I think you read a lot, you talk to a lot of people. I think one thing that's very specific at Confluent that I've really liked that we do is we have our monthly market observations meeting. And so what we do is we basically have amongst the technologists people on this team, we rotate month to month. And so every month it's someone's responsibility to go through a curation of a lot of RSS feeds. And we have market intelligence Slack channels, and we look at just other key, we'll say leaders in different parts of the market that we try and look at the information from for data streaming, but also adjacent things, as you say, like GenAI. And we write up a really cool report for the entire company. And then we spend an hour briefing the report every month. And it's my absolute favorite meeting every month, because we go through what's new in the space, what others have come up with, and then we provide provocative, opinionated takes.
And I think that's really important too because again, sometimes, sometimes we're right, sometimes that's where they were going. Sometimes we find out six months later that we were completely wrong about why that acquisition happened or why that technology got released or how much that feature really took off or didn't. And so I think that making it part of the job and part of the organization and then sharing it, actually creating a product from it and not just passively consuming, for us has I think been really critical to keeping our pulse on what's going on everywhere else.
David Joy:
That's a practice I feel would be great for multiple organizations to have, especially if you're in a role like strategy where you have to be aware, but also not somebody who just sees a tweet or sees a post, but you also really think about how this applies to my technology. That's amazing. Well, Andrew, I know it's been a great 50 minutes of us talking. I just realized we hit 50. It's been an absolute pleasure having you here. Where can people follow you and do all the amazing things that you guys are doing? Tell us a little bit about that as we close.
Andrew Sellers:
I am most active on LinkedIn on social media, so please reach out there. It'd be great to connect with people. And I'm a frequent writer on the Confluent blog, and then I blog as in a guest capacity, other places around as well. But yeah, please, the Confluent blog is a great place to check out all the things that we're doing and a lot of what we feel is important in data streaming. And so we would love to have you along for the journey and really excited again to help data streaming add business value to whatever it is you're doing.
David Joy:
Once again, thank you for coming on, Andrew. It's been an absolute pleasure. And everyone who stuck around listening to this episode, it's been a pleasure. I will catch you in the next one. Have a great day.
A podcast for architects and engineers who are building modern, data-intensive applications and systems. In each weekly episode, an innovator joins host David Joy to share useful insights from their experiences building reliable, scalable, maintainable systems.

David Joy
Host, Big Ideas in App Architecture
Cockroach Labs
Latest episodes

Introducing Cockroach Continuum | A Big Ideas in App Architecture Exclusive
Tara Shankar Jana "TJ"
Senior Director Product Marketing @ Cockroach Labs

A Love Letter to the Database: Industry Shifts, Lessons Learned, and What's Next with Perry Krug
Perry Krug
Manager Solutions Architecture at Baseten

The Everything Trap: Building AI Software That Lasts with Sam Hilsman
Sam Hilsman
Co-founder and CEO of CloudFruit

Why Inference Engineering Is the Next Big Role in AI with Philip Kiely
Philip Kiely
Author of Inference Engineering | AI Education @ Baseten

Distributed Systems, Linkerd, and the Cost of Network Calls with William Morgan from Buoyant
William Morgan
CEO @ Buoyant, creators of Linkerd

Making Software as Durable as Data with Peter Kraft from DBOS
Peter Kraft
co-founder of DBOS

Breaking the Pillars: Rethinking Observability with Charity Majors
Charity Majors
Co-founder and CTO of Honeycomb.io and co-author of Observability

How to Transform Dev Workflows with CI/CS and AI Agents with Tomer Karin
Tomer Karin
Embedded Software Architect

AI, Market Cycles, and the Systems Built to Outlast Them with Cockroach Labs CEO & Co-founder Spencer Kimball
Spencer Kimball
CEO & Co-founder Cockroach Labs

How to Scale Data Infrastructure from Startup to Enterprise
Nishant Raman
Data Engineer at FinTech Company

How to Build an AI-Native Organization
Peter Mattis
Co-founder and CTO/CPO at Cockroach Labs

Inside Infrastructure as Code with Pulumi’s Founder & CEO
Joe Duffy
Founder/CEO at Pulumi

Inside Ericsson: How AI and Automation Are Shaping Telecom
Anand Bajaj
Chief Architect - 5G Network Slicing at Ericsson

Unboxing the Cloud: AI, Microservices, and Resilient Databases
Jim Hatcher
Solution Engineer at Cockroach Labs

Strategic AI and Cloud Solutions: GitHub’s Blueprint for Modern Development Success
Ari LiVigni
Senior Cloud Solutions Architect at GitHub

Cloud Architecture in the Public Sector: Balancing Innovation and Security
Nick Mayer
Principal Cloud Architect at Maximus

GenAI Meets Celebrity: Inside Cameo’s Journey from Startup to Stardom
Dom Scandinaro
CTO at Cameo

The journey from mainframe to adopting generative AI with Equifax’s Senior Network Architect
Samarth Shah
Senior Network Architect at Equifax

Modernizing your cloud strategy with OneStream’s Senior VP of Cloud Architecture
Ryan Berry
Senior VP Cloud Architecture at OneStream Software

Driving digital transformation with Chief Architect at Altimetrik, Ignacio Segovia
Ignacio Segovia
Chief Architect at Altimetrik

Discussing the Patterns of Distributed Systems with Unmesh Joshi
Unmesh Joshi
Principal Consultant at Thoughtworks and Author of Patterns of Distributed Systems

How to simplify your software architecture
Rob Reid
Technical Evangelist at Cockroach Labs

Behind the scenes with Vimeo’s Director of Enterprise Architecture
Sachin Joshi
Director of Enterprise Architecture at Vimeo

How to leverage real-time data processing for enterprises
Andrew Sellers
Head of Technology Strategy at Confluent

Inside the Mind of the Chief Architect at Index Exchange
Joshua Prismon
Chief Architect at Index Exchange

Solving for Scale: Real-time Retail Experiences with Endear's CTO
JP Grace
Endear

Data, Acquisitions, and AI: Insights from FiscalNote's CTO
Vlad Eidelman
CTO and Chief Scientist at FiscalNote

Discussing Data Trends in the AI Era
Gajanan Chinchwadkar
CTO at Hypermode

Unwrapping Moonpig: Architectural Insights into Personalization and Scalability
Alexis Lowe
Principal Engineer at Moonpig

Solving for data intelligence at scale
Madalina Tansie
Chief Technology Officer at Collibra

Simplifying solutions architecture with Brian Johnson of Booz Allen Hamilton
Brian Johnson
Sr. Solutions Architect at Booz Allen Hamilton

How to make your applications smarter
Rod Senra
VP of Engineering at Loadsmart

Scaling for 2 billion events per day with Principal Software Engineer at Red Ventures
Majid Fatemian
Principal Software Engineer, Data Platform at Red Ventures

The data behind digital marketing: A conversation with Bluecore’s Software Architect
Mike Hurwitz
Software Architect at Bluecore

A Lesson in Scaling: How Kami handled 25x growth with CTO and Co-Founder Jordan Thoms
Jordan Thoms
CTO & Co-Founder at Kami

Mastering Multi-Cloud with PwC’s Erol Kavas
Erol Kavas
Director at PwC Canada

From FedEx to Five Guys: Designing digital experiences with Yext’s VP of Software Engineering
Matt Bowman
VP of Software Engineering at Yext

Reliability and scalability in a data-driven world with Fivetran’s VP of Platform Engineering
Mike Gordon
VP of Platform Engineering at Fivetran

Enabling a data-driven and innovative engineering culture at Amplitude
Shadi Rostami
SVP of Engineering at Amplitude

How Estée Lauder scales strong engineering culture
Meg Adams
Executive Director of Platform Engineering at Estée Lauder

Can I take your order? Building conversational AI to improve the customer experience
Akshay Kayastha
Senior Engineering Manager at ConverseNow

Engineering resilient systems: Rescuing old treasures and unleashing modern capabilities
Marianne Bellotti
Author, Engineering Leader, Systems Geek

The Full Package: How Route architects its all-in-one post-purchase platform
Siddhartha Sandhu
Engineering Manager at Route

A historical journey in developer technologies
Mike Willbanks
CTO at Spark Labs

From Legacy to Cloud: Success stories from migrating mission-critical applications
Kishore Koduri
Senior Director of Enterprise Architecture at Ameren

Building purpose-driven engineering cultures
Jason Valentino
Head of Engineering Enablement at BNY Mellon

Modernizing Insurance Application Architecture at New York Life
Mike Murphy
Corporate Vice President and Life Insurance Domain Architect at New York Life

Innovation and Disruption: How Materialize pioneered a new era in data streaming
Arjun Narayan
Co-Founder and CEO at Materialize

Stories from an SRE: How Hans Knecht builds better developer experiences
Hans Knecht
Cloud Consultant at Knechtions Consulting (Ex: Capital One; Ex: Mission Lane)

Inside Chick-fil-A’s infrastructure recipe for a perfect customer experience
Brian Chambers
Chief Architect at Chick-fil-A Corporate

Modernizing from the Mainframe: An Exploration of Distributed Systems
Chris Stura
Director, PwC UK

IoT Standards & Data Mesh: Utility Facility App Architecture
Grant Muller
Vice President, Applications and Technology Architecture at Xylem

Relational Data Problems: Doubble Dating Application Architecture
Mattias Siø Fjellvang
CTO & Co-Founder at Doubble

From Legacy Systems to Limitless Scaling with Paycor’s Systems Engineering Fellow
Adam Koch
Systems Engineering Fellow at Paycor

How to Understand Problems & Build Better Software with Technical Leader Joe Lynch
Joe Lynch
Technical Leader

Observability in the Cloud & Dataflow Modifications with Yolanda Davis from Cloudera
Yolanda Davis
Principal Software Engineer, Data Flow Operations

Early Days at Google & Building CockroachDB with Peter Mattis
Peter Mattis
Co-Founder and CTO of Cockroach Labs

Database Benchmarking Efficiency with OtterTune’s Andy Pavlo
Andy Pavlo
Associate Professor of Databaseology at Carnegie Mellon and Co-Founder at OtterTune

Observability & Statelessness with TripleLift’s Chief Architect
Dan Goldin
Chief Architect at TripleLift

Understanding AI: PubNub CTO Stephen Blum’s Key to Faster App Development
Stephen Blum
PubNub

Building reliable systems with DoorDash's Matt Ranney
Matt Ranney
DoorDash

Real-Time Data Capturing: The Future of Fitness Technology
Paul Lawler
Head of Software at Wahoo Fitness

Building Efficient App Architecture with Alloy Automation’s Gregg Mojica
Gregg Mojica
Co-Founder and CTO Alloy Automation

Unleashing the Power of Hiring Software with Greenhouse CTO Mike Boufford
Mike Boufford
CTO at Greenhouse Software

Decoding Data Warehousing: Insights from Ken Pickering, SVP of Engineering at Starburst Data
Ken Pickering
Senior Vice President of Engineering, at Starburst Data