
46
Discussing the Patterns of Distributed Systems with Unmesh Joshi

Unmesh Joshi
Principal Consultant at Thoughtworks and Author of Patterns of Distributed Systems
Ever wondered how the world's most resilient applications stay up and running, no matter what?
This week, Unmesh Joshi, Principal Consultant at Thoughtworks and author of “Patterns of Distributed Systems” joins me for a conversation on the technical intricacies that keep modern applications running smoothly.
Join as we discuss:
Key concepts behind distributed systems, including fault tolerance and consistency models, consensus algorithms, and more.
Unmesh’s inspiration for “Patterns of Distributed Systems” and how he aims to simplify complex concepts for developers.
Technical predictions for distributed systems over the next five years.
David Joy:
What is up everyone, and thanks for tuning in. In today's episode of the Big Ideas in App Architecture podcast, we speak to Unmesh Joshi, who is a Principal Architect at Thoughtworks, and the author of the book Patterns of Distributed System. Unmesh and I get into his book, and talk about some various patterns he has observed running distributed systems. We also get into some really interesting findings, and his advice to folks who are running, operating, and want to learn these distributed systems. So pump up the volume, and get ready for an intriguing conversation with Unmesh Joshi. Welcome to the podcast, Unmesh, how are you doing today?
Unmesh Joshi:
Great, great, thank you David.
David Joy:
For everyone listening, Unmesh Joshi, you are a Principal Architect at Thoughtworks. Why don't you tell the people a little bit about yourself?
Unmesh Joshi:
My name is Unmesh Joshi and I work as Principal Consultant at Thoughtworks. Been with Thoughtworks for over 16 years now, and overall coding software for about 24 years. I started coding when UML was big, late '90s. And for last I think several years, I'm keen on understanding, and teaching and exploring distributed systems. So the outcome of the last four years of effort in figuring out how to explain distributed systems, I documented them as patterns and now it's available as a book.
David Joy:
Everyone listening, Unmesh wrote this book called Patterns of Distributed Systems. And it's a book about understanding how these different systems that are there in the world, say that use Paxos or Raft, these are consensus protocols, how they're very similar. And while in the literature there is an understanding on how this works, but the actual implementation is something that nobody gets to talk about. And that's what I was fascinated [inaudible 00:02:08]. Coming into the conversation. I did not know your background or what you have been working on. And what I could attach myself to immediately was my own experience. For the last 5, 7, 8 years I've been working with distributed system. I was working with the Cassandra Community. I've been an avid user of Kafka for almost eight, nine years. I was using Spark for my data AI stuff, and then now I work at Cockroach Labs, and we use Raft and we have a distributed system. And we call ourselves the distributed SQL leader. So it's been fascinating for me myself to experience all of this, and it was really interesting to read your perspective and understand this book. So why don't you explain, or tell me a little bit about what inspired you really to start going through and writing this book itself?
Unmesh Joshi:
Yeah, so to start with, for last I think tenish years, 10, 15 years since the cloud movement started, all of us software engineers there, we use various cloud services and products like Kafka and Cassandra and Mongo and stuff like that. And I was always curious to understand more about why now, and how they're built and stuff like that. But about I think six, seven years back, I was engaged in a very interesting project to build software for an optical telescope, which is called a 30-meter telescope. It's an ongoing work, it might be live in 2035 probably, but we were writing core software systems for that. And what happened was, a lot of decisions we had to take, how to component registration, and how to do discovery, and how to make sure the work is distributed across multiple components. All of that was very similar to what a typical distributed data system would do, a Kafka, or a Cockroach DB or a Kubernetes. But obviously, we couldn't use any of that off the shelf. So we had to write our own code, which I thought would look very similar to what these other code bases do.
And I was obviously trying to make sure that the decisions that we take, they're not bad decisions. They're actually production-tested and stuff like that. And obviously when you need to understand distributed systems, the path is difficult because you need to go through several academic papers, or very theoretical books. And that's probably because of the history of distributed systems, that obviously proving something is correct and something will work with the assumptions you have. It's difficult. And the focus always had been to make it theoretically accurate, and mathematically provable and stuff like that. But the challenge with that literature and with academic papers, is it's very hard to see how all that connects with the code that you need to write. And one thing I tried doing then was really read open source code bases. Because one of the good things in last 10, 15-ish years is there are a lot of open source products like CockroachDB, Kafka, YugabyteDB, Kubernetes. And reading that is helpful. But obviously when you read a massive code base, that's the other challenge, because you don't know where to start and what to look for.
David Joy:
Needle in the haystack problem. Yeah.
Unmesh Joshi:
Yep. Yep. So there are two extremes. One is read the code, and one is read a theoretical description of something. So I thought because every distributed system is so similar, but there are differences. And when you need to explain this similarity, not the exact implementation, but a similar implementation, and give some context behind why something is implemented the way it is, patterns is a very good approach. And what I tried doing was, I tried changing my role from reading other people's code to becoming an implementer myself. So I changed that role and said if I need to implement something like CockroachDB let's say, or something like Kafka, where would I start? And how would I start implementing this? And I created my own miniature toy versions, obviously not production ready. But what I made sure was that the core structure remains similar to what I see in a GitHub label.
And I used this for teaching distributed systems, because we had some failed attempts at teaching distributed systems to developers, and that worked pretty well because you're not talking about theoretical abstractions and maybe just more real. Yeah, you talk to people saying, "Let's implement Kafka together." And it doesn't nee to be production ready. Let's start with one problem at a time. And maybe some of that came naturally because ThoughtWorks is a big proponent of extreme programming, and practices like test driven development, incremental design and stuff like that. So I think that focusing on one problem at a time, trying to solve that, and then look at next problem and then try to solve that. I think that came naturally, and that worked pretty well with people when we were discussing several implementations. And I talked about this with Martin Fowler and he liked the idea.
He said, "Yeah, distributed systems is something hard to learn, and hard to explain and hard to think about. So why not we maybe document it a little more formally as patterns?" And that's how it started. So that was the inspiration, and pairing with Martin was a great experience. He is one of the first people who was part of this pattern movements. So that worked pretty well, and we had been collaborating for nearly three years to incrementally publish. So we published this pattern, and we wanted to make sure that what we write about is actually validated. So we published all these patterns over three years on Martin Fowler's site, Martinfowler.com. And they were read by a lot of people. We got reviews. And after we got a significant chunk, around 30 patterns and we thought now they're good enough to explain a particular system. If you need to explain how Kubernetes works, or how CockroachDB implements some things like with hybrid clocks, this is a good enough chunk to be published. And that's when it's published in Martin's signature series.
David Joy:
When I was researching you, I ended up also researching Martin in the process. It seems like the way the book is structured, is also that you don't necessarily need to read it page one to end. You can go between patterns, so it's easy for somebody who wants to understand a specific topic.
Unmesh Joshi:
Right. Right. Absolutely. And that's I think one good thing, and also a difficulty with patterns books. Because you can't read patterns books from start to finish. Let's say exploring CockroachDB code base, and you need to figure out how versioning is implemented in CockroachDB, and you know it uses hybrid clock. Then you can probably go and read hybrid clock pattern, or you read a version value pattern where you see how in [inaudible 00:11:11] a storage you implement version data. So it helps that way. So what we have though, is we have first two chapters, and then second chapter is just a whirlwind tour of all the patterns. So it's a big 50, 60 page chapter. But we have it there to help readers so people can read, let's say first and second chapter, and then visit a pattern based on what they want to figure out.
David Joy:
Yeah. And what I enjoyed, and the reason why I ordered the book personally, is because I haven't been a user of the products. So I've been using Cassandra, and I've been using say Kafka, I've been using Kubernetes, and [inaudible 00:11:55] XCD, you know Raft is there, Cassandra I used to use as back source. And I would do the necessary understanding required, but I would not go way deep, or go into the pattern. And what I feel your book reveals, is an understanding on how these systems are similar, and would give anybody who wants to learn a little bit more a step into that world.
Unmesh Joshi:
Yeah, yeah, no, and that's one of the important motivations as well. Because one of the things that has also become problematic, I would say in the cloud era, is that there is a strong boundary between people who build products, and people who use these products to implement solutions. And there is always a problem when this kind of thing happens, because the people who implement solutions using these products, there is no way for them to learn the technical principles. They highly rely on various certifications that product vendors, like AWS certification on any number of services that AWS has. And that I think in general, creates a problem. Because when you build a solution, you don't use just one product, you use multiple of those. And you obviously can't do certification on every product you use on every project. So you need a way to understand the fundamental technical principles with which these products are built, and you understand it. Again, one thing I liked about the patterns approach is, you don't give just a theoretical description, or you just don't draw diagrams. You can show actual code. And a code which is not exactly the same, but very similar to what you would expect in an actual product.
David Joy:
Right. So it's an implementation as a small MVP?
Unmesh Joshi:
Yeah, absolutely. And other thing it helps with, is you start getting questions. And I think getting right questions is the first thing to start with a good implementation, because only when you get questions you'll look for answers.
David Joy:
And also it's interesting, first of all, this book that you have explained, it's you have put at least four or five years into doing this, and validating this through feedbacks. And that's a lot of work. And I think for any technical person, sometimes the source of truth sits in when you actually write the code yourself. You can watch a code, somebody writes it. But when you really write the code yourself, and when you see the output, you're like, "Okay, this is why it works like this. Okay, I need to do something else here." Yeah.
Unmesh Joshi:
Absolutely. And what happens is, particularly with a topic like distributed systems, there is fear in the mind of developers. So one thing you start talking about Paxos to someone. One is only, let's say, 2% of the developers know about pack sauce or something like Paxos. But people who know about facts Paxos, only thing they know is, "Okay, that's way hard to understand." It doesn't need to be that way. When you do simpler implementation of Paxos, it's not that hard. It's not really hard to understand. The way Paxos is explained, the thing is that it talks about... Again, and that goes back to having a theoretical abstraction and proving that it works. So all Paxos literature primarily is about proving that there is agreement on single value. And agreement on single value is never what you would do in let's say a key value store.
You'll have multiple values obviously. And to bridge that gap, what Raft essentially did is important. I think one of the things, or one of my hopes is that once you see this real code and see how Paxos is implemented, you're no longer afraid of trying something out. And then it also gives you the questions, as I said, right? Because Paxos, the way it's described, and the way even if you implement it, what you can achieve at the most is you set a single value. And it's actually designed for that. And obviously the question then is, if you need to update the value, what do you do?
And when you read that, okay, Cassandra implements Paxos, but let's say Cockroach DB implements Raft. Now what does Cassandra do if in a database you need to update a value? You can't initialize only one anything for a given key and value. And what Cassandra does is implement something outside of basic Paxos description. They handle committed values in phase one of Paxos, which is now undocumented. But you get that question, and then you say, "Oh yeah, even if they're saying they implement Paxos and they don't need Raft, they actually need to do something else." Yeah, other stuff required.
David Joy:
Yeah, there's something else. There's some other stuff required. Yeah, but I think you brought up a good point, and I was going to ask that question. We started talking about it a bit earlier as well, Paxos in literature. When I was reading the white paper myself a few days ago, I was like, I stopped reading it after a few pages. I don't get it. But when I started reading Raft, I started getting it. Because it's way more simpler. And I think there's a pattern too, right? Raft is more open source friendly. A lot of open source projects went the Raft route.
Unmesh Joshi:
Yeah, I think it's more implementation friendly I would say. Because a lot of decisions that you need to take while implementing something, they're explicitly documented in Raft. But one of the things where I think this patterns approach works is, you read Paxos and then there is this concept of ballot there, and you read a draft and there is this concept of term there. And you read view stamp replication, and then there is this concept of view number. And they're exactly the same concepts, but no one explains what that concept is in isolation. So it is called as fencing token sometimes, but essentially it's to make sure which request is latest, and which is stale, or which request is originated with the latest reader. And it's a slightly enhanced concept or implementation of Lamport's clock, because you increment that number with quorum voting.
So instead of calling that as a fencing token, I call that as a generation clock. So before you start communicating, you need to establish yourself as a leader for a particular generation. And for doing that, you need to establish this value from generation clock. And once you know that, that you need to establish this unique number, which is monotonic, which is higher than any of the previously assigned numbers. And you need to do that with quorum voting, because there is no other way if you're not relying on some external source. And you need to do that because there are partial failures. That's one of the most difficult things with distributed systems. And with partial failure can always be an older leader, which is sending requests, and you might receive those requests late, and after a new leader is there. And you need to cut off that old leader, or you need to ignore those requests. And for that, you need this monotonic number. And I have explained that as a pattern, because it is a pattern.
David Joy:
What chapter is this, again, in the book?
Unmesh Joshi:
It's a pattern called Generation Clock.
David Joy:
Okay.
Unmesh Joshi:
Yeah. And once you know that, then if you look at, for example, Kafka's implementation with Zookeeper. And when you use something like Zookeeper, or generic product like Zookeeper, you need a controller which orchestrates your cluster coordination. And the same problem that a Paxos proposer, or a Raft leader faces this controller faces now, right? Because controller is sending, and figuring out how to assign partitions, how to select partition leaders and stuff like that. And all the requests it sends, they need to be tagged with this monotonic number, because it can always be possible that you have an older controller and a new controller is elected. So then it becomes, I would say, easier to now see the Paxos, and Raft leader, and the Kafka controller or Kubernetes as well I would expect to do similar things. But the problem is same, and then it becomes easier to see what's happening there.
David Joy:
Correct. Yeah. Yeah. And what I liked about what you said is if you look at it at a... Because that's the level at which we operate sometimes. The problem to solve is essentially the same. How do they coordinate? And this goes back to the question I was going to ask you, is also distributed systems, we're trying to solve a problem that centralized systems that couldn't scale solve. But we have to solve the same things a centralized system does, now at a distributed scale. So that's where the complexity adds comes up. So my question to you is, what was one of the things that really surprised you when you were trying to go through these patterns? When you started reading, the myths started going away and you're like, "Okay, I get what's going on." What surprised you?
Unmesh Joshi:
Oh yeah, absolutely. You rightly said a lot of this work is demystification of concepts that look very foreign to you. So I think the most surprising thing, why is this not explained earlier by someone else? Because distributed systems is taught for at least 20, 25 years in academia, and somehow it's not a common knowledge. Even now, even after the cloud is everywhere and almost ubiquitous, everyone uses distributed systems for the last 10 years. People don't know, and there is no easy way to learn as well. That's other amusing thing. And that's not just with the masses. But what was surprising was that before Raft thesis was written, view stamp replication algorithm was documented. And it's almost Raft, but it was not well known even in the academic circles, and you needed a PhD thesis to document something.
David Joy:
Yeah, that's the surprising part. I feel like, there's a lot of these implementations, these papers were around for a while, but I think the systems and the supporting systems did not exist in a way that they could be built in a way. I wanted to jump to this question. Now that you have spent this time writing the book, and learning these patents. How do you explain fault tolerance and consistency models to people now? Your understanding is so much more different now, right?
Unmesh Joshi:
Yeah. So one of the things that's pretty obvious, if you look at these patterns a little more closely, is that fault tolerance and consistency, particularly consistency is always in the context of what the expectation of the consumer is. And as long as you can manage that, it's fine. Whatever happens on your storage service, for example, it doesn't matter, as long as it's not visible to the visibility the consumers. So I think that focus on what the expectation of your consumers is, and to understand that, understand that well. I think that becomes an important talking point, rather than focusing on wrong things, I would say.
David Joy:
I think it's interesting, because we have to solve the problem at scale. So in a distributed system, you're looking at a lot of patterns, requirements like scalability, fault tolerance, you want it to manage the latency. And then you have to make sure that there is encryption, and communication between these nodes in the cluster and a bunch of things related to that. And then we hear the concept of trade-offs, right? A system has a trade-off on consistency. Another system says, "Okay, I'm going to trade off partial partitionality." And then we go back to the cap theorem in a sense, yeah.
Unmesh Joshi:
So one of the reasons why I brought up consumer expectations, is that for some situations you cannot tolerate eventual consistency. And I think for that reason, I don't necessarily like cap theorem. Cap theorem gives you some thinking tool to explain why certain decisions are taken in certain systems. But from a consumer's point of view, for example, if you look at your cluster metadata where it is stored now, cluster data, metadata needs to be consistent. And it cannot be eventually consistent, because you can't have two nodes in a cluster having two different views of how your cluster looks. And this is a problem problem even in systems like Cassandra. And one of the interesting things is that Cassandra is implementing now something called a transactional metadata, which is very similar to what something like CockroachDB or YugabyteDB already has. You have a smaller core with all the metadata. And when you are solving that problem as someone who is implementing the cluster, it has to be consistent. You cannot tolerate that to be inconsistent, or eventually consistent for that matter.
David Joy:
Yeah, so it seems like my understanding of that is, it came out when I was leaving the company as well. I was joking with my Cassandra, "You bring consistency when I'm leaving?" It was funny, but it seems like from what I've read, and one of our friends, good friends is also a PMC chair. And he was telling me that the Accord paper essentially simplifies the Paxos implementation, and brings in some of the things that people have been asking for with Cassandra. But one of my own personal experiences was, the reason why I started using CockroachDB and eventually needed that, was because I was trying to solve a use case personally, and I needed consistency. And I tried to use multiple systems and it was just not working. I tried eventually, it did not solve the problem. And then I fell in love with the implementation and the simplicity with which they implemented CockroachDB. And also I like the concept of follow reads. Going back to your consumer expectations, expand on that a little bit about. I know you have written a whole chapter on follow reads. Yeah.
Unmesh Joshi:
Right, right. And that's another interesting thing when you talk about consumer expectations, because generally what people think is once you implement Raft and two-phase commit, you always have consistency. And consistency in the sense if I write something, a very simple thing, read your own writes, for example. That you make some changes and then you read. So you expect to get the same value back. But with Raft or any other consensus mechanism, if you are reading from two different nodes, you read it from leader and then you read it from follower, you don't necessarily get the same value back. Unless some extra precaution is taken, and you need to take that precaution. And you need something extra beyond basic consensus. And then you need to know in what situations you can possibly tolerate that, so that you can distribute your reads to the nodes, which are probably closer to you, or closer to your users.
But otherwise, you need to have something extra. Make sure you manage some token, let's say with every request, and then you either direct the request to the leader to make sure that the user gets the values that the user expects to make. But again, that's a trade-off. And I'm using that word, but yeah. And particularly for the distributed systems that we are seeing now, which are really globally distributed. And I think CockroachDB, all these modern new SQL databases, you can literally have database spanning across geographies, where latency, network latency can kill you really. And that again, most use cases are read heavy, and that's the nature of most use cases. And if you need to take advantage of that observation, that you can probably optimize on reads, and you can give a better experience to your users by directing their reads to the closest node which is in their geography, so that you don't get hit by 100 millisecond latency to other continent. That's doable. And I think that's what most databases do.
David Joy:
You can tune the latency a little bit depending on department. One of the things that maybe somebody should write a book on, the patterns of the time actually is a symptom of-
Unmesh Joshi:
Yeah. So I do have a section on that in the book. And the trigger for that was again, you know the famous Spanner paper, and Google's true time discussion and stuff like that. And you see in practice CockroachDB, and YogabyteDB, and MongoDB even. MongoDB doesn't exactly implement this, but these are globally distributed databases working in production. But no one explains what exactly happens, because you read a Spanner paper and you get the impression that, "Okay, only Google can do it. We are common men, we can't achieve that." But essentially the trick that's there, and one key interesting thing to understand is, even with Google and TrueTime, you cannot achieve exactly same timestamp on two nodes which are globally distributed. But what you get is a idea of what the maximum clocks queue can be, and that data, and clock bound.
David Joy:
The difference between.
Unmesh Joshi:
So the clock bounds you get, and that's what Google achieves. But in your implementation, let's say based on your observations, you see that okay, the clock bound is not more than, let's say 100 milliseconds. Then you can implement very similar implementation a bit inverted, because rather than waiting on commits, let's say you figure out if there are conflicting things going on while you read. So you don't wait, but you retry. But once you understand what's happening, you can think through. Okay, so with TrueTime there essentially doing this thing. And if I don't have TrueTime, I can do the same thing.
David Joy:
What else can we do?
Unmesh Joshi:
Yeah. In a slightly different way. So my focus should not be on making sure my clocks are exactly the same, and going for atomic clocks, and GPS and stuff like that. But I need to-
David Joy:
But that's such a fascinating problem to solve. When I came to Cockroach Labs, and one of the first papers that was recommended to me is an article written by Spencer, who's our CEO, around atomic launch. It was fantastic to read. That took me a few days, a few weeks and months to really understand it. But I was curious to talk to you about some of the things that you mentioned just before. It's the scale at which we're operating is because of the times we're in. In the sense, again, a time-related idea. But it's like users back in the day, or when we were designing systems like say for example, you have a catalog, and you want everyone in India to see the same catalog as you. So you can have a distributed system that runs in say Mumbai, and Delhi, and Calcutta, but it's a global metadata. Everybody sees the same thing, but when you try to update that, it'll take time. So those are some interesting patterns I've seen, and I think we've reached a point in expectations by the consumer spread, they want performance to be low. So what you said, within a region, you can have more copies of the data, and you can get consensus more quickly, and get quorum and get reply backs and things like that. Yeah.
Unmesh Joshi:
No, absolutely. And whatever techniques you use, or how much advanced your infrastructure becomes, you cannot fight physics. So the speed of light is the limit for any message to go from one place to other, and you can't fight that. I was having some interesting discussion about disaster recovery, because one of the fascinating things about CockroachDB or YugabyteDB, and essentially architectures which use hybrid clock and not rely on Oracle. But one of the advantages of that is there's flexibility in the various topologies that you can have, cluster topologies I'm saying, cluster topologies. But essentially, when we say globally distributed, you can't have your rights going to three geographies, and expect then your throughput also to go up with distribution. And that understanding I think you need to have, and then how to tune. So you get various options with these database architectures, but you need to be clever enough to use them appropriately based on your use cases.
David Joy:
Right. And you brought up a point that when I started reading the sample on your book, what I really liked about the way you started making it simpler for people where you start by saying, "We have to consider four variables in any system, a network, compute, memory and storage." And then you made it very real by saying that if you have a one gig network bandwidth, that means you can do 25,000 transactions or requests per second, or something like that. And we don't think like that. A developer doesn't think like that, but as an architect, when you're designing a system and you're thinking about where business is going, you want to consider these variables that we don't understand. So I really enjoyed that you wrote the book with some validation around what it really looks like in logic. Yeah, very cool. Well, I wanted to ask you as we come to a close to this particular episode, where do you think distributed systems are going now? In the next five, six years, what do you think are some emerging trends that you are observing? And let's make some bold predictions right now.
Unmesh Joshi:
Yeah, it's tricky to predict always because things change. But one of the things, again, any design, or any architecture that's proven or that's used in practice is, it boils down to what kind of hardware restrictions you have. For example, current computer architecture is, you have CPU, memory, and disk as the slowest component. And any process to process communication, you do it over a network. But let's say if those assumptions go away, you have something like a persistent memory, or let's say within your data center region or zone, you can use techniques like RDMA, Remote Direct Memory Access. So you don't go over the network, but you use something like RDMA.
And then things can be very different, change drastically. And I was following up with some research implementation called Derecho, University of Cornell. And they have essentially implemented Paxos using RDMA. And I was experimenting a bit with Intel Soft in a couple of years back. They shut down the product, but it was an interesting idea, have persistent direct access memory. Essentially, that changes your implementation upside down because you no longer require something like write ahead log, in some cases. But these are some. So again, I think that's an important thing to understand, that any software architecture that's done either by your database or any other product, it's within the constraints in which it operates. You have constraints of the computer systems that your product operates in. And if some of those constraints they change, or they go away, the implementation will be very different.
David Joy:
Very different. And I think you brought up a good point, because we work with customers, and we keep seeing some wild edge cases in this [inaudible 00:41:48]. And that's a good point to bring up in this context, is because when you take something from literature and you start building an actual system, then that consumers and customers have to start using, we come across certain failures or certain limitations within the literature. And then that's where smart engineers come in, and we build it like a system like what we have so many amazing systems in the market today. I say this to, I have had folks on the podcast and we talk about this idea that 20 years ago, all you had was a few systems, and all of them were closed, you had to pay licenses, and for learning was not. But the last decade is an amazing time for us to use technology.
And what you have done with this is a very extremely amazing use case. Where you've looked at what's available open source, learned it, turned it into such a great education for everyone. So it's really awesome. So for anyone listening, if you really want to dig into this a little bit more, go check out Unmesh's book, which is on Amazon, Patents of Distributed System. It's really good. I feel like it's good for people like us who even are involved in talking about our own individual product with customers. It gives us a much more deeper depth without going and doing a PhD thesis. So thank you so much for that. I'll leave everyone with this last question for you, Unmesh. What's your advice to practitioners of these systems, and who are trying to build these systems and implementing them?
Unmesh Joshi:
Yeah, my advice is there is so much of an amazing source code available for distributed systems. So don't limit yourself to theoretical books. And one of the goals for my book, or when we documented patterns, was that it should be easier for people to read the source code. And I think that's what everyone should do. So code reading is a great way to learn.
David Joy:
And in your book, you have put some code and you've tried to explain that as well. I think it's written in Java, which-
Unmesh Joshi:
It's written in Java. I mean,
David Joy:
Most people understand Java.
Unmesh Joshi:
Only reason was most people understand Java.
David Joy:
Right. And now it's way more easy. Just take a photo of the code, and just ask Chat GPT to convert it into Python or something. Way more easy. Awesome. Well, we should come back again and talk more. I think there's so much to unravel here, and I've learned so much just spending this time with you. Thank you so much for coming on. Wish you best with your book, and all the other projects that you're doing. And hopefully we'll have you again on the podcast. Thank you again.
Unmesh Joshi:
Yep. Thank you. Thank you, David. I enjoyed the discussion.
A podcast for architects and engineers who are building modern, data-intensive applications and systems. In each weekly episode, an innovator joins host David Joy to share useful insights from their experiences building reliable, scalable, maintainable systems.

David Joy
Host, Big Ideas in App Architecture
Cockroach Labs
Latest episodes

Introducing Cockroach Continuum | A Big Ideas in App Architecture Exclusive
Tara Shankar Jana "TJ"
Senior Director Product Marketing @ Cockroach Labs

A Love Letter to the Database: Industry Shifts, Lessons Learned, and What's Next with Perry Krug
Perry Krug
Manager Solutions Architecture at Baseten

The Everything Trap: Building AI Software That Lasts with Sam Hilsman
Sam Hilsman
Co-founder and CEO of CloudFruit

Why Inference Engineering Is the Next Big Role in AI with Philip Kiely
Philip Kiely
Author of Inference Engineering | AI Education @ Baseten

Distributed Systems, Linkerd, and the Cost of Network Calls with William Morgan from Buoyant
William Morgan
CEO @ Buoyant, creators of Linkerd

Making Software as Durable as Data with Peter Kraft from DBOS
Peter Kraft
co-founder of DBOS

Breaking the Pillars: Rethinking Observability with Charity Majors
Charity Majors
Co-founder and CTO of Honeycomb.io and co-author of Observability

How to Transform Dev Workflows with CI/CS and AI Agents with Tomer Karin
Tomer Karin
Embedded Software Architect

AI, Market Cycles, and the Systems Built to Outlast Them with Cockroach Labs CEO & Co-founder Spencer Kimball
Spencer Kimball
CEO & Co-founder Cockroach Labs

How to Scale Data Infrastructure from Startup to Enterprise
Nishant Raman
Data Engineer at FinTech Company

How to Build an AI-Native Organization
Peter Mattis
Co-founder and CTO/CPO at Cockroach Labs

Inside Infrastructure as Code with Pulumi’s Founder & CEO
Joe Duffy
Founder/CEO at Pulumi

Inside Ericsson: How AI and Automation Are Shaping Telecom
Anand Bajaj
Chief Architect - 5G Network Slicing at Ericsson

Unboxing the Cloud: AI, Microservices, and Resilient Databases
Jim Hatcher
Solution Engineer at Cockroach Labs

Strategic AI and Cloud Solutions: GitHub’s Blueprint for Modern Development Success
Ari LiVigni
Senior Cloud Solutions Architect at GitHub

Cloud Architecture in the Public Sector: Balancing Innovation and Security
Nick Mayer
Principal Cloud Architect at Maximus

GenAI Meets Celebrity: Inside Cameo’s Journey from Startup to Stardom
Dom Scandinaro
CTO at Cameo

The journey from mainframe to adopting generative AI with Equifax’s Senior Network Architect
Samarth Shah
Senior Network Architect at Equifax

Modernizing your cloud strategy with OneStream’s Senior VP of Cloud Architecture
Ryan Berry
Senior VP Cloud Architecture at OneStream Software

Driving digital transformation with Chief Architect at Altimetrik, Ignacio Segovia
Ignacio Segovia
Chief Architect at Altimetrik

Discussing the Patterns of Distributed Systems with Unmesh Joshi
Unmesh Joshi
Principal Consultant at Thoughtworks and Author of Patterns of Distributed Systems

How to simplify your software architecture
Rob Reid
Technical Evangelist at Cockroach Labs

Behind the scenes with Vimeo’s Director of Enterprise Architecture
Sachin Joshi
Director of Enterprise Architecture at Vimeo

How to leverage real-time data processing for enterprises
Andrew Sellers
Head of Technology Strategy at Confluent

Inside the Mind of the Chief Architect at Index Exchange
Joshua Prismon
Chief Architect at Index Exchange

Solving for Scale: Real-time Retail Experiences with Endear's CTO
JP Grace
Endear

Data, Acquisitions, and AI: Insights from FiscalNote's CTO
Vlad Eidelman
CTO and Chief Scientist at FiscalNote

Discussing Data Trends in the AI Era
Gajanan Chinchwadkar
CTO at Hypermode

Unwrapping Moonpig: Architectural Insights into Personalization and Scalability
Alexis Lowe
Principal Engineer at Moonpig

Solving for data intelligence at scale
Madalina Tansie
Chief Technology Officer at Collibra

Simplifying solutions architecture with Brian Johnson of Booz Allen Hamilton
Brian Johnson
Sr. Solutions Architect at Booz Allen Hamilton

How to make your applications smarter
Rod Senra
VP of Engineering at Loadsmart

Scaling for 2 billion events per day with Principal Software Engineer at Red Ventures
Majid Fatemian
Principal Software Engineer, Data Platform at Red Ventures

The data behind digital marketing: A conversation with Bluecore’s Software Architect
Mike Hurwitz
Software Architect at Bluecore

A Lesson in Scaling: How Kami handled 25x growth with CTO and Co-Founder Jordan Thoms
Jordan Thoms
CTO & Co-Founder at Kami

Mastering Multi-Cloud with PwC’s Erol Kavas
Erol Kavas
Director at PwC Canada

From FedEx to Five Guys: Designing digital experiences with Yext’s VP of Software Engineering
Matt Bowman
VP of Software Engineering at Yext

Reliability and scalability in a data-driven world with Fivetran’s VP of Platform Engineering
Mike Gordon
VP of Platform Engineering at Fivetran

Enabling a data-driven and innovative engineering culture at Amplitude
Shadi Rostami
SVP of Engineering at Amplitude

How Estée Lauder scales strong engineering culture
Meg Adams
Executive Director of Platform Engineering at Estée Lauder

Can I take your order? Building conversational AI to improve the customer experience
Akshay Kayastha
Senior Engineering Manager at ConverseNow

Engineering resilient systems: Rescuing old treasures and unleashing modern capabilities
Marianne Bellotti
Author, Engineering Leader, Systems Geek

The Full Package: How Route architects its all-in-one post-purchase platform
Siddhartha Sandhu
Engineering Manager at Route

A historical journey in developer technologies
Mike Willbanks
CTO at Spark Labs

From Legacy to Cloud: Success stories from migrating mission-critical applications
Kishore Koduri
Senior Director of Enterprise Architecture at Ameren

Building purpose-driven engineering cultures
Jason Valentino
Head of Engineering Enablement at BNY Mellon

Modernizing Insurance Application Architecture at New York Life
Mike Murphy
Corporate Vice President and Life Insurance Domain Architect at New York Life

Innovation and Disruption: How Materialize pioneered a new era in data streaming
Arjun Narayan
Co-Founder and CEO at Materialize

Stories from an SRE: How Hans Knecht builds better developer experiences
Hans Knecht
Cloud Consultant at Knechtions Consulting (Ex: Capital One; Ex: Mission Lane)

Inside Chick-fil-A’s infrastructure recipe for a perfect customer experience
Brian Chambers
Chief Architect at Chick-fil-A Corporate

Modernizing from the Mainframe: An Exploration of Distributed Systems
Chris Stura
Director, PwC UK

IoT Standards & Data Mesh: Utility Facility App Architecture
Grant Muller
Vice President, Applications and Technology Architecture at Xylem

Relational Data Problems: Doubble Dating Application Architecture
Mattias Siø Fjellvang
CTO & Co-Founder at Doubble

From Legacy Systems to Limitless Scaling with Paycor’s Systems Engineering Fellow
Adam Koch
Systems Engineering Fellow at Paycor

How to Understand Problems & Build Better Software with Technical Leader Joe Lynch
Joe Lynch
Technical Leader

Observability in the Cloud & Dataflow Modifications with Yolanda Davis from Cloudera
Yolanda Davis
Principal Software Engineer, Data Flow Operations

Early Days at Google & Building CockroachDB with Peter Mattis
Peter Mattis
Co-Founder and CTO of Cockroach Labs

Database Benchmarking Efficiency with OtterTune’s Andy Pavlo
Andy Pavlo
Associate Professor of Databaseology at Carnegie Mellon and Co-Founder at OtterTune

Observability & Statelessness with TripleLift’s Chief Architect
Dan Goldin
Chief Architect at TripleLift

Understanding AI: PubNub CTO Stephen Blum’s Key to Faster App Development
Stephen Blum
PubNub

Building reliable systems with DoorDash's Matt Ranney
Matt Ranney
DoorDash

Real-Time Data Capturing: The Future of Fitness Technology
Paul Lawler
Head of Software at Wahoo Fitness

Building Efficient App Architecture with Alloy Automation’s Gregg Mojica
Gregg Mojica
Co-Founder and CTO Alloy Automation

Unleashing the Power of Hiring Software with Greenhouse CTO Mike Boufford
Mike Boufford
CTO at Greenhouse Software

Decoding Data Warehousing: Insights from Ken Pickering, SVP of Engineering at Starburst Data
Ken Pickering
Senior Vice President of Engineering, at Starburst Data