
51
Cloud Architecture in the Public Sector: Balancing Innovation and Security

Nick Mayer
Principal Cloud Architect at Maximus
Modernizing government agencies with cloud technology involves navigating a delicate balance of challenges and opportunities. In this episode, we explore the world of federal cloud architecture and innovation with Nick Mayer, Principal Cloud Architect at Maximus.
Nick shares insights on overcoming the complexities of federal procedures and slow innovation in highly regulated industries, while harnessing the transformative power of cloud computing in the public sector. He offers practical advice on cloud migration challenges, leveraging AWS Lambda, and understanding the key differences between public cloud and GovCloud.
Join us as we discuss:
The importance of challenging norms and embracing change in regulated fields, despite slow procurement and risk aversion.
The varying approaches to cloud adoption among agencies, highlighting the need for robust security in both on-prem and cloud environments.
The pitfalls of "lift and shift" migrations and the benefits of utilizing cloud-native features like auto-scaling for enhanced efficiency.
David Joy:
Welcome back to another episode of the Big Ideas in App Architecture podcast. In today's episode, we're going to be speaking to Nikolas Mayer, who is the senior cloud architect at Maximus. It's a company that's known for the work they do with governments and federal agencies. Now, Nik and I talk about his journey as a cloud architect and how he fell in love with the cloud and modern infrastructure, and how he and his team today help federal agencies with modernizing to the cloud and the challenges that are associated with it. So pump up that volume and enjoy listening to this episode, and thank you once again for coming back and listening to the episode.
All right, awesome. Well, Nik, thank you so much for taking the time to come to the Big Ideas in App Architecture podcast. How are you doing today?
Nik Mayer:
I'm doing good. How are you doing, David?
David Joy:
I'm doing all right, man. Every time I see you, as I was saying, this look is like so tech, so this is awesome.
Nik Mayer:
Thank you.
David Joy:
Yeah. So tell us a little bit about yourself and what your role at Maximus is, for the people who are listening.
Nik Mayer:
Sure. So I'm a weird oddball, not so much these days, but back in the '80s, '90s when I grew up, I was that kid in high school in the back corner, I graduated 25 years ago. But back late '90s, I was the kid in the back of the classroom cranking on a computer, the only one, I had... Notebooks had DVD ROMs, that was the big deal, and the battery lasted 35 minutes and you're always tethered to a wall, that's how far I've been in tech. My freshman year in college I was working at Hewlett Packard in the technical computing lab, where we were in charge of virtual multi-tenancy from a on-prem data center perspective. So when I say I've been doing cloud as long as clouds been around, that was private cloud, not public cloud, but that was private cloud, I've been doing that since 2000. So I've been dealing with cloud for a very long time, even before that, I'm one of the few people that have dropped a three-meter rack on its face full of servers, and oops.
But dealing with a lot of this since then. And since doing that my freshman year in college, done a lot of enterprise, generalized architectures, for large companies that everybody knows, you drive one of their cars probably or knows somebody who does, and if you've ordered food from some restaurants over the pandemic, you've run through platforms that are back end microservice based architectures. That system grew from 15 to 45 services over 18 months, with a fully automated prod pipeline, with less than 1000 failures in that two year process that we were running it. We processed 3,000,000 million, less than 1,000 failures due to platform. And this happened right at the beginning of the pandemic, so this wasn't a... That was 2020 we launched and the world went weird, and being able to enable companies to stay afloat and survive the pandemic, was an awesome feeling.
And then coming to Maximus, came over here to help... It was a program for Customs Immigration Service, where we basically came in and created a whole new process for them, for the adjudicators, the hardest process of the immigration process, the humans. So they have to go through adjudicate individuals and manually check on forms, and rather than logging into all the various systems, we created them a single pane of glass, grabbed in all the data, and really allowed everybody to know what's going on and enable them to work much faster and try to help fix that logjam. Because I wanted to stop building... I didn't care about selling widgets, I wanted to do something that actually mattered and helped impact the world, in a good way. Having done this for long enough, you get to the point of like, "I'm done building widgets, I want to do something that's impactful, leave the world a better place." And that's why I came over here to Maximus.
So now I'm a principal cloud architect. I deal primarily in the justice and homeland area, is a lot of the agencies that I work with. There's a few of us that span a lot of it on our federal side. Maximus does more than just federal business, basically... Comments, the statement I'm going to say, basically for almost 50 years, Maximus has partnered with state, federal, and local governments to provide communities with critical health and human service programs. We basically work with governments to implement programs rapidly, scalable operations and automated systems for Medicare and Medicaid, to welfare to work programs, modernization and comprehensive solutions that help governments run effectively and achieve the mission and goals. And some cool things that we do, we are the leading administrator of Medicare enrollment brokers for services in the US, we answer 7,000,000 per month, through our contact center contracts that we do. And we've performed more than 1,600,000 assessments, annually worldwide. So that's really what we do. We do a lot of work for helping the citizens of the government, and helping the government get the missions done, both at a federal, and state and local level.
David Joy:
Wow, that's awesome, dude. I mean, I would love to dig into so much that you mentioned, especially around Maximus, and we are going to obviously talk about that in your work. But as a cloud architect, how did you fall in love with this? What really happened, you know?
Nik Mayer:
How did I fall in love with cloud? It kind of fell into my lap. I mean-
David Joy:
But I... Or let me say this, why did you fall in love with the cloud, is probably the best question?
Nik Mayer:
Okay. Why? Really, when public clouds started coming out and I started understanding that it was what I had done so many years ago, and knowing how... I mean, I was the person that was in charge of rack and stacking. So when somebody wanted a new server, I knew it took one to three weeks, and that was at HP, at headquarters, getting computers, getting hardware, stacking it, maintaining it, managing it, flashing it, the whole thing, and what that management meant. And the fact that I could push a button and it would spin up in a matter of minutes, I went, "This changes everything. This is the game changer." Not just from a technical perspective, but also from a business perspective, because an oddball in the tech world, I actually have an MBA and a marketing degree. Because when I went to school, I was in technical computing lab at HP and I was doing things, I mean, we had motherboards with fans on a testing SP64 for their workstations. We were load testing them, with just massive fans on it to keep it cool, because there was no case, there was no cooling, it was just motherboard, proc, Ram, GPU, fan. And doing it for real and then sitting in class and being told I'm wrong. And I went, "Okay, I can't do this."
So I transitioned more into... I knew I was going to be doing stuff in tech my whole life, I knew that, it's been ingrained in me. But then getting into more of what got me to love it was understanding where it's come from and knowing those small little tricks that come from the old school, of... In a different world, the .com boom, I lived through it. I mean, and so I grew up with a single parent and he had a tech firm, he was an entrepreneur and had a tech company in the '90s. So I grew up... My after school activities were going to the office and hanging out and watching what was happening, and marketing discussions.
And so I always knew I was going to be in tech. How it actually came about was just... I love cloud because it reduced the barrier to entry. CapEx, "I'm going to buy this server that costs $50,000, somehow, because I need it for five years from now. Okay, cool. I can buy that server, or I can buy a developer, I can rent a space. There's all the stuff associated, now I can just push a button and it turns on and I have it, and I just pay for what I use?" It really changes the game, in my opinion, to commoditizing, and helps innovation. And so probably my big passion is innovation, what can change? What can be different? And so in doing that, cloud really allows that innovative process and that innovation cycle to happen, because depending on the cloud provider, as Google says, you can run out of money before they run out of capacity. For, I think it's Spark, their high throughput, they'll give you as much capacity as you can spend, and you'll run out of money before they run out of capacity.
David Joy:
Oh, yeah. I mean, is it Google or Microsoft have their data centers on the ocean as well now?
Nik Mayer:
They're fantastic. I mean, great cooling, water cools 25 times faster than air, so dump it in the ocean where it's really cold.
David Joy:
Exactly.
Nik Mayer:
Kind of smart.
David Joy:
A pretty smart decision.
Nik Mayer:
Except if there's a leak-
David Joy:
Yeah, exactly.
Nik Mayer:
... water and electronics don't mix.
David Joy:
Yeah, that's where you need some multi-region resiliency built into the service.
Nik Mayer:
Yep.
David Joy:
Yeah.
Nik Mayer:
Some good DR practices and employments.
David Joy:
Yeah. It's funny, I bring this up because I used to work for this company called, I don't know if you know, Sungard Availability Services, it was an infrastructure company. And all they did was they would have these racks and mainframe systems, and Z-Series PCBs, I-Series, just sitting in their data center, waiting to be provided to folks who needed them, in case a disaster happened. And that was a pretty lucrative business, to 2007, 2008, 2009. And then I think after the beginning of the cloud, it affected that business because you always had all of this infrastructure available, you could connect to that, you had DR, all these kind of set up. So the landscape kind of changed, but I do agree with you, the cloud essentially came at a time when everybody was innovating and wanted these new paradigms to innovate faster, and go through different iterations, and use what you needed and then come back to it when you need it again, that kind of thing. So I'm glad you fell in love with that. I enjoy the cloud because of those particular reasons as well.
But yeah, let's pivot from that and get into, one of the things that you mentioned is that, when you're an engineer or when you're somebody who's in tech, we can build a lot of things, but what really gratifies sometimes is building something of value that makes a difference. And when you mentioned some of the projects that you worked on, especially through the COVID, allowing end users and customers to experience no issues, as well as what you're doing right now at Maximus. Tell me a little bit about how the work that you do at Maximus, especially with the federal use cases, what are the challenges that you come across? And how is it different from what a typical, say a company like Netflix, is going to deal with?
Nik Mayer:
Sure. I mean, it's an easy answer, but it's also very hard. And the best... You probably hear this a lot because you talk to architects, but it depends. That is the ultimate architect answer, it depends, because everything is different. But I'll just take private versus public sector, so private being any company that isn't part of either federal, state or local government. For the private sector, it's really a choice of what's your risk from a company perspective, if you're in a regulated industry, so we'll take Netflix, because it's not regulated super easy, you use them, and their real choice... The only thing they got to worry about is shareholders. And in most of the private sector, you're beholden to your shareholders if you're public, if not, you're beholden to your CEO or whoever owns the company, and so your risk tolerance can be much higher than for federal government. So if your PCI-compliant, something happens, there's a data breach, some credit cards get lost, yes, that's a big deal, but in reality, it's not. You notify customers that it happened, credit card companies have their checks and balances in place to try to mitigate that risk.
But now let's switch over to the public sector and let's say that it's the IRS. A data breach happens to the IRS, that's a big deal, that is, I call those sacred data sets. There's like PII, special PII, HIPAA, that level of data at Citizenship and Immigration Services, USIS, or let's say IRS, SEC, these agencies, if data breaches happen there... And let's not forget about DOD, now we're talking about things that are threats to national security, that if a data breach happens, that's a problem. So the biggest difference is the risk tolerance in public sector is so low because of the data set that's being used, as compared to private sector, which has a little bit more risk tolerance. So that also drives innovation, and innovation in government and federal work is a lot different because of those requirements of a lot of those pieces.
So I've been talking to a lot of agencies, and lot... GenAI is here, it's here to stay, it's not going away, LMS, GenAI, it's being incorporated to the foundational level, down to cell phones. And when that happens, it's here to stay. Which is awesome this time, because I think it's actually going to... It matters, it helps. I watch my kids be able to interact with the computer without being able to type because they can talk to it. And the startup companies that have... I can have third-graders writing code by talking to a computer, that straight-up sci-fi from 30 years ago. The matrix was a mind-blowing thing in the '90s, we're kind of living there now. Not the fact that we're jacking into our brain stems and we're batteries, but talking to computers and having them talk back in a way that isn't robotic.
I mean, the chatbots that are out there on call centers, sometimes you don't know they're a bot, but they are. I mean, the voice... So all that stuff is here to stay. But where I'm getting to in all that and bringing it back down is, the risk tolerance is low. So GenAI is here, and I've talked a lot to public sector agencies recently, about how can we use it, and we work a lot with the big three hyperscale cloud providers. And then so how do we work with them to ensure that data doesn't flow out of the model? How do we ensure that AI model ops, is a way that works that can keep that data contained, so there isn't data leakage out? I mean, would you want your data to be stuck in ChatGPT right now, that somebody could prompt engineer themselves into finding out SPII?
So that's where a lot of the challenges come from, is a lot of sitting in that space where, you want to do innovation, you want to bring innovation, but you have to move at a very cautious state, depending upon the agency. And so there's agency risk tolerance, there's then... So I deal a lot with Homeland Security, that's 22 agencies compressed into 14 different areas, that each have different priorities and requirements. So though it's just one agency I'm dealing a lot with, it's actually 14 agencies with 22 components, that have overall DHS, and then each component and the national labs. And I mean, it's wild the amount of requirements that change, and then an executive order could come out and change everything.
David Joy:
Right. Yeah. I mean that's important too. So how has it changed though? Traditionally, these institutions that are governments, have preferred to stay on-prem or have their own infrastructure, but I think that's not scalable with... I mean, most of these systems are also not GenAI ready, so I mean path to the cloud, also provisions or modernizes them to start using some of these capabilities. So tell me a little bit about how the landscape has been and how has it changed? In a sense, is it more cloud-friendly now? Or there's a... I know AWS has GovCloud and some of these requirements are coming up, so...
Nik Mayer:
Yeah, it depends. But it really... Yes, there are some agencies that are very modern, and some agencies that have a very high appetite for wanting to go to cloud. But it's that balance, so yes, everything was on-prem, really easy, you could secure your data center, secure in and outs, you could touch the stuff, you had air gaps, and if the... I mean, I forget which movie it is now, Transformers was a really good example, cut the hard line. I mean, they have air gaps in place, so if you pull the power, you have little fiber channels, that is a literal air gap that if something's happening, you air gap it, it's a secured system, you're now isolated, and have a literal break between you and the internet. Running on generators, and you are completely broken from the outside world and you can protect your data.
In cloud, you can't do that. So you have to look at it more, because it's software based, you have to figure out how to build those pieces in, and yes, there's GovCloud. The challenge with GovCloud as well is, it's not in parity with public cloud. So we'll just talk AWS for a moment, GovCloud, is always sitting, behind because they have to lock a version, have it go through all the security checks and check all the data, about where the data flows, and how it flows and where it go, and go through that whole security process, then it's certified and then it's able to be used. So when Lambda first came out... Lambda is awesome, I love Lambda for... That big order process was all Lambda based, 45 services. Fun fact-
David Joy:
Part of your microservice is running on Lambda, right?
Nik Mayer:
And those weren't just Lambda... They were all Lambda process. So for us, the reason why we were so stable, when we do our CCICD process, we would build the entire infrastructure, test against, tear down, because we could, we were running two tests. So it cost us very little money to do a full end-to-end test, but we could stand up the entire infrastructure, just to do testing in a CICD pipeline. That's some of the... And that piece, also in my opinion, brings innovation, because now I can build new stuff, test it against not a prod-like environment, but literally the prod environment, because whether Lambda runs it here or here, it's the same, as QS is the same. All these AWS managed services, whether you use A or B, it's the AWS managed service, and so you can take that SLA and that guaranteed ability and usability, and just use it. So you want to test your system, automated testing? Yeah, why not? I don't have to... And so I'm not double buying my infrastructure, with Lambda, you don't need to have really... Besides having it deployed in another region for DR, it's an active-active environment at all times. If you have Dynamo Global Tables and Lambda, you're literally an active-active and can be anywhere on the planet.
David Joy:
You mean for a NoSQL workload, essentially?
Nik Mayer:
Yep. Yeah, yeah. Yeah, If you're going to run Dynamo with NoSQL, because Dynamo has the global tables now, and that parity within seconds globally. But I mean, so Azure has their Cosmos DB, it's Cosmos no, it used to be Document DB, now it's Cosmos. That, same type of thing, global parity within seconds, and using that for massive throughput workloads. We did something that I was pushing 1,250,000 documents per minute, 24/7, 365, through that, for a project that needed to have a data feed for real time transit data. And we just used that because of TTL, and all the pieces, and all the search qualities and everything that we needed, but we had a collection of 10,000,000 documents, and a production workload that was getting hammered, and it kept up and worked.
David Joy:
I mean, that's the advantage, right? Yeah. So I think we pivoted to this, but you were talking about how things have transitioned now, and GovCloud, and how there's preference to the cloud between these agencies.
Nik Mayer:
Yeah, yeah. Sorry. So the other thing that GovCloud gives you is the ability to keep those security pieces in place too, because it's certified at a certain... So there's something called FedRAMP and ITIL, those are the two standards, and so anybody out here listening or watching, if you want to work in public sector, make sure... It is a process, don't get me wrong. I mean, it takes time just to get your environment certified to be able to put a workload in, and so that's one piece. But that's something that we did as a company, to help us, we actually host workloads for public sector. Because we have that FedRAMP certified space, and then all you have to do is get that workload certified, rather than the environment and the workload. So that's another piece that we help a lot of public sector agencies, federal, state, local, to be able to rapidly put things up, and have something in that kind of protected space, that they don't have to worry about multi-year process to get it up.
David Joy:
Right. Yeah. So this begs a question now, as a cloud architect, when this is where the space is going, how has the role of the cloud architect changed? Because you have to take the role of a consultant, obviously, to also educate the people you're working with, to tell them, okay, how these things still remain, but from a software point of view. So do you feel like you have to wear those hats as well, while you're also architecting a solution for them?
Nik Mayer:
Yeah. This is actually my second time being in government. So I went from school, I started working in public sector with the Department of Homeland Security: Science and Technology Directorate, which was this really cool kind of conglomerate of national labs and agencies to look at it. I got to talk to the person that put the 3-1-1 rule in place, the actual director of that division from S&T, and going, "Why do I have to put all these silly things in a bag? Why does it matter?" And then he showed me the video of why it mattered, and I went, "Okay." Because you cannot get the right collection of the right chemicals, in that bag, in those different sizes, to make something that can explode, that was the reason why. We didn't have to take our shoes off until somebody stuck a whole bunch of explosives into a flip-flop, then we had to take our shoes off.
But figuring these things... So that was that group. I worked with them at first, and when I first worked with them, you basically had to be a really smart full stack engineer, that knew all the components and knew how to build it. Because solution architects were essentially really smart engineers, and if they were in that done space of, "Well, I've been an engineer. I've been a senior engineer for way too long, I guess now I'm an architect." That was the initial thought back in 2008 to 2015. 10 years later, the architect is now actually an architect, we are the people that have to come in and understand holistically what this means, and what's this new piece of tech and how is it going to implement. So it's a hybrid role now between business consultant and tech. My other piece that I pull in, is security, because that's just the next evolution beyond cloud.
Now, security is actually top of mind on everything, so much so that the administration put out the executive order, about everything has to be zero trust. Not knowing what that actually means, because, "Okay, let's take some COBOL systems and make them zero trust." But how? Yeah. But the idea is correct. How you implement it, then becomes the challenge, and then you have to make abstractions and all that fun stuff that we all have to do. But that's really some of the pieces of... The architect now has to understand, beyond just how to build it, but what's the cost implications? Because your choice to use tech A versus tech B... I mean, so working with new Relic for the New Relic One platform, we were one of the first test users for serverless on that, not knowing how chatty it got, and the cloud watch logs bill, was four times our compute bill, because of how much it wanted to get to telemetry data.
Because it had never done serverless before, it had just been on... If you have an agent running on a server, it's always pinging out, you don't think about that. But we turned it on and in a week we burnt a month worth of our normal cost on CloudWatch, there was 30,000,000 hits over the month, which is just astronomical for a non-prod system. We were just doing test in prod, it wasn't fully out to prod yet, and we were cranking on it like you would... Even when we were running full production, we never hit that many a month. But when you're running a couple of thousand Lambdas and they're all pinging once every minute, that starts collecting up. I mean, that's 60,000 hits an hour.
David Joy:
Yeah, that's why. Yeah.
Nik Mayer:
So that's why the architect has to start thinking about cost. So it's a cost benefit analysis that we have to do a lot. So my approach that I love to use is I love AWS managed services, they cost more money, yes, but if I need to get there tomorrow, I know it's going to work, I have a SLA, I can build things against it, I can get it to work, build a POC and be done. Instead of having to go through all the individual steps along the way and build out everything, and build it full productionized. But if I need a quick POC, use the managed services because there is that inflection point, where the cost to build it and manage it, you're using it so much you need to do that. But if you build a POC and realize it's happening and it's happy, that's your zero day, not [inaudible 00:25:40] really, that's zero day, from how long you have to get it done. And hopefully you can build it in enough time that you hit that break before you hit the breakeven point.
So that's why I look at it as, yes, you're paying more money upfront to get you where... Maybe it's a new feature for business needs, maybe it's that new feature that all of a sudden is a revenue generator, that makes your company go from eh, to now a hit, maybe it's that new feature in the app, whatever it is, but you can get there tomorrow. I mean literally you're... Everything also with the cloud, change developer roles from just, "I'm going to write code." To, "Now I have to think about how this operates in a cloud environment." Junior devs now are beyond what would be considered architects 10 years ago, because an architect was a good engineer who kind of understood some cloud stuff because it was just in the infancy.
But now your junior devs that have four years or less experience, because they've also grown up in cloud and understand it, but now they're thinking of, where's my workload and how does it work within that cloud environment? And DevOps has been great and developer enablement, and that, now as a developer, I need this resource. Cool. Spin it up, use it. Go for it. As an admin on the DevOps side, on the infrastructure side, I don't want to have a ticket to build a server, so you can test it. Just test it, it's going to get torn down at the end of the day, I'm not going to let it burn their zero.
David Joy:
Yeah. I think it's also good for developers and just generally for folks who are doing these evals and POCs, to understand what the projection looks like. Like, "Okay, I've added this feature, it's a bunch of Lambda functions and it works really well, but it's going to cost us this much." It helps you project things out, and it's good for business, because back in the day, you did not know how to plan for scale and plan for peak hours because a lot of workloads that we know have hockey sticks, where they're flat and then suddenly there's like, Elon posted something on Twitter, or somebody, it just went cuckoo, right? So I think it gets... All these DevOps and all these key technical paradigms that we are used to now, I think have made things so much more better for folks like us who are trying to develop things, it really does.
Nik Mayer:
Yeah. And a weird thing that nobody thinks about, I love Lambda for very specific applications. If you need heavy log levels... Because once Lambda has done is, I consider it as ephemeral hyperscale containers. If you just want to think of it as a simplistic... Think of it as Kubernetes, it's throwing containers away every time you're done running it. So you have to store the data, so usually that's going to go in your queue. So though it might be a lower compute cost, you have to think about the total cost of it, because now you have queue costs, you have logging costs, you have compute costs. So yeah, your compute's really low, but my bill doesn't change, why? Because the cost has shifted. But if you need really heavily logged things for a regulated industry, or a regulated agency, maybe it's smarter to stick with more of a containerized application, that runs on EC2, or EKS, or whatever, Kubernetes for the cloud platform, because you need that extra logging ability and you have to run those agents.
So it gets back to the other consideration of, I don't always get to use the coolest tech because requirements require me to log everything, of every time of every bit moves through the system. Okay, so that's not going to be done on Lambda because, no. So yeah, we're going to run it containerized on EC2, or ECS, I mean, I prefer ECS, because I like the management because they'll push the logs, I don't have to think, "Did the log streamer work?" But whatever. So it runs an ECS instead of EKS, because we don't want to manage a cluster, or maybe that agency has a large EKS cluster that they manage and all we can then throw workloads on it. So now we don't actually manage the EKS cluster, but we have to manage the workloads on EKS cluster, that's managed by a different team that we don't touch and can't talk to. So when I say it gets really weird and messy, it does, because the requirements for every workload that I work on and every contract is different.
David Joy:
Oh, yeah. I'm pretty sure. I mean, they will have these buckets, these people can only touch these environments. Folks can't do it, they just can push, but they can't pull, or they don't have visibility and things like that, all these guardrails have to be posted, so it completely makes sense. One of the questions that I wanted to get into, from previously, what you were mentioning, was COBOL and the idea that you have these applications that are extremely complex, they were written 10, 20 years ago by somebody, those people are retired. And so tell me... Or 30 years ago. So tell me this, what is easier to migrate nowadays for you to think about, is it migrating the data or is migrating the application?
Nik Mayer:
Well...
David Joy:
Or what have you sensed in the ecosystem?
Nik Mayer:
Yeah, how to answer this right? So I mean, we all know that in enterprise, there's always that computer under the desk that runs something that is a critical piece of something. All of us in this industry know that if we've run infrastructure, at one point, it's like, "Oh, it went down. Why? "Oh, we unplugged that computer." "Okay, cool. How do we... Okay." And that person's been gone for 10 years, and, "Uh-oh, now what do we do?" And I liked how the hyperscale for writers, so I'll just pick on AWS again, when they brought out Lambda, they brought it out for Java and Go, and PHP, I believe, I forget the first three. But it's three basic languages are very wide. I'm like, "Okay, cool. Well, what about those of us that need the other weird stuff?" And then they went, "Oh yeah, okay. Well fine, here's one that you build your own runtime. It just cost you more money." "Okay." But that was really an interesting way, and that was just leveraging the power of containers, in a very interesting way that nobody had thought of. Because it's like, "Oh, it's just a container. Spin it up, run what you need. Fine. Cool. Now you're going to want a really lightweight version of whatever."
So getting into that realm of... I think the easiest thing is always to migrate data, just because data is data, and now if you're going to transform it from SQL to NoSQL, sure there could be some challenges, but if you're migrating from MySQL to Postgre, there are now lots of industry standard tools that are from hyperscale providers and very specific companies that could do that. Migration of data works. I mean, you could even write languages to do it. And I can go on ChatGPT or OpenAI, whichever I want go with, Bedrock, Copilot, and say, "Write me a Python app, to transform data from MySQL to Postgres, that's going to go from on prem to cloud service." And it will give me all the stuff, I fill the three things, I could be done with that in a day.
Now I need to go validate that data, so there's more time spent on the validation side, but I can automate the validation, that's fine. As long as I'm fine to hammer that DB as hard as I need to validate. I mean, I can validate every cell within that, every field, every cell, it doesn't matter, all that matters is time and how long you can be without that data. But I mean, sure, I'll take a snapshot of the database at this time, migrate it, build up my own infrastructure outside of scope of production business use, and hammer the heck out of it, whatever. Fine. I'll tell you, and I can guarantee the data is the same, it's a one-to-one parity at that point. So that's why for me, data is easier. Workloads, because when you transform languages, that just gets weird. I mean, because a function that works in Java changes in Go because it's an abstraction, and then you're going to Python it in between. That brings up a point where, you get me thinking about lift and shift versus... A lift and shift migration.
David Joy:
That's where we are going, yeah.
Nik Mayer:
Okay. Okay, cool. But a lift and shift migration, that's great, but you're building a tech debt. I actually have a quote that I want to pull up because this was something that I found. McKenzie did a study on this, and they found that tech debt can cost 20% to 40% of your total value... It says, the direct quote, "Tech debt can account for 20 to 40% of the value of a company's entire technology estate."
David Joy:
That's a lot.
Nik Mayer:
That's massive. So if we're talking $1,000,000 tech estate, that's $250,000 to almost $500,000, that if you just lift and shift, it's going to cost you that later. Not to mention, then they start talking to operational efficiencies, because if it's old tech running on new stuff, can you really take advantage of DR if you don't actually have it put in place? So they estimated that 10% to 20% of technology budget, for new products, is often redirected to address tech debt. So now we're talking about a 20% to 40% cost, another 10% to 20% cost, and then we're really starting, this is the big one, this is the opportunity cost that you can't really put a number on, well, they kind of did, but... Your strategic flexibility, if you have an old system, that hasn't been modernized to at least base level of what cloud can offer, you can't take advantage of auto-scale, like simplistic things of auto-scaling. So when you hit that hockey stick, you're done. You either have to have a massive bare metal server burning at $4,000, $6,000 an hour, just for that what if, as compared to spending your couple hundred dollars on your normal workload and auto-scale it up as you need.
David Joy:
We'll get right back to today's episode in just a moment. But before we do, I want to let you in on a secret, big ideas like those you hear on this podcast every week, don't need big databases to start. With CockroachDB Cloud, you can bring your apps to life quickly and without upfront cost. Architected on the same resilient, distributed SQL platform that industry leaders trust with their mission-critical workloads. So sign up for free and start bringing your big ideas to life today, at cockroachlabs.com/bigideas.
Yep. No, I agree. Yeah, and that those are the patterns that I myself get into conversations with, and I think that's where the key goal to... Should we lift and shift or should we lift, tinker, shift? You and I were talking about this when we first met, and I have this theory, that not many people want to do it, but I think that's probably the right thing is, migrate the data, but build a completely new application that can understand this data. But figure out a way to... And the problem with the mainframe systems or these old system, is that knowledge management was not done properly, so many times you lost what's connected to what. And so it becomes very difficult for people to build something completely new, and from an application level with all the same functionalities, and I think that's where most of the work goes in. So yeah, that was my experience, what about you?
Nik Mayer:
Yeah, 100%. The difference for my thinking is, I take that as a chance to look at it and say, "What isn't needed?" Because when you migrate, if you're going to lift, tinker and shift, tinker with it, cut out unneeded... If you have functionality that's been there just because it's been there, why? Can we pull it out? Because when you lift and shift it, A, you're not on that production workload yet, so this is one of the only opportunities you're going to get to be able to have it break and not impact business. So push it up there, take some things out. Because if you're migrating, you have the source AMI files, if you're talking AWS, or the source images, push it up, take some stuff out, see what happens. Give yourself that... If we're talking a 30% to 50% cost, take that 10% and play.
Let your engineers play. Let your engineers go in and rip stuff out. See if it works. I mean, you can push it to a new database, push it into a new VPC that has nothing to do with what you're doing, but give yourself a tiger team. Take your top engineers that can get things done quickly, and let them go rip at it and find it... Every time I work on a project, I say, "This is the only time I get to have my first eyes on it." So I dig in deep. I look at security architecture, system architecture, if I have the time even application logic architecture, data architecture, where are you doing things? Why are you calling a database 16 times, instead of calling it once and caching it? Could we optimize... Where's the optimizations? And that tinker also comes to optimizations, so you could potentially find savings.
There was one project, I walked on, that I looked at one thing and went, "This is a 64 node cluster that's running at 10% average workload. Why?" "Well, because that's what we have on prem." "Okay, so let's see what the actual usage is. Let's cut it back to 32, and see what happens." And during our migration, I pulled it back to 10 on the cluster, because it was for an elastic cluster. And I pulled it back to 10 and just let it go, and people were freaking out, but it had nothing to do with... It didn't impact anything production, it was all in the background. But let's see what happens when we start hammering it and seeing what happens, in our automated test suites. And yeah, it spiked to 95% a few times and it pinged out at 100% for a little bit, nothing timed out. But because of the way that... It's like, "Okay, fine, so we don't need 10, maybe we put 12. Or we sit at 10, and when we notice that we hit 80% load, we know we have the five minutes, spin up five more servers. Fine. Spin it up, shard it out, we're good. Let's keep going." Or spin up a read replica just to serve the recalls off of, where you have your write... Just associate write and read replicas, so you write to...
David Joy:
Yeah, no, I like that pattern too. I feel like it's a good pattern for high throughput applications, where you can do your wrights and reads on your main infrastructure and then move non-operational queries to another cluster, and you just read out of there as much as you want and don't bother the application.
Nik Mayer:
Yeah. I mean, that's what the AWS got so right with Aurora. Aurora, it's expensive, but man, it is... That has to be one of the coolest hyperscale DBs available. Because I mean, automatic primary failover, as an infrastructure guy, we get calls at two o'clock in the morning because the database went down, because primary went down, reads were fine, you could read them all day long, you just couldn't write, that just handles it and it's done. And it boots that primary off, the next one comes up as that primary, and they add another read at the bottom and you got six versions that are there. So be it, it just makes life easy.
David Joy:
Yeah, no, I have a lot of respect for... I mean, I worked with like AWS, this is one of my partners I work with, but how much do you know about CockroachDB? Because we do... Yeah, so the biggest difference, I mean, the reason why I want to bring it up, is because you said... When it comes to CockroachDB, we don't even have a primary, secondary cluster, we just have one always on cluster you can read and write from. So most of these scenarios that Aurora even gets into, Cockroach never has that. And those are critical for certain workloads and use cases. That's [inaudible 00:41:17].
Nik Mayer:
Yeah, yeah. Oh, and that was one of the things I loved when I first heard about... I'm like, "Wait, you don't have..." When it was just read, write everywhere and everything's there, and it's on the surface, and I hope I'm not saying the wrong thing, but it's very blockchaining in its concept, in that the ledgers everywhere, the data's everywhere, and read, write, it happens.
David Joy:
Right, yeah. And I think those are some of the patterns. Where I think in Aurora's case, I think I've enjoyed using Aurora, but there were certain workloads where we were like, "We cannot have writes going down." And this, we are talking about highly, highly, highly mission-critical workloads. So those kind of cases make sense. But dabbling into the data stuff, just because you mentioned Aurora, I mean, I think it's really cool tech, when you think about data migration for these federal use cases, what mindset do you adopt? So what are some of the best practices that you go through, that you can share with the listeners and help them understand, "Okay, this is how I should be thinking about this problem."?
Nik Mayer:
My first piece to say, is to remove the, it's always been so I can't. In highly regulated industries, when you walk in, "It's always been done this way for 25 years, so we can't do it." "Okay, well if you're wanting us to modernize, then let's talk about what we can do, not why we can't." So that's that first mindset, for me as an architect to walk in and try to make it work better, is just to remove that, I can't. Because When everybody sits in the, I can't, you can't build, you can't invent, you can't talk about what's next, you're just going to talk about what is. And what is, is what we're trying to fix and trying to make better. So I love hearing about it, but I'm always the person that goes, "Why? Why is it always done that way?" "Oh, because it's always been." "Oh, so there's no real need? Cool, let's change it."
So I like to challenge the normality of, it's always been done that way, "So why don't we change it?" "Because we can't." "Why?" Now if there's a law saying we can't do it that way, sure, okay, fine, then we have to do it that way. But in the most part, a lot of the times I'll walk into it, and, "Why?" "It's just always been done that way." "So why don't we change it?" Because that runs into people risk appetite, because if people's jobs are on the... Another fun thing that I get to deal with a lot of public sector, is if there's a big wrong choice, it could be considered misappropriation of funds, which means people can go to jail. So a lot of people don't have a large risk appetite because it's the potential freedom of choice. So these are all considerations that you have to take into place of, what's the business cost? What's the business need? What's the risk appetite? What's the personal risk appetite for that project or program manager, versus the company versus the agency?
So it really is this... It's almost like playing three-dimensional chess for a... Or any sort of... Thinking of it as quantum physics or astrophysics of, how do you shoot a rocket out to meet a planet in orbit, as we're both moving in different speeds, it's that level of calculations in real time of, "Okay, cool." So one of my best practice, I poke at the edges, because if nobody's poking at the edges... My job, especially as a consultant, is to challenge and question. Ultimately, it's my client's projects, my client's money, I will do what they want, but it's my job to give them options and to ask the questions that nobody else will. Because imagine somebody who's been a career federal employee, they've been working there for 22, 23 years, they're a director of a program, they got two years till retirement on a full pension, do you think they're going to try to do something that might cost them their job?
David Joy:
Maybe, no. So, yeah.
Nik Mayer:
Yeah, I'm not saying it's wrong, I mean, that person is considered an expert in the field of what they've done for the last 25 years. So for me to come in and ask and go, "Well, what about these couple things?" Also, if it's out of scope... The other big challenge is the scope of the contract that you get assigned or awarded, can you change that? Not usually, not easily. So what we're talking about now is, for the next five-year contract, because essentially you have... And so bringing new innovative technologies is challenging, because it can take two to three years to get that through the approval process.
David Joy:
And by the time you get these approvals done, the technology has gone behind because there's something new.
Nik Mayer:
Yep, yep, yep. That's one of the reasons I like working with DHS, they're the newest agency, so they built some of those things into it. DOD's notorious for buying five-year-old laptops because that was what was on contract. But your contract is, you will provide X notebook with X specs, and that's how IBM, Toughbooks are out there for forever, because IBM would build that version of laptop for 20 years. Or Oracle has a database that they'll support for 20 years. It's that kind of challenge with a industry that moves at the pace that we do where, I mean, Moore's Law is stating, every 18 months technology is doubling. We're hitting that point where it's slowing down a little bit, but it's still... Okay, cool. We're talking a three-year process between when they get an idea and put it out for procurement, that's still changing at least once, maybe twice, as much as three times, between when you talk about it.
So that's another fun piece and challenge with that, is how do you deal with the innovative challenges of the procurement process, and getting in the way of innovation? But it should as well, because... So I don't know if you know about the federal procurement process, so first you get an idea, you then send out an RFI as an agency, the RFA will take three to six months for industry to go out and respond back to you, and then you have a month or two to read it. So we're talking three-quarters of a year just to get RFI responses, digest it, figure out what that means. So now a year later you're going to go, potentially, put out an RFP that's going to take six to nine months to award, so we're talking a two-year process. Usually it's like a one to three year life cycle, from a new idea to being able to buy it.
David Joy:
But why is it so slow?
Nik Mayer:
Because the federal government is the largest enterprise on the planet. I mean, if you get down to it, it is the largest enterprise business on the planet. Because there are laws involved and a lot of lawyers, and that's not even if you're going to talk about a protest, if somebody has a contract and they lose and they protest it, that could stall for another one to two years beyond that.
David Joy:
Making the job of the cloud architect way more difficult.
Nik Mayer:
Yeah. Well because now, I got two years of a company not trying to update and modernize. They're just trying to maintain, because that contract... They're going to follow the literal contract, for the time that that protest happens, and then we got to take it over. So what we're expecting and understanding, what was there is no longer there, what's changed in the last months or years, so then it's coming in and reassessing everything again because it's been two years later. And these are potential critical workloads that can't go down, because they're mission-critical to Homeland security or DOD.
David Joy:
Right, yeah. How is the performance expectations though? Because back in the day... And I grew up in India, and this is this funny story, you would try to book a train ticket to travel from City A to City B, and the train system was basically owned by the government, and right when you want to book it, the system would crash, and then you basically can't book the ticket for forever. So what people would do is, go and actually try to book the ticket from the counter, and they would end up there at 4:00 AM in the morning, for a counter that opens at 8:00 AM. I've done that actually, my dad and I went and did that, waited four hours, because my dad was like, "We won't get the ticket otherwise." I'm like, "Okay, cool. But government systems are always notorious, especially growing up in India, and over here is like, but not being the most innovative, most robust systems because there's so much bureaucracy not just that, I mean laws and things and people's focus is not on innovation, but to just provide something, maybe oversimplifying that. But how is it now? At least with federal, how is resilience performance important now?
Nik Mayer:
Yeah, yeah. I will say that over the past 10 years, things have gotten better. As agencies and the federal government, I'll just talk from a federal perspective, they have gotten more of an appetite for cloud, to give them the elasticity and availability that they need, to handle increased workloads. Like simple case in point, healthcare.gov. That website gets almost no traffic for nine months of the year, and for three months of the year, it gets hammered. Traditional purchasing for a contract like that, would purchase max capacity for the five years of the contract. So you're spending money on infrastructure sitting there on idle, but you have to buy it because you can't just click a switch and charge more, because you have to know what you're going to spend five... Now, this is also some of the challenge, I'll put it, you're buying five years ahead of your capacity.
David Joy:
A tricky problem.
Nik Mayer:
How often can you be right predicting a five-year usage trend for something, without having the information on it? It's like if you're doing work on your house and you have a contractor, and you don't like their work, but you got to find a new contractor, you say, "I want you to finish the work, but you can't look at my house. You can't take a look at anything, you can't talk to anybody and give me bit." So everybody's guessing. So that gets to be another challenge, in that it's better now. So now with cloud spend, it's more of a, "Cool, we think we're going to use about this much, and you got to keep it around within it." But it isn't like, "Okay, cool, I need X server with X specs, for five years." And you're buying for a five-year unknown capacity, and, "Maybe we'll get there."
So that kind of level of change over the last 10 years, with more cloud adoption and... The ability is actually saving... I love it, because me as a taxpayer means I'm paying less money for more... I'm still paying the same amount of money, sure, but I'm getting more functionality out of it. So healthcare.gov doesn't... When it launched, sure it went down, but who could have guessed? I would have... I mean, I would have over provisioned like 3X what I thought I could have, just to make sure it didn't go down. But was that in the contract? I don't know. Could have been that contracts that limited to, you have this much and make it work. So it had to go down to go, "See, we told you so. We said it was going to be more, it didn't work." I don't know, I wasn't a part of that contract or any of those things. But just from a user experience, I look at that and go, "That was likely a contracting challenge, and that was a very early cloud adoption piece."
David Joy:
The problem is you don't really know, because I think these kind of initiatives in the federal space, is associated with some law or some innovation that the government introduces, and you don't really know how people are going to really take it.
Nik Mayer:
But throw it in an auto-scaling group. Do I have to over provision? No, but throw it in, and if it gets hammered at 100% for five minutes, spin up 2X the servers, just for at least that first few days. But I mean, that's the great thing with cloud, I can automate my pipeline to say, push to this many servers and put it to all ludicrous number that you don't think you'll ever hit. If the mandate is, have the system up and available, if you're running 10 servers, turn your max auto-scale to 40. And as an infrastructure person, if I was managing that infrastructure, I'd be watching it when it turned on for that whole day, and I just have that screen up watching the CloudWatch dashboard and watching my metrics, and seeing what's happening.
David Joy:
Agreed. I mean, same thing happens, I mean, I'm just... Other examples are like IRS, right? January to... Not even January, I think it blows up in April, I think that's what happens, right? And then right after that, it has to process things, so maybe it's a little bit more up for a couple of months and then it slows down.
Nik Mayer:
Their heavy times of the year is January through May, and then they get another big hit in October, because October is the deadline for extensions. But that's their big thing, so they're building all year, testing... And that's a system and a collection of an astronomical number of smaller systems that talk to each other, that do very specific things within it, that I wish I could go into more detail. But it is an interesting system in the way that it's been built over the years. And-
David Joy:
We'll probably need two, three hours of us just-
Nik Mayer:
At least, and some extra clearances and you know?
David Joy:
Exactly. Yeah. We are even just scratching the surface and not even going in depth. Very cool. I mean, I didn't even realize you're at 50 minutes now. But what's cool about this is, you talk from... What I loved about the way you connect with your experience, Nik, is, it's so real for you, you've worked on these problems and these problems exist in these ecosystems and agencies, and it's a different word from what private enterprise is of trying to solve. This is a very funny example, I don't know about you, in 2017, Venmo was pretty new, and Super Bowl happened and crashed their database. They just couldn't handle the right... Because everybody started sending money, add pizza, add chicken wings and all of this. And then they had to literally get out of their database, I think at the time it was MongoDB, and then I think they kept some of the use case on Mongo, and then moved to Cassandra. But they can make those changes quickly because the way they have to prove to the shareholders that the money that they've put in is going to get returned and stuff like that. But the federal agents don't operate like that.
Nik Mayer:
No, it isn't the contract. And then if that case happened for our program, the federal government come back and go, "You didn't perform. You didn't pass the performance to meet the SLA, we want our money back." So again, it gets that risk appetite of, do you want to push the boundaries too far? Sometimes it's needed. If it's broken, fix it, and then it depends on the agency and their appetite for the openness. And it depends how long you've been there, so there are some workloads and some agencies we're working with, where they implicitly trust us, because of the people that have been there for 20, 30 years. And so getting to be that trusted advisor, it really gets to help because it takes twice, it really... I mean, and this is something to the audience, that really, if you want to become a trusted advisor, make the right choice for the customer twice and they will trust you.
David Joy:
I agree. I agree. I just realized that we are hitting almost an hour. So first of all, I would like to ask you, as we close I'll come to the end is, how do you stay up to date with what's up in the market, with technology? And how do you keep up with it really, in [inaudible 00:57:12]?
Nik Mayer:
So I'm weird in that I do... Because if you're a pure technologist and you're focused on writing Java code, that's an easy answer. So it kind of gets into are you going to specialize or are you going to generalize? Because I've been working in tech so long and have gone across the different specializations, that I've become more generalized just because I transitioned through the different areas. From building websites in colleges as a side gig, and full stack development from that, and then getting into application architecture, data architecture, now system and security architecture. So for me, it's been a generalized piece, but if you choose the generalized route, you can't have FOMO, because there's too much happening, there really is.
So part of the way that I like to stay up on things is by giving myself time in a week to do it. So I like going to smaller conferences, that aren't... AWS re:Invent, it's great, I'm not knocking it, I've been there a few times, I don't want to go again, it's too many people. You go to AWS re:Invent to meet with clients or to learn about what's the new technology, so if I need to go learn, I'll go spend two days and go to the exhibition hall. The last time went was in '19, so there was the ARIA exhibition and the Venetian, I went to the ARIA first and walked that for a whole day, just to learn what was there and then went to the Venetian, because the Venetian one is the massive huge expo center. But if you need to go learn info, go walk the booth and go walk the outer edge, the big boost in the middle, the Splunks, the New relics, it's all the stuff on the website.
But I want to talk to those small innovative companies, especially when I'm talking private sector, because then I can talk to them and go, "Oh, you got a cool little tech that's going to work with this thing." And always remember that, keep a log. So I'll keep a log of what I found at the show and the cool stuff, and just keep that in a digital file. And then when I go... Then somebody asks me a weird question, I just go in there and search for it and I'll find it. So I have an old Excel file that... People have them for passwords, I have them for cool tech I found over the years. So that's what I do is that old school rudimentary database, go in and search it and find what's there. And I'm like, "Well, is that company still there? Or has it been acquired or what's happened?
That's how I do it is, if I have to go to a convention to support a client, or a customer or my company, I try to block out at least a half a day, maybe the last day, and just cruise the floor to see what's new. I also set up email lists that will be that data flow in of what's happening, what's new, to get those. And I also look at things outside of tech. So I don't just go overloaded with just techno babble, I try to give myself 10% time to learn about other stuff, artistic things. I started down a rabbit hole of learning what it took to do glassblowing, because I saw a show on Netflix that was a competition about glassblowing and I'm like, "That's cool. It's very scientific, it's very rudimentary, they don't have high-tech tools, but they make these amazing things that I know I can't do."
So I try to always pick up at least four skills a year to... Am I going to go do a glassblowing thing? No. But at least I understand the concepts and the tech, and what's used to give my brain an extra space of learning new stuff. Because if I learn new stuff outside of tech, that gives the tech a chance to go and take that normal stuff and maybe take that widget, instead of it looking at this way, you go, "Okay, now it's like a Rubik's cube." And you sit on the corner and now it's that perfect thing for your workload that you need. So that's what I use to keep... I more try to keep my innovative thinking of, how do I do things differently, than what's out there.
David Joy:
Yeah, no, I love that. I think sometimes we can be too tunnel vision, sometimes in our world, and having different experiences, reading books on different philosophies or watching something different, allows you to look at something that can be learned from somewhere else and applied in tech. Or what do you learn sometime in tech, you apply in your personal life, I get my wife to follow process at home sometimes because I'm like, "Why can't you just follow the process?" But anyways, so it's-
Nik Mayer:
But I got Scrum certified and it made my workflows and homework a lot easier, because I'm like, "Let's run this a Scrum team. All right, we're going to run on sprints, we're going to do with this." And my wife was laughing at me the whole time, she's like-
David Joy:
The only yellow belt is that.
Nik Mayer:
Yeah. Yeah.
David Joy:
Amazing. Nik, well it's been an absolute pleasure, man, having you here, it's such a fantastic conversation. What I wanted to understand was, where can people follow you, and track some of the awesome stuff that you and your company is doing? Tell us a little bit about that.
Nik Mayer:
I'm on LinkedIn. Because I sit in the security space, I'm not massively public social mediad, that's a whole different discussion and a whole other... We could spend two hours talking about that. But I'm on LinkedIn. It's Nikolas Mayer, you'll see me, I have a big beard, there's young Tony Stark working on a Iron Man head at the top image, can't miss me. Maximus on there too, and that's really where I post anything business related.
David Joy:
Well that's awesome. Well, I hope to see you at some re:Invent. But yeah, it's been a absolute pleasure having you here, Nik. It was a great conversation and hopefully you had a good time talking about your world.
Nik Mayer:
This has been a blast, David. I'm happy to do it anytime you want.
A podcast for architects and engineers who are building modern, data-intensive applications and systems. In each weekly episode, an innovator joins host David Joy to share useful insights from their experiences building reliable, scalable, maintainable systems.

David Joy
Host, Big Ideas in App Architecture
Cockroach Labs
Latest episodes

Introducing Cockroach Continuum | A Big Ideas in App Architecture Exclusive
Tara Shankar Jana "TJ"
Senior Director Product Marketing @ Cockroach Labs

A Love Letter to the Database: Industry Shifts, Lessons Learned, and What's Next with Perry Krug
Perry Krug
Manager Solutions Architecture at Baseten

The Everything Trap: Building AI Software That Lasts with Sam Hilsman
Sam Hilsman
Co-founder and CEO of CloudFruit

Why Inference Engineering Is the Next Big Role in AI with Philip Kiely
Philip Kiely
Author of Inference Engineering | AI Education @ Baseten

Distributed Systems, Linkerd, and the Cost of Network Calls with William Morgan from Buoyant
William Morgan
CEO @ Buoyant, creators of Linkerd

Making Software as Durable as Data with Peter Kraft from DBOS
Peter Kraft
co-founder of DBOS

Breaking the Pillars: Rethinking Observability with Charity Majors
Charity Majors
Co-founder and CTO of Honeycomb.io and co-author of Observability

How to Transform Dev Workflows with CI/CS and AI Agents with Tomer Karin
Tomer Karin
Embedded Software Architect

AI, Market Cycles, and the Systems Built to Outlast Them with Cockroach Labs CEO & Co-founder Spencer Kimball
Spencer Kimball
CEO & Co-founder Cockroach Labs

How to Scale Data Infrastructure from Startup to Enterprise
Nishant Raman
Data Engineer at FinTech Company

How to Build an AI-Native Organization
Peter Mattis
Co-founder and CTO/CPO at Cockroach Labs

Inside Infrastructure as Code with Pulumi’s Founder & CEO
Joe Duffy
Founder/CEO at Pulumi

Inside Ericsson: How AI and Automation Are Shaping Telecom
Anand Bajaj
Chief Architect - 5G Network Slicing at Ericsson

Unboxing the Cloud: AI, Microservices, and Resilient Databases
Jim Hatcher
Solution Engineer at Cockroach Labs

Strategic AI and Cloud Solutions: GitHub’s Blueprint for Modern Development Success
Ari LiVigni
Senior Cloud Solutions Architect at GitHub

Cloud Architecture in the Public Sector: Balancing Innovation and Security
Nick Mayer
Principal Cloud Architect at Maximus

GenAI Meets Celebrity: Inside Cameo’s Journey from Startup to Stardom
Dom Scandinaro
CTO at Cameo

The journey from mainframe to adopting generative AI with Equifax’s Senior Network Architect
Samarth Shah
Senior Network Architect at Equifax

Modernizing your cloud strategy with OneStream’s Senior VP of Cloud Architecture
Ryan Berry
Senior VP Cloud Architecture at OneStream Software

Driving digital transformation with Chief Architect at Altimetrik, Ignacio Segovia
Ignacio Segovia
Chief Architect at Altimetrik

Discussing the Patterns of Distributed Systems with Unmesh Joshi
Unmesh Joshi
Principal Consultant at Thoughtworks and Author of Patterns of Distributed Systems

How to simplify your software architecture
Rob Reid
Technical Evangelist at Cockroach Labs

Behind the scenes with Vimeo’s Director of Enterprise Architecture
Sachin Joshi
Director of Enterprise Architecture at Vimeo

How to leverage real-time data processing for enterprises
Andrew Sellers
Head of Technology Strategy at Confluent

Inside the Mind of the Chief Architect at Index Exchange
Joshua Prismon
Chief Architect at Index Exchange

Solving for Scale: Real-time Retail Experiences with Endear's CTO
JP Grace
Endear

Data, Acquisitions, and AI: Insights from FiscalNote's CTO
Vlad Eidelman
CTO and Chief Scientist at FiscalNote

Discussing Data Trends in the AI Era
Gajanan Chinchwadkar
CTO at Hypermode

Unwrapping Moonpig: Architectural Insights into Personalization and Scalability
Alexis Lowe
Principal Engineer at Moonpig

Solving for data intelligence at scale
Madalina Tansie
Chief Technology Officer at Collibra

Simplifying solutions architecture with Brian Johnson of Booz Allen Hamilton
Brian Johnson
Sr. Solutions Architect at Booz Allen Hamilton

How to make your applications smarter
Rod Senra
VP of Engineering at Loadsmart

Scaling for 2 billion events per day with Principal Software Engineer at Red Ventures
Majid Fatemian
Principal Software Engineer, Data Platform at Red Ventures

The data behind digital marketing: A conversation with Bluecore’s Software Architect
Mike Hurwitz
Software Architect at Bluecore

A Lesson in Scaling: How Kami handled 25x growth with CTO and Co-Founder Jordan Thoms
Jordan Thoms
CTO & Co-Founder at Kami

Mastering Multi-Cloud with PwC’s Erol Kavas
Erol Kavas
Director at PwC Canada

From FedEx to Five Guys: Designing digital experiences with Yext’s VP of Software Engineering
Matt Bowman
VP of Software Engineering at Yext

Reliability and scalability in a data-driven world with Fivetran’s VP of Platform Engineering
Mike Gordon
VP of Platform Engineering at Fivetran

Enabling a data-driven and innovative engineering culture at Amplitude
Shadi Rostami
SVP of Engineering at Amplitude

How Estée Lauder scales strong engineering culture
Meg Adams
Executive Director of Platform Engineering at Estée Lauder

Can I take your order? Building conversational AI to improve the customer experience
Akshay Kayastha
Senior Engineering Manager at ConverseNow

Engineering resilient systems: Rescuing old treasures and unleashing modern capabilities
Marianne Bellotti
Author, Engineering Leader, Systems Geek

The Full Package: How Route architects its all-in-one post-purchase platform
Siddhartha Sandhu
Engineering Manager at Route

A historical journey in developer technologies
Mike Willbanks
CTO at Spark Labs

From Legacy to Cloud: Success stories from migrating mission-critical applications
Kishore Koduri
Senior Director of Enterprise Architecture at Ameren

Building purpose-driven engineering cultures
Jason Valentino
Head of Engineering Enablement at BNY Mellon

Modernizing Insurance Application Architecture at New York Life
Mike Murphy
Corporate Vice President and Life Insurance Domain Architect at New York Life

Innovation and Disruption: How Materialize pioneered a new era in data streaming
Arjun Narayan
Co-Founder and CEO at Materialize

Stories from an SRE: How Hans Knecht builds better developer experiences
Hans Knecht
Cloud Consultant at Knechtions Consulting (Ex: Capital One; Ex: Mission Lane)

Inside Chick-fil-A’s infrastructure recipe for a perfect customer experience
Brian Chambers
Chief Architect at Chick-fil-A Corporate

Modernizing from the Mainframe: An Exploration of Distributed Systems
Chris Stura
Director, PwC UK

IoT Standards & Data Mesh: Utility Facility App Architecture
Grant Muller
Vice President, Applications and Technology Architecture at Xylem

Relational Data Problems: Doubble Dating Application Architecture
Mattias Siø Fjellvang
CTO & Co-Founder at Doubble

From Legacy Systems to Limitless Scaling with Paycor’s Systems Engineering Fellow
Adam Koch
Systems Engineering Fellow at Paycor

How to Understand Problems & Build Better Software with Technical Leader Joe Lynch
Joe Lynch
Technical Leader

Observability in the Cloud & Dataflow Modifications with Yolanda Davis from Cloudera
Yolanda Davis
Principal Software Engineer, Data Flow Operations

Early Days at Google & Building CockroachDB with Peter Mattis
Peter Mattis
Co-Founder and CTO of Cockroach Labs

Database Benchmarking Efficiency with OtterTune’s Andy Pavlo
Andy Pavlo
Associate Professor of Databaseology at Carnegie Mellon and Co-Founder at OtterTune

Observability & Statelessness with TripleLift’s Chief Architect
Dan Goldin
Chief Architect at TripleLift

Understanding AI: PubNub CTO Stephen Blum’s Key to Faster App Development
Stephen Blum
PubNub

Building reliable systems with DoorDash's Matt Ranney
Matt Ranney
DoorDash

Real-Time Data Capturing: The Future of Fitness Technology
Paul Lawler
Head of Software at Wahoo Fitness

Building Efficient App Architecture with Alloy Automation’s Gregg Mojica
Gregg Mojica
Co-Founder and CTO Alloy Automation

Unleashing the Power of Hiring Software with Greenhouse CTO Mike Boufford
Mike Boufford
CTO at Greenhouse Software

Decoding Data Warehousing: Insights from Ken Pickering, SVP of Engineering at Starburst Data
Ken Pickering
Senior Vice President of Engineering, at Starburst Data