
49
The journey from mainframe to adopting generative AI with Equifax’s Senior Network Architect

Samarth Shah
Senior Network Architect at Equifax
In this episode, our host David Joy speaks with Samarth Shah, Senior Network Architect at Equifax to learn about the steps they took to modernize from the mainframe to the cloud, as well as the impact AI has on the industry.
Join as we also discuss:
The limitations of the mainframe and what enterprises need to consider when moving towards a cloud architecture.
How to address modern security challenges including IAM controls, network segmentation, and DDoS protection.
Innovative use cases for generative AI in network security, including how to automate defensive actions during cyber incidents.
David Joy:
What is up everyone and thanks for tuning in.
In today's episode of the Big Ideas in App Architecture Podcast, we speak to Samarth Shah, a senior network architect at Equifax. Samarth and I talk about the evolution of network security in the light of his learnings working at Equifax and modern standards of network security layer required to secure and protect data today.
He also talks about his experience modernizing applications from IBM mainframe into the cloud. It's a fascinating conversation filled with interesting bits. So pump up that volume and enjoy listening in on this episode.
So welcome to the podcast, Samarth. I know we had a bunch of rescheduling to do, but we finally made it. How are you doing today?
Samarth Shah:
I'm doing great. Yes, finally we made it today. And I'm on a vacation this week. So yeah, I'm in a very different space than what I'm normally in.
David Joy:
Well, awesome. Well, I'm glad that you're getting some time off and thank you for taking the time to come on and like chat with me on the podcast.
Now, for everyone listening in, you know, Samarth works at Equifax. And what I would like without butchering his intro is, like, Samarth, why don't you tell the people who are listening in right now a little bit about yourself and your role at Equifax?
Samarth Shah:
Absolutely. So, I'm Samarth Shah. I work at Equifax as a Senior Network Architect.
It's been close to six years now since I've been with Equifax. But I started at Equifax in early 2019 as a Lead Engineer at that time because the company decided to go through a cloud transformation journey. And as part of that, there was a lot of restructuring, new team building.
And also we were coming off from a security breach. So there was a lot of restructuring and putting the right people in the right place all across the world. As part of that effort, there was a new position created and then I got an opportunity at Equifax and I came, as I mentioned, I came as a Lead Engineer.
After two years, I was promoted as Cloud Squad Leader, running a team of 10 people, including global regions. And then a year and a half ago, I was promoted as an Architect. So there was a journey along with the journey of Equifax and cloud transformation.
Even I got a chance to grow, which was the best part so far. What I do... So in the current role, what I do is I review the business cases with all the stakeholders from different business units, the apps, security, compliance, budgeting, finance. I've worked with all of the stakeholders on what the business requirements is.
It starts from there and then pretty much work with security, application team to understand the requirements and then make sure that whether we have all the security controls in place. Equifax is big, big on security now. After, I mean, this happened, there just really no appetite to go through or even have a small incident on that list.
So working with security after that, working on the engineering, design, architecture, get approval from all the stakeholders. Then do the POC, proof of concept is done, and the budgeting of that environment, what is needed, resources, technology, a combination of everything, then the rollout. So I work closely with engineering teams for the implementation, not just in U.S., but also across the globe.
We have 10, I mean, 24 countries point of presence, but really our infrastructure span across five global regions, Europe, Canada, Latin, India, Australia, and of course, U.S. After that, then I help the engineering team to hand off to the operations team because then it's all BAU and what the observability and monitoring looks like so that we keep maintaining our SLAs with our customers, internal and external both.
So this is the cycle that I live through here. For any new business case, it evolves into what it means for the application, then security, and then POC, engineering solutions, design, architecture, then budgeting, getting the products, rollout, and everything we do here is through automation. So that automation lifecycle has to be also part of it.
So I started in one area when I came out of Equifax, which was just lead engineer for the engineering and implementation. And then I kind of got a chance to evolve from that, seeing this end-to-end lifecycle with the leaders that I worked with for so long. So that's my role, and that has been our journey of me at Equifax.
David Joy:
Yeah, that's a lot of interesting things that you've touched on. So starting off as a person working on engineering to leading a team, and now being one of those decision makers who help with what technology and what the security requirements need to be and what protocols need to be followed and some things that they have to make sure that they have in place. And it's a pretty big responsibility now.
So that's good.
Samarth Shah:
Yeah, question.
David Joy:
For people listening in, I know Equifax is a kind of, in the North America at least, is a pretty well-known brand. We know what some of the things that Equifax does. I mean, from a credit score point of view. Can you elaborate that a little bit about what Equifax does and maybe some things that we don't hear about Equifax?
Samarth Shah:
Yeah, so you're right. I mean, as a general, it's a pretty big one background, Equifax is a credit reporting agency because there's big three, TransUnion, Experian, and Equifax. Big three. So that's very general that, okay, yeah, Equifax is a company that is a reporting agency. But actually, to be very honest, and I'm glad I got this platform to answer that question because we are a lot more than that, especially the cloud transformation journey that started in 2019. We have evolved from a fintech company to a data analytics company now. We're just not a financial industry. Yes, we are. The core is in fintech. But we have evolved from that to a data analytics. And I'll explain that in a second.
So that's the first thing. And second thing is, yes, other than credit reporting, we are a data company. Basically, we collect data, and then we aggregate all the data, and then we sell it.
So that's pretty ... and we have a pretty niche there. So depending on who our customers are, if it's healthcare, banks, auto, we aggregate the data and then sell that data, and then they can make a business model out of it and whatnot. That's one area which has grown significantly.
This space has really evolved in the last four or five years because we are not spinning our head on what we can do with this data, we are just aggregating and creating models. Okay, this can be this, this can be part of this industry, and then we really partner with our ... Sometimes it turns into partnership. For example, recently, I don't know if you have seen the news, Equifax signed a partnership with Workday. What that means is today, for unemployment, if somebody ... What is the process? Like, you apply for a job online, then they reach out to you, you get through the interview. Everything is done. You get off later. Then what is the next step is you need to have a clearance, background check, income verification. All that happens, right, which is another two weeks process, depending on what the nature is. Workday has streamlined that process for Equifax.
When you apply a job through Workday, it already takes care of that process. We are not waiting an additional two or three weeks. So, you know, it saves so much time and it frees up the time for HR to do something else.
So that's one area, and we are making partnerships like that. Then we are also working with the banks' customers directly, where we can work with them on creating ... Let's say, for example, Chase is coming to us that, "Hey, you know what? We are pulling this data for auto. We pull this for housing mortgage, and all that stuff." Going forward, we want to be able to separate the soft inquiry and hard inquiry. Because the soft pull does not affect your credit score. Hard pull does. So we don't want to have two hits and then affect the consumer's score. So help us to separate that."
There's a lot of things like we do, and I don't know if you remember, but in COVID, Equifax was one of the first companies who decided and partnered with the banks that, you know what, we are going to freeze the credit score for the next six months.
David Joy:
Right. Right.
Samarth Shah:
So things like that. Yeah.
David Joy:
Yeah, that's pretty cool because when you said that, I can relate with that so much because every time I have to do something, I'm like, I'm going to ask somebody, "Is this a hard pull or a soft pull?" You know, because at the end of the day, I don't want it to affect my credit score.
But overall, I think it's really good to know that, yeah, you know what, Equifax is all about. It's not just a credit scoring agency, but you do so much more. And that brings value to people like us who are using these different solutions.
Like I use Chase. Chase is one of my accounts. But I did not know in the back end those guys do this, and they have that intentionality towards how they are designing everything. And there's a bunch of companies that work together. That's really cool.
Samarth Shah:
And one last thing I would like to add real quick is that for Equifax, because it's a very common perception and very generic perception that Equifax is, like I said, you know, data analytics company now. But with COVID, you know, the mortgage business and the unemployment business went both high.
Unemployment went up and then mortgage people were buying homes because of low mortgage rates. So they both were peaking the business. But after that, the mortgage business kind of declined. And we were really, we were not struggling, but we really did not hit our numbers. And there was a significant visibility that how much we rely on mortgage industry. So our leadership and while we were going through this transformation journey, we, you know, our strategy was to invest more into another growing space, which is, I'm sorry, another growing space, which is cyberspace, prod risk monitoring. So basically we offer two products, credit monitoring and then risk protection on your social security. If there's any attempt of stealing identity or somebody's trying to use your credit insurance number.
Because remember, you give your credit score all the time and everywhere, you know, on the phone, you're talking to some banks, they ask for credit score, I mean, social security all the time. So how do we protect that? So we have really invested in that space and tried to create more niche products for the consumers so that we also eliminate the dependence on mortgage business as well as we diversify our business from that standpoint that we are also selling the security products. And with that mindset, we acquired 12 M&As in the last three years. And most of them were focused on the risk products that can be integrated with Equifax products.
David Joy:
Nice. Well, that's awesome.
So I know pretty much all of these areas, even when you do M&A, companies bring in their own software, they bring their own cloud solutions. You have your existing legacy systems and you're also on a cloud transformation journey. So I know there's a lot that goes on within an organization such as yours in that scale.
Tell me a little bit about right now, which area are you focused on? Like, is it the credit score side or is it these products that you just mentioned or is it all of it?
Samarth Shah:
Yeah. Good question, actually, because it can be so overwhelming at times when so much is going on. So my idea is basically, I will try to simplify this. My role stays within the first, since we are talking about M&As, how do we integrate M&As? And when I say integrate, I do not just migrate them to our environment. That's a simple thing. A lot of M&As that we have acquired, their security posture is not there where we want it. And we invested so much over the last couple of years. We want all this M&As to be at the same level. So I do an assessment of all M&As of their environments. What does their application security posture looks like and the infrastructure posture. So if it's matching to us, we integrate them. I build a ... you know, I come up with a design, a pattern that can integrate them with Azure's environment and give them a path to move their workload to our environment.
One thing we also, at Equifax, have raised is that we are not entertaining lift and shift migration. We only want to do ... We want to modernize our application. So depending on what the app architecture and environment is, we definitely do an assessment, not just on security controls, but how they are deployed. What is their app requirement? How do we bring them in our [inaudible 00:14:02] and establish pattern? So that helps in compliance too, down the road. Because what happens is that if you miss this train here, you're dealing somewhere in compliance, operationally, missing the TRs, technical requirements. It's hurting you somewhere. My role is to do that assessment thoroughly and then come up with a plan, determine that this is what it looks like.
Second is, yes, we invested, I mean, we are the biggest customer for Google in Southeast. I mean, we invested $1.5 billion with Google Cloud. That's our partnership. So the goal is that, yeah, we're going to be hybrid On-Premise now, we're going to go hybrid, but we're going to get rid of our 24 data centers now globally and then be 100% cloud with some small footprint in AWS. So that's the big goal. We have gotten 80% done. We still have some data centers that we need to shut down. So that's my another area.
The drive application team, drive the dependencies from the data center to move it to the cloud. And the biggest hurdle right now that I'm facing and I'm trying to solve is IBM mainframe. A lot of these companies, they were using IBM mainframe in '90s and it's custom written code its skill set that doesn't exist today.
How do you unwrap and wrap it around? And how do you, and again, I said, lift and shift is not an option. How do you understand and how do you build something [inaudible 00:15:35] in the cloud? So that's the second, like really, I'm really knocking my head out on that one. So that's the second area. So that way, once we solve that, then yes, everything else is pretty much [inaudible 00:15:48].
And the third thing is the new ... What I realized, David, and I think this will come up, I'm sure, in the next question. I thought once you build an environment, right, network connectivity to the cloud and everything is done, you provide the connectivity to promote their workload, you're done. Now it's more like compliance and run well at operation and servicing. No, actually nowadays, it's more than that because cloud is coming up with the new features every day. So you have to keep very close eyesight on them that, oh, what's your standard pattern was like A, B, C, but now they have simplified to A and B. Maybe I need to simplify that in my environment too, to keep up. So that's another area. I'm constantly optimizing my architecture, my security posture, trying to keep it simple. That's the goal and that's across the board.
David Joy:
Great. Great. Yeah.
Samarth Shah:
So hopefully ... Yeah.
David Joy:
No, I was going to say, I wanted to talk to you about network and network related activities that you do towards security. However, you opened up Pandora's box by mentioning mainframes. It's just something that I'm also myself thinking about and actively, like, we are a database company, right? So obviously we work with a lot of folks who are in the financial sector and we're working on existing mainframe systems. And then you see an app that's like 1 million lines of code and is running on IMS or DB2 and they're like, you are trying to fit that into the cloud architecture and it's just a big, big, big migration. So tell me, I want to get your perspective on this. How are you tackling it? Like, how do you look at the mainframe stuff because you brought, and especially in the context of modernizing to the cloud, how do you look at all of that? Yeah. Let's go deep into-
Samarth Shah:
Yeah. Yeah, you're right. So the way ... The exercise that we are doing right now is that we call it Fabric, Fabric Mods. So because in the product name that we have, which is on main premise InterConnect. Basically that's where all the data is today, the 140 years data sitting there. Right? So it basically sorts ... The call goes to the IBM mainframe and then they have ... The code is written in such a way that it goes to different Fabric. So what we are trying to do is, we are picking one use case at a time.
So after so many months, we have finally understood how many use cases are there. So I guess my answer to your first question is, part of the first answer is, try to do the discovery on how many use cases you're dealing with here. So then now you know what ... you are understanding your own dependency in a way. So that definitely took a lot of time that, because we did, like I said, we built a new environment in Google. We moved a lot of dependency from mainframe, but there was still significant dependency on IBM mainframe that we discovered and we don't know what that is. So that's why we did that exercise.
Let's understand what is this being used for? So we came up with the use cases and now we know which ones are the sensitive ones, which one we cannot touch. Like we prioritize that, we touch these use cases according to their sequential number. First, so the less priority, the less impactful we can start from there. And now we have identified three use cases where it's not a significant, I mean, I shouldn't say revenue, but the size of the pool is not as much. They do after hours, it's pretty much idle on the weekend, we do all that analysis. When is the peak hours, when it's not the peak hours, so now we have that data. So now we feel there is a larger window to do anything with this particular three use cases.
So that's the second step we did. Let's see what is our change window here. How much time do we have to play around it? After that, we engaged our developers that if we just want to eliminate these three use cases and what we need to do, even if you don't understand this code written on the IBM mainframe, but what it's doing, what will be the equivalent application that we need to create in-house, homegrown, and tell us the effort for that. And before you tell us the effort for the last year, is this even possible? So let's build a custom thing. And then we came, our developers came, they asked for the time that this is what we need. This is the first infrastructure. Everything was already in the cloud. Build that.
Then we transitioned the calls and the code one at a time. And there were times where the call was still failing because there's so many sets of instructions in there that you miss. There is a dependency of so many interdependencies call because even in Equifax, there is a lot of inter apps where the call goes. And if you miss that, then it's failed.
So we are doing that exercise as we are speaking. We haven't solved those three use cases completely, but this is the exercise we are going through. And I feel this is the safest way and also understanding what we have.
David Joy:
Right. Right. Yeah, I like what you said. Like, one of the bigger issues with mainframe is that, especially with folks, like you and I have not worked on these systems, right? We have inherited it. And some people, and I have a lot of respect for folks who did work on mainframe and built all of this at a time when we didn't have all of this, right? And going back and basically it's like trying to figure out and reverse engineer what those guys were thinking when they wrote all of that code and when they built all of the system.
So I agree with you. Like, that's the approach. But one of my questions I always ask is, why leave mainframe? Like, why do you think ... Because these systems are great, they'll probably keep working even after 50 years, right? I mean, look at, AS400s or PCs or Z series is obviously mainframe. They keep going on for a while. But what's the reason to leave mainframe? I know I've heard cost. So if we remove cost, what do you feel is the main reason that he's leaving the mainframe?
Samarth Shah:
That's a great question.
David Joy:
Yeah.
Samarth Shah:
That's a great question. And as an engineer, I always relate with what you just said. Why move from something that is working? Well, like a well-oiled machine, I mean, what's the reason to create work out of it, right? I agree with you.
And this is something you always, and I'm glad I have a team that I work with is, we always try to think outside the box. And just like you said, there are people with us who have never worked on IBM mainframe. So to them it's a new thing, but we still admire what's out there. Wow, there's something. And imagine somebody wrote this code on his own. But let me get back to your question.
If you ask me, my assessment of this question that you just asked is, yes, I think cost is, yeah, it's there, but more than cost, I feel like it's a skill set. How many people do you find who can really work on IBM mainframe today? I don't think, and it just amazes me, for something, an IBM main product, and I think PenTec company has IBM mainframe, and there's so much reliance of big, big industries on it. Why there was no such awareness of, like, for example, I'm a network guy, Cisco CCNA, everybody knows about. You name it, if anybody's in the networking industry, Cisco, CCNA, [inaudible 00:23:46]
David Joy:
The CCNA.
Samarth Shah:
I'm sure, yes, that's the first thing they know. So that branding or awareness was somewhere ... It was lost for IBM mainframe, I feel like. So I think the leadership, the companies who ended up, you know, taking this, the accountability in leadership roles at these companies where there's IBM mainframe, they were not comfortable, I feel like.
We have something that I don't know, and it can just break one day. Maybe it's not going to break one day, but you cannot innovate on top of it. So you cannot monetize on top of it. It's literally, you're just taking care of the lights that is on. And I think there's a big drive here that let's move to something that we can understand. There's a skill set out there for that, and we can innovate. And, you know, infrastructure, as an infrastructure guy, you know, we make changes all the time.
David Joy:
I know [inaudible 00:24:43]
Samarth Shah:
So it's important that we know our platform, that what is the risk level. I think that's my assessment. That's why people really wanted to.
David Joy:
Yeah, yeah, yeah. I mean, I wasn't looking for Equifax point of view here, but I wanted to get your perspective. I think you got, you kind of covered some of it, right? Like the modularity, the ability to be flexible, the ability to scale onto the cloud and be able to, like, do different things.
Like, I feel like I've spoken to a bunch of people who've done this mainframe, and we are sharing opinions here, right? Some of the guys I spoke to, they were like, well, back in the day, they just wanted to solve one problem. And they were like, they were not anticipating what else will happen with the existing data or the way that things are going. Today, if anybody wants to do data analytics, they need data. Some of this data sits in IBM mainframe, but you want to run Python programs along with generative AI models. You can't do these on these systems.
Samarth Shah:
Compatibility, yeah.
David Joy:
Yeah, the compatibility is a big issue. So especially where the world is going, these systems are sort of obsolete, right? And then the modularity and the compatibility is not there. And then obviously you said the skill set, right? And I noticed these patterns pretty much everywhere.
Samarth Shah:
But problem is-
David Joy:
That's such a challenging problem. I have worked at four companies. I was working at one of the companies in the past that was providing IBM mainframe for people whose mainframes would go down, you know, for the disaster recovery company.
And there was this funny, funny story of once they were moving data centers, and then they found an IBM machine, like an AS400, not a mainframe machine, that was sitting under somebody's desk. And apparently was there for 32 years. They did not know what it was doing.
So then they plugged-
Samarth Shah:
[inaudible 00:26:26]
David Joy:
They plugged the AS400 out and they got a mail from somebody in Arizona saying that, "Hey, my shop's computer doesn't work." It's like the craziest thing that could happen, you know, because the one thing you have to give processor these systems, IBM machines are like they run and they are always available, require low maintenance. However, where we are going, it seems like they're losing what they could do, you know, in the future.
So that's why when you mentioned it, I was like, it is one of my fun topics to talk about. So I brought it up a little bit more.
Samarth Shah:
Well, I know you're right. I mean, and again, like you just said, I mean, modularity is a big problem because you cannot do any innovation and more innovation you have to take. And who's going to do it on IBM mainframe? People who wrote the code, is not there, the skill set is not there. I don't think, you and I wouldn't mind to learn that, but who's going to show us that?
David Joy:
Yeah, I wouldn't mind learning that. The problem is, I mean, I just, it's not just fun for me. It's no longer fun enough.
Samarth Shah:
It is no longer-
David Joy:
If it was three years ago, yeah, but... COBOL was introduced, it was like the thing, right? Everybody wanted it. And now people want to do something else. They want to do Python or they wanted to go, you know, the more programming, you know, languages that are more interesting for people nowadays.
So it was good to dive into that with you a little bit, but then let's go in. I mean, obviously I wanted to talk to you about the security part of things, right? And I know Equifax had a big breach and everything that happened too is, like, they came out and your company came out with a lot of protocols and process and how you wanted to change that from happening and you were integral to that. And I know you won an award with... you also won an award with Google Next or at Google.
Samarth Shah:
Yeah.
David Joy:
Yeah. So tell us a little bit about what was all that about? Yeah.
Samarth Shah:
Yeah, absolutely. So, yeah, when the breach happened and, like I said in the beginning, there was a revamp of restructuring of leaders, teams, new skill set, everything happened. But if you ask me, yes, when I came on board, I could feel that there is so much pressure on everybody's job in a way that even if you have a seven letter password, you will get an email, "Hey, your password is not meeting compliance requirements." I guess people were just so terrified that let's not take any chance. That's just an example. I'm just trying to give you what the vibe was when I came in. Right?
And on top of that, now, I'm part of network team, our job was to, "Hey, we want to go to cloud. Our systems are old. We cannot do..." And that's the other problem with the complex, the whole system that you cannot have so many security controls on it. It's very basic and there's only so much you can do.
So anyways, the decision was made, let's go to Google and AWS. At the time, AWS was also our primary provider, but we changed to Google later on. Let's go to Google and now, let's map it out what that environment looks like.
Now, nobody knows, cloud is a new thing. I mean, if you ask me, let's go to cloud. Me, okay, I mean, all I knew is that VPC, that's it. But how do you get to VPC? What should be in the VPC? It's all unknown to me. And now we're talking about security controls. So I need to understand that product because then I can secure something. Right now, I don't know what that is. How can I secure something? So, you know, that's the security rule.
So what we did, the security stakeholders, the security architects, network architect, all of us, for the first few weeks, all we did is let's write down the TRs, technical requirements. Let's forget about design. Let's forget about what Google can do, what AWS can do. Let's forget all the vendor capabilities for a second. What do we want? What security layers do we want? What sort of security controls do we want? Layer 7 inspection, inbound decryption, SSL-powered decryption, IAM controls. Okay, IAM controls, but at what level? Live service account, then creating the entitlements for every function. Right? Network segmentation, big, big, big on it, right? How do we segment MPs, UATs, production environment, period.
David Joy:
Right, right. 100%.
Samarth Shah:
Like, period, with GKE clusters, you know, you can have multiple apps running on the prod and that becomes a cost-saving approach very easily. Hey, you know what, if I'm building a GKE cluster that I'm paying so-and-so money and if it's an MP and UAT, why not run multiple apps? And that's how it starts.
David Joy:
Yeah, yeah, yeah.
Samarth Shah:
And then you start missing things, right? So we came up with that process that what is strictly allowed and what is not allowed. So segmentation across the board and then granularity of how we enforce that. So network segmentation can be done through firewall rules, security groups, backers, network policies for the GKE clusters.
And mind you, I'm not a GKE guy. I understand what it is, what you need to deploy, what the infrastructure looks like, but I'm not a [inaudible 00:32:05] on by any means what the GKE configuration looks like. But network policy, I had to educate myself to a point. Network policy is something and namespace policy for the authorization. So all these things.
And then MFA. We want MFA period. So not just to sign in, but also the application, calls, making a call to another application. If a service account has a certain more than ... I mean, I need a role on more than certain access. How do you make sure that it's coming from a legitimate application call? So we invested in Ping ID, so MFA. Everything has to go through Ping ID.
So that's another thing. So this, so again, we created a list of the network TRs and security TRs, what these things we want to keep it in our environment. And then that went through multiple reviews, revisions. And then the next effort was, how do we implement that? Just a mapping of the products, not the architecture. Okay, IAM controls. We decided not to use cloud provider rules. We created custom rules.
And to begin with, let's just understand from application team what they will need. Let's create custom rules. If they don't have a certain access, we'll add it, but let's not give them blanket access and then figure it out.
So those things, that's where we spend the most of the time, David, for the first few months, I would say.
David Joy:
Yeah, yeah, yeah. I guess-
Samarth Shah:
I mean, the three. And then we got OSM educated, the mappings of the product, and then we started talking to vendors. "This is what we want. This is how we want. What are your options?"
And then we started creating the draft architecture. What the network segmentation looks like in order to have this network segmentation? We need to have, for example, I'm going to speak about Google. Have a separate service project, I mean, a host project for non-prod, UAT, and prod.
We have no exceptions on that. You're going to deploy NET and NP host project, and then you're going to create multiple service project, [inaudible 00:34:25], UAT and prod. You will never be able to have access from production environment to your lower environment. Never. Either way around is fine. If you're promoting your code to NP, to UAT and prod, that's allowed, but not the other way around.
Your application will not be exposed to internet services from your NP and UAT, only from production.
So those things, we came very strictly about it, and then creating multiple attacks. On top of that, what we did is we reinvested, and that's where my role was very critical, work with all the vendors. And again, Google was not as mature at that time. So this controls that we had, they said we cannot meet, but we have a roadmap. So that's one thing, you know, it's cloud there. So they are learning from you too, at the same time, from your own environment.
David Joy:
Yeah, yeah, yeah. Yeah.
Samarth Shah:
No, I mean, it's totally fair.
David Joy:
Yeah. Yeah. One of the things I liked about some of your responses, so realistic, right? Like it's ... What you are sharing here is a very, very, very, very high bar of what needs to be done to set up these level of processes and protocols in place.
Like a lot of times, what I've realized about networking, and I'm not a networking guy, when we spoke, I told you how much of an anti-network person I am. I've no ... I understand enough to be able to set up my infrastructure and, you know, cut the internet out and think over the internet traffic and things like that, and set up my route tables and stuff like that. But what I'm going behind that point key is that if you look at developers when you start developing something, they're single-mindedly focused on I just want to build something, I want to see my MVP. If this works, okay, I'll take it to the next level and I keep building on top of it.
But now with the way things are and the state of the world is, you cannot do networking like that. But you have to be very clear about what you are going to do, what your network, what your security protocols are from the very beginning, or at least have a very high-level idea of what it needs to look like, and then build on top of that. Right? And what you presented in just a few very good explanation to what you did is tell what is required.
And I remember when you [inaudible 00:36:40] the company, we add production round, we talked to non-production round, we're pushed from non-prod to production. It's nice. And it was a small company and we had all of that window. It was not like Equifax. We didn't have like a PI data or anything. We just had user information. Right? Maybe tokens and things like that. But what you're saying is so critical.
Samarth Shah:
Yeah, it is, it is. And that's what it is.
So when I was working with this is that we always hear about breaches, right? And if you really do the study of those companies' breaches, right? It's very simple thing. You miss somewhere something. It's not a big thing in a way.
So for example, our environment, now these are the controls that we wanted to have from the application perspectives. And it's funny that you brought up how frustrating it is that what is your requirement because I'm trying to find out my requirements, but my requirements somewhat depends on how you're going to deploy your applications. So if you're going to need, I need you to put network policies or you need to put default deny in the beginning, but then you need to know what your application is going to talk to.
So that puts more thing on you. And so I think there were some, I wouldn't lie, there were some frustration in collaboration like that because like you said, you explained that I can set up my environment, I can go live. But that was not happening.
And now everybody was supposed to learn their own thing in very basic foundations. And then, so that was a frustrating exercise for sure at some point, but in the end, I'll tell you this, it was the best exercise. And we all really grew so much with each other's support.
And now whenever I design or think of another pattern or application or architecture, I can see and learn because of the understanding I got from everybody on their app side. I mean, for example, I don't know what the standard services does, but it took me a while what it does after I learned from them. I don't know what Cloud SQL or Composer for that matter.
What are the Composer constraints from the networking perspective? I don't know that, but I learned all that from them and that helps us to design an architecture where we will not hit certain limit from Google. Scalability is gonna be there. We don't have to re-enable our architecture that we put in place at that point and we can still accommodate the growth.
So that was the...
David Joy:
One thing I'm curious about is, I'm just like... I think we have three categories. Like AWS, GCP, Azure, top three. Each of them have some benefit the way they are. Like, when I look at the portfolio of tools and services, I see AWS has a lot more services than Azure or Google. Google has a unified network.
You don't have to really... There are some VPC related stuff, some networking capabilities that you have to do with AWS or Azure that you don't have to do with GCP. So when you started getting into GCP, and in just a very specific GCP question, what did you like about GCP's network compared to AWS's?
Samarth Shah:
Great question, and one of my favorite too. Because as I mentioned in the beginning, when we mapped out all the products for our controls, I was speaking to... I remember we had an AWS Pro Server and Google PSO team on site for a couple of weeks every month just to do this whiteboarding sessions, right, for what we are trying to do. And I'm sure you have heard of this.
David Joy:
Yes.
Samarth Shah:
So AWS has called something, for example, Transit Gateway, TGW, right? And there's a transitive routing. Managed service. And then they, of course, added more features. Eventually, TGW PA, regional, the prefixes limit, attachment to DH gateway.
Anyways, I like that design. So I thought, "Okay, this is good. You know, I don't need to manage my virtual appliance. And it's simple. I can scale up and down. It's everything I see. Launch it. That's it." Turns out Google is not there. But I was very clear in my mind that I want something that is ... a design that is same across the cloud. Not just, oh, AWS has its own flavor and Google has its own flavor. Operationally, how can you support? And when things go south, how do you expect application team to understand everything too? So then I realized, okay, then I was a little disappointed with Google. Come on, man, this is the first thing you should have invested in, in my opinion. Because you know, connectivity is a big thing. But then they shared something with me, which was VPC is global in Google. Which I really liked it because that really simplified my region scalability.
I don't need to create a bunch of VPCs for each region and the route tables and everything. Everything can be managed. Basically, my region was egress centralized host project for network security controls. Right? Now that perfectly fits in as well. And I can spend my multiple regions in the same VPC within the same host project. And I don't have to keep creating this. And it's going to be one spot.
So I really like that. And the second thing about Google is the security products. We happen to use Cloud Armor product, which is managed DDoS service provider on [inaudible 00:41:52]. And then the global load balancer, which is global in nature. So that also helps us to enable our business to go live or inside a cloud because that's what we use to export the services to the internet. And because of the global nature and its integration with Cloud Armor and SSL policies so that we can restrict the ciphers, certificates, PLS 1.2. So I felt like all of those security components were pretty simplified. And Google was very ... They invested heavily in that first. And with AWS, I felt a little bit ... Everything is very ... It looks you have a lot of products and mature products, but then you are adding more puzzle to it. And to keep it simple, it's going to be challenging. Versus in Google, I felt like I can really scale with my design.
David Joy:
Right. Right. Right. Yeah, I mean, I've heard similar stuff from people who love the networking side of things about Google. I think that's been their advantage, personally, in my opinion as well. Very cool. So let's switch gears.
I know we didn't realize we are at 45 minutes already. But what I wanted to ask you was like, you mentioned some of this stuff already, like how you are laying the things out. And one of the things I really liked about is you removed vendors, you removed application folks, you just sat together with your folks and decided what is it that you want? And what is the business requirement? What are the technical requirements? How do we want this to be structured? So in many ways, you laid out what the best practices are. So where do you think folks get things wrong? Like maybe a few, if you can mention, a few things that people get constantly wrong, and it may be super easy to think about when they're designing the network.
Samarth Shah:
Absolutely. So I'll give you an example. This is what happened, and I'm not going to name the consulting company who was with Equifax, but Equifax had a consulting company for helping us to build the environment in cloud. When we came on board and they will not go after that, what they did is, now Equifax is a global company, as I mentioned, five geos with two regions in each geo, right?
What is the one thing that even everybody, even you would answer that question if I asked you, that what is that one thing you need for networking to build your workload, [inaudible 00:44:24], is your site or subnet. You need that routable subnet, right?
Now that company didn't realize and did not plan or estimate that, you know what, this is a global company, this is a global region, and did not really work with that team that, hey, what will be your GK cluster size? How much range you're going to use for your master, Slash 25, like, doing, and then how many regions do you have? So now you're going to have MT, UAT and prods for network segmentation, if I just give you one block and Slash 25 and you start creating MT, UAT, prod from that small Slash 25, firewall rules, routing, and compliance, and it's just crazy to do so much in that one thing.
So the planning of IP CIDR, what we do today is, we pull out a block from Infoblox, you know, I mean, that's what I plan basically. We created a Slash 15 container. For example, I mean, 10.8.0.0/15. .8/16 will be used for MT, .9/16 will be used for prod. Now you use, as that will be mapped to the region, UST score.
So when I now I do all my routing, network connectivity, firewall rules, audit, now you do what you need to do behind the scenes, it's your problem, you manage it, and as long as you keep using this UST score non prod subnet, then I'm cool with it, but also I'm going to hold you accountable that you cannot come up with ... you cannot tell me in two months, "Hey, I used that block already." And I'm like, "Dude, you used /16 in two months, so can you explain?" So yeah, if you have a business justification you can show me that IP exhaustion happened on your parts. Yeah, fine. But so those process were put into place, but this was completely missed.
How did it affect me and my team? Now, basically, we wanted to move from VPN to DEREF. That migration was very challenging, just for the simple reason, the subnet masking was not done properly. They created a bunch of blocks, slash 27. So within that slash 16, they use for 10 different regions. So I had to do, like, we had to be so creative that we, if it was done just for one slash 16, for one region, there was no overlapping. I could have moved one region from VPN to DEREF and easy. And for environment. So I think it's a simple basic thing you miss out sometimes.
David Joy:
Very basic. It's very basic. Yeah. But you said that I was like thinking, "Oh, man, I did that eight years ago." But I completely get it.
Yeah, I mean, everyone who does ... Yeah, like IP exhaustion is like one of the big problems, right? Like, especially when you start playing around with Kubernetes, you're like-
Samarth Shah:
Yes.
David Joy:
I want to like, scale, scale, scale, scale. And suddenly, you realize why is it that it's failing, and then you go back to see a configuration [inaudible 00:47:28] 288 or 287, super low.
Samarth Shah:
I mean, imagine David, [inaudible 00:47:28] I would just like to share this. I don't know how many companies out there are feeling this. But I also go to a lot of meetups. And it's been five years and our GKE, I mean, EKS and GKE, we have AWS and Google Boards. We have run out of routable space from our internal address space. And also keep in mind, we acquired 12 M&As. I have to integrate them too.
So there's an overlapping from that in mind. So right now the biggest challenge is to solve that basic problem because the application team is coming to me, "Hey, we're gonna add more computing power for the prods and we need more capability of adding routable subnet because now we need to talk to Composer, BigQuery." "Dude, we don't have it." So now we're working with IPv6.
David Joy:
Well, that's the only option available. Yeah.
Samarth Shah:
Yeah, I mean, I can see how it can affect so many companies and industries because how the application is being built On-Prem. I don't think it was like that. I mean, you just build your hypervisor on top of that you have VMs and there was no such thing like that.
David Joy:
Yeah. Yeah. Yeah. No, it's fair. I mean, it's a good point actually that you brought up around IP exhaustion. It's very basic, but it like happens all the time. Great stuff, man.
I know we have a little bit more time, but I'm always curious and it's probably because of my personal interest in this. I've been following a lot of generative AI stuff that's happening in the industry. And I got to speak to somebody about six months ago who built a large language model. It's on Hugging Face now. It's called WhiteRabbitNeo. Okay. I don't know if you know about it. And one of the interesting thing about that model is it's been trained on offensive attack and defensive attack. And the intent of the creator, I spoke to the guy, was to help folks learn about building offensive strategies and defensive strategies to protect their networks, protect their application ecosystem and things like that.
From that, I want to ask you, are you seeing any generative AI penetration in the networking space? If so, like how... is that there or not? Or what are your point of views on generative AI in networking?
Samarth Shah:
I'm still trying to figure out how it's going to embed with cyberspace network security because gen AI is definitely helpful in the event. For example, if you're under attack, if you have written the things that, 10 things that we do today can be embedded into gen AI. And we get that output right away and we can make decisions. Like, we skip that step of doing those 10 things that we do today. Instead that just gets to us whenever an attack happens and we have an observability tied to it, which launches Lambda function or Cloud Function or whatnot, then we can make quick decisions. That's one thing I can visualize right now. And we are trying to get there with our research team, who is playing around to do this kind of analysis and testing. But that's so far, that's what I've gotten. And then you can really be on the offensive side.
And the other exercise we are doing is that we, I mentioned Cloud Armor, DDoS service is what we use today, but we are getting attacks all the time. It's been protected by Google, but we have started analyzing those logs and we are creating patterns of it. That way, what type of attacks and what other patterns these attackers are using today.
David Joy:
Right, right, right. Yeah.
Samarth Shah:
So that we can build our signatures to proactively, basically, block, not only block, but also identify that this is this kind of attack, not just an attack.
David Joy:
Fair. Yeah. Yeah. Yeah.
I mean, that's awesome actually. Yeah. But check it out. I mean, if you're interested in it, it's called WhiteRabbitNeo.
Samarth Shah:
Yeah, I will check it now.
David Joy:
Yeah, check it out and you know, one of the first things I did with that was I got access to it. We have this folks, the creator, and then I went and checked it out. And I said, "Okay, write me a DDoS attack script." Within like 10 seconds, reduced a DDoS attack script per week. And I was like, "I could have actually used it to on my own network and see if it worked or not, you know."
And it was a good way for me to learn. And I asked like, "What are the best defensive strategies to protect my network? And this is the kind of network I have." It started giving me some really good stuff.
So even if... Like, what I realized is for me, even if I'm not a network expert, or I don't attack, understand things completely, I could use a model like this to kind of learn. I did the same thing with ChatGPT and then ChatGPT didn't respond to me at all. Because I think it's been processed to keep things, you know, at this level and not give everyone information.
Samarth Shah:
Not give [inaudible 00:52:43].Yeah. Yeah. Yeah.
David Joy:
Yeah. Yeah. Yeah.
Samarth Shah:
But Gemini is doing... Gemini is better now, I think it's good.
David Joy:
Yeah. Yeah. Yeah. No, no, go ahead. Yeah.
Samarth Shah:
No, I was just... I started using Gemini now. So, I moved on from ChatGPT. But I would say one thing on the security strategies is that for the last couple of years that I've been at different industries, and mainly attack surfaces, the more you have attack layer surface, for example, the things that I mentioned to you, but our... My product name is Global Network Security Stack in the cloud. It's just not Cloud AML, it's just not VPC firewall rules, security IAM controls, or network segmentation, it's more than that. We really have observability of every transaction, and layer seven inspection, web application firewall, basically, IPS. So we have that on top of it. And it's every layer. So if your one layer gets compromised, if you have a good alerting system and monitoring system in place, at least you know something changed. By the time it gets to another layer of attack, you have time to react.
David Joy:
Correct. Correct. Correct. No, that's a good point, too. I'm pretty sure it was part of your technical requirement document like this, how you want to design stuff.
Samarth Shah:
I guess, yeah. I think over the period of time, I've learned from my leaders and how they, and what worked and what didn't work. And you know, you have your do's and don'ts to this.
David Joy:
Yeah, that's very cool. You know, I'm really stoked about some of the stuff I learned from you actually, because I never thought about it like that, especially what you said, adding monitoring controls along with each of the layers to add time because you want to be able to respond, right? It's a pretty, pretty, pretty good stuff.
You know what, you and I can start geeking out on this topic and keep going on for two, three hours.
Samarth Shah:
Yeah.
David Joy:
I want to be respectful of your vacation, you know. So, Samarth, I mean, this was really awesome. Where can people connect with you or follow you or know some amazing things that you are doing?
Samarth Shah:
So I will say I'm a pickleball fan. So you can always find me on pickleball halls. But anyway, jokes apart. Yeah, I'm very active on LinkedIn. And people can always reach out to me on LinkedIn. And I'm very active. Like I said, I go to a lot of meetups, I meet a lot of people. I love connecting with people. Because that's where you grow. And that's how I honestly keep up with my industry. And there's only so much you can read online. And I mean, if you don't even know a language that exists out there, how are you going to Google or learn about it? So you have to know what's going on. And then you know, this is what I need to look into it.
David Joy:
Yeah, I agree. I think I have the same philosophy. I mean, I've noticed that I learn more by talking to people who've had a bad day.
Samarth Shah:
Good point.
David Joy:
And I... with somebody who's having [inaudible 00:55:46], you know. As I go, like, I go talk to my wife and I, "How's your day?" She's like, "Good day." I'm like, "I'm flying back up." You know.
Samarth Shah:
[inaudible 00:55:53] after that.
David Joy:
Yeah, like [inaudible 00:55:56]. So that's like my philosophy. I generally check with people as to like, "It was really bad." I'm like, "What happened?" And then they tell me and you learn so much from people's experiences, right? And that's what this podcast is all about.
And thank you everybody for tuning in. And thank you so much for coming. This is an absolutely fun episode. I hope you enjoyed and thank you once again for coming.
Samarth Shah:
Absolutely. I hope to see you soon.
A podcast for architects and engineers who are building modern, data-intensive applications and systems. In each weekly episode, an innovator joins host David Joy to share useful insights from their experiences building reliable, scalable, maintainable systems.

David Joy
Host, Big Ideas in App Architecture
Cockroach Labs
Latest episodes

Introducing Cockroach Continuum | A Big Ideas in App Architecture Exclusive
Tara Shankar Jana "TJ"
Senior Director Product Marketing @ Cockroach Labs

A Love Letter to the Database: Industry Shifts, Lessons Learned, and What's Next with Perry Krug
Perry Krug
Manager Solutions Architecture at Baseten

The Everything Trap: Building AI Software That Lasts with Sam Hilsman
Sam Hilsman
Co-founder and CEO of CloudFruit

Why Inference Engineering Is the Next Big Role in AI with Philip Kiely
Philip Kiely
Author of Inference Engineering | AI Education @ Baseten

Distributed Systems, Linkerd, and the Cost of Network Calls with William Morgan from Buoyant
William Morgan
CEO @ Buoyant, creators of Linkerd

Making Software as Durable as Data with Peter Kraft from DBOS
Peter Kraft
co-founder of DBOS

Breaking the Pillars: Rethinking Observability with Charity Majors
Charity Majors
Co-founder and CTO of Honeycomb.io and co-author of Observability

How to Transform Dev Workflows with CI/CS and AI Agents with Tomer Karin
Tomer Karin
Embedded Software Architect

AI, Market Cycles, and the Systems Built to Outlast Them with Cockroach Labs CEO & Co-founder Spencer Kimball
Spencer Kimball
CEO & Co-founder Cockroach Labs

How to Scale Data Infrastructure from Startup to Enterprise
Nishant Raman
Data Engineer at FinTech Company

How to Build an AI-Native Organization
Peter Mattis
Co-founder and CTO/CPO at Cockroach Labs

Inside Infrastructure as Code with Pulumi’s Founder & CEO
Joe Duffy
Founder/CEO at Pulumi

Inside Ericsson: How AI and Automation Are Shaping Telecom
Anand Bajaj
Chief Architect - 5G Network Slicing at Ericsson

Unboxing the Cloud: AI, Microservices, and Resilient Databases
Jim Hatcher
Solution Engineer at Cockroach Labs

Strategic AI and Cloud Solutions: GitHub’s Blueprint for Modern Development Success
Ari LiVigni
Senior Cloud Solutions Architect at GitHub

Cloud Architecture in the Public Sector: Balancing Innovation and Security
Nick Mayer
Principal Cloud Architect at Maximus

GenAI Meets Celebrity: Inside Cameo’s Journey from Startup to Stardom
Dom Scandinaro
CTO at Cameo

The journey from mainframe to adopting generative AI with Equifax’s Senior Network Architect
Samarth Shah
Senior Network Architect at Equifax

Modernizing your cloud strategy with OneStream’s Senior VP of Cloud Architecture
Ryan Berry
Senior VP Cloud Architecture at OneStream Software

Driving digital transformation with Chief Architect at Altimetrik, Ignacio Segovia
Ignacio Segovia
Chief Architect at Altimetrik

Discussing the Patterns of Distributed Systems with Unmesh Joshi
Unmesh Joshi
Principal Consultant at Thoughtworks and Author of Patterns of Distributed Systems

How to simplify your software architecture
Rob Reid
Technical Evangelist at Cockroach Labs

Behind the scenes with Vimeo’s Director of Enterprise Architecture
Sachin Joshi
Director of Enterprise Architecture at Vimeo

How to leverage real-time data processing for enterprises
Andrew Sellers
Head of Technology Strategy at Confluent

Inside the Mind of the Chief Architect at Index Exchange
Joshua Prismon
Chief Architect at Index Exchange

Solving for Scale: Real-time Retail Experiences with Endear's CTO
JP Grace
Endear

Data, Acquisitions, and AI: Insights from FiscalNote's CTO
Vlad Eidelman
CTO and Chief Scientist at FiscalNote

Discussing Data Trends in the AI Era
Gajanan Chinchwadkar
CTO at Hypermode

Unwrapping Moonpig: Architectural Insights into Personalization and Scalability
Alexis Lowe
Principal Engineer at Moonpig

Solving for data intelligence at scale
Madalina Tansie
Chief Technology Officer at Collibra

Simplifying solutions architecture with Brian Johnson of Booz Allen Hamilton
Brian Johnson
Sr. Solutions Architect at Booz Allen Hamilton

How to make your applications smarter
Rod Senra
VP of Engineering at Loadsmart

Scaling for 2 billion events per day with Principal Software Engineer at Red Ventures
Majid Fatemian
Principal Software Engineer, Data Platform at Red Ventures

The data behind digital marketing: A conversation with Bluecore’s Software Architect
Mike Hurwitz
Software Architect at Bluecore

A Lesson in Scaling: How Kami handled 25x growth with CTO and Co-Founder Jordan Thoms
Jordan Thoms
CTO & Co-Founder at Kami

Mastering Multi-Cloud with PwC’s Erol Kavas
Erol Kavas
Director at PwC Canada

From FedEx to Five Guys: Designing digital experiences with Yext’s VP of Software Engineering
Matt Bowman
VP of Software Engineering at Yext

Reliability and scalability in a data-driven world with Fivetran’s VP of Platform Engineering
Mike Gordon
VP of Platform Engineering at Fivetran

Enabling a data-driven and innovative engineering culture at Amplitude
Shadi Rostami
SVP of Engineering at Amplitude

How Estée Lauder scales strong engineering culture
Meg Adams
Executive Director of Platform Engineering at Estée Lauder

Can I take your order? Building conversational AI to improve the customer experience
Akshay Kayastha
Senior Engineering Manager at ConverseNow

Engineering resilient systems: Rescuing old treasures and unleashing modern capabilities
Marianne Bellotti
Author, Engineering Leader, Systems Geek

The Full Package: How Route architects its all-in-one post-purchase platform
Siddhartha Sandhu
Engineering Manager at Route

A historical journey in developer technologies
Mike Willbanks
CTO at Spark Labs

From Legacy to Cloud: Success stories from migrating mission-critical applications
Kishore Koduri
Senior Director of Enterprise Architecture at Ameren

Building purpose-driven engineering cultures
Jason Valentino
Head of Engineering Enablement at BNY Mellon

Modernizing Insurance Application Architecture at New York Life
Mike Murphy
Corporate Vice President and Life Insurance Domain Architect at New York Life

Innovation and Disruption: How Materialize pioneered a new era in data streaming
Arjun Narayan
Co-Founder and CEO at Materialize

Stories from an SRE: How Hans Knecht builds better developer experiences
Hans Knecht
Cloud Consultant at Knechtions Consulting (Ex: Capital One; Ex: Mission Lane)

Inside Chick-fil-A’s infrastructure recipe for a perfect customer experience
Brian Chambers
Chief Architect at Chick-fil-A Corporate

Modernizing from the Mainframe: An Exploration of Distributed Systems
Chris Stura
Director, PwC UK

IoT Standards & Data Mesh: Utility Facility App Architecture
Grant Muller
Vice President, Applications and Technology Architecture at Xylem

Relational Data Problems: Doubble Dating Application Architecture
Mattias Siø Fjellvang
CTO & Co-Founder at Doubble

From Legacy Systems to Limitless Scaling with Paycor’s Systems Engineering Fellow
Adam Koch
Systems Engineering Fellow at Paycor

How to Understand Problems & Build Better Software with Technical Leader Joe Lynch
Joe Lynch
Technical Leader

Observability in the Cloud & Dataflow Modifications with Yolanda Davis from Cloudera
Yolanda Davis
Principal Software Engineer, Data Flow Operations

Early Days at Google & Building CockroachDB with Peter Mattis
Peter Mattis
Co-Founder and CTO of Cockroach Labs

Database Benchmarking Efficiency with OtterTune’s Andy Pavlo
Andy Pavlo
Associate Professor of Databaseology at Carnegie Mellon and Co-Founder at OtterTune

Observability & Statelessness with TripleLift’s Chief Architect
Dan Goldin
Chief Architect at TripleLift

Understanding AI: PubNub CTO Stephen Blum’s Key to Faster App Development
Stephen Blum
PubNub

Building reliable systems with DoorDash's Matt Ranney
Matt Ranney
DoorDash

Real-Time Data Capturing: The Future of Fitness Technology
Paul Lawler
Head of Software at Wahoo Fitness

Building Efficient App Architecture with Alloy Automation’s Gregg Mojica
Gregg Mojica
Co-Founder and CTO Alloy Automation

Unleashing the Power of Hiring Software with Greenhouse CTO Mike Boufford
Mike Boufford
CTO at Greenhouse Software

Decoding Data Warehousing: Insights from Ken Pickering, SVP of Engineering at Starburst Data
Ken Pickering
Senior Vice President of Engineering, at Starburst Data